What if modern AI could run on a computer from 1997?

You'd probably laugh. I did. Then EXO Labs actually did it.

They took a Pentium II from the late '90s (total eBay cost: £118.88), loaded a Llama 2 AI model onto it, and watched it generate coherent responses at 39 tokens per second. This says something important: the future of AI doesn't actually need billion-dollar data centers.

How they actually did it

128 MB total RAM available  •  39 tokens/sec on the 260K-parameter model  •  1 second to respond on the 15M-parameter model

This was EXO Labs, led by AI researcher Andrej Karpathy. You might know him from Tesla, where he led the Autopilot team, or OpenAI, where he was a founding member.

The team called their approach “llama98.c.” It's based on Karpathy's llama2.c, which is 700 lines of pure C code that handles AI inference on minimal hardware. They adapted it for Windows 98, compiled it with Borland C++ 5.02 (from 1999), and made it work.

The hard part wasn't running the model. It was the infrastructure: getting files onto a machine with no USB support. They used FTP. The keyboard had to go in port 2, the mouse in port 1. Everything was broken in ways that required real engineering to fix. But they fixed it, and it worked.

Why this actually matters

You could dismiss this as a stunt—a fun demo that proves nothing. Except EXO Labs is pointing to something that's about to reshape AI infrastructure: BitNet.

BitNet is a transformer architecture that uses ternary weights instead of floating-point numbers. Every parameter in the model is one of three values: negative one, zero, or one.

Here's the impact: a 7-billion-parameter BitNet model needs only 1.38 GB of storage. Compare that to a standard 7B model that needs 13 GB. It runs on CPUs instead of GPUs, and it's roughly 50% more efficient than full-precision models—while running at near full-precision accuracy on language tasks.

The math is striking. BitNet b1.58 delivers 2.71 times faster inference and 3.55 times less memory usage than standard FP16 models. Energy consumption drops by 55 to 82% per token depending on hardware. A 30-billion-parameter BitNet model has similar performance to a standard 7-billion-parameter model, but with dramatically lower resource requirements.

What this means in practice

Your device gets smarter. Your laptop stops being slow. A 100-billion-parameter BitNet model runs at normal conversation speeds (5–7 tokens per second) on a single CPU. No waiting, no lag.

Your data stays put. Everything runs locally. No uploading, no cloud processing. Your files, your emails, your work stays on your machine—faster, simpler, more secure.

Hardware you already own gets better. That 2018 MacBook or old phone can now run AI locally. No need to buy anything new; existing devices just work better.

The edge computing shift

This connects to something bigger happening in 2026. Hardware companies are no longer racing on clock speeds. They're racing on efficiency—specialized AI accelerators, neuromorphic chips, quantum co-processors. The goal: enable devices to run sophisticated models locally without relying on remote data centers. That's the shift, from cloud-centric AI to device-centric AI.

Meta is building this into consumer hardware. Google is embedding it in Pixel phones. Apple is running models locally on iPhones. And OpenAI is reportedly launching its first consumer hardware device in late 2026, designed by Jony Ive, that emphasizes local inference over cloud-dependent operations.

The Windows 98 demo is a proof of concept. The actual play is happening on billions of devices that already exist: your phone, your laptop, your tablet, your old MacBook gathering dust. Each one becomes capable of running AI locally.

The EXO Labs vision

The team's stated goal is to build open infrastructure to train frontier models and enable any human to run them anywhere. That's not just engineering, it's philosophy: the alternative is a future where AI runs exclusively in mega data centers owned by a handful of companies.

EXO Labs runs a Discord community with a “Retro” channel where people experiment with running LLMs on old Macs, Game Boys, Raspberry Pis, and other limited devices.

What you should know

The Windows 98 demo works, but it's glacially slow on larger models—a 1-billion-parameter model runs at only 0.0093 tokens per second. That's not practical on its own.

But BitNet changes the equation. With ternary weights, smaller models perform like much larger ones. You don't need billion-parameter models running locally; you need purpose-built, fine-tuned models. This is the real edge AI opportunity: not running giant models on your phone, but running the right model for the specific task on any device.