Run Claude Code for free on your own machine
If you've been watching your API bill grow every time you use Claude Code, this is worth ten minutes of your time.
Ollama, the tool for running open-source language models locally, added full compatibility with the Anthropic Messages API in version 0.14.0 (released January 16, 2026). That single change opened a door most developers didn't expect: you can now point Claude Code at a local model running on your own machine, skip the API entirely, and pay nothing.
What actually changed
Why this works: Claude Code talks to Anthropic's API using a standard message format. Ollama now speaks that exact same format locally. Result: Claude Code thinks it's talking to Anthropic. It's actually talking to a model running on your own hardware.
Ollama handles the heavy lifting: model downloads, memory management, GPU acceleration. The Anthropic API compatibility layer is the new piece that makes the Claude Code connection possible.
What Ollama supports through this bridge includes streaming responses, tool calling, system prompts, image input, and extended thinking—which covers most of what Claude Code actually needs to function properly.
The setup (takes about 10 minutes)
Step 1: Install Ollama
Download and install Ollama from ollama.com. It runs on Mac, Windows, and Linux.
Step 2: Pull a local coding model
Open your terminal and run one of these based on your machine's RAM:
# 8GB RAM (lightweight start)
ollama pull qwen2.5-coder:7b
# 16GB RAM (best balance)
ollama pull deepseek-coder-v2:16b
# 32GB+ RAM (highest quality)
ollama pull qwen2.5-coder:32b
Qwen 2.5 Coder 32B scores highest on coding benchmarks among local models, making it the top pick if your hardware can handle it. DeepSeek Coder V2 16B is a strong middle ground, delivering most of that performance at half the RAM requirement. For more on the strategic advantages of compact architectures, read our deep dive on why small AI models are the real gold rush.
Step 3: Install Claude Code
# Mac / Linux
curl -fsSL https://claude.ai/install.sh | sh
# Windows (PowerShell)
irm https://claude.ai/install.ps1 | iex
Check the official Anthropic Claude Code documentation for the latest CLI version and system prerequisites.
Step 4: Connect Claude Code to Ollama
This is the key step. Set two environment variables to redirect Claude Code away from Anthropic's servers and toward your local Ollama instance:
# Mac / Linux
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_BASE_URL=http://localhost:11434
# Windows (PowerShell)
$env:ANTHROPIC_AUTH_TOKEN = "ollama"
$env:ANTHROPIC_BASE_URL = "http://localhost:11434"
Step 5: Launch it
Navigate to your project folder in the terminal and run:
claude --model qwen2.5-coder:7b --dangerously-skip-permissions
Note: the --dangerously-skip-permissions flag turns off the repetitive permission pop-ups inside Claude Code so it can work through file edits and tests without interrupting you. Always review actions in active Git repositories.
That's it. You now have a fully local coding agent with file access, terminal commands, and multi-step task handling.
What you should know before going all-in
This setup is genuinely useful, but it's worth being honest about the trade-offs.
| Factor | Cloud Claude | Local (Ollama) |
|---|---|---|
| Cost | Paid API ($20/mo Pro or usage fees) | Free ($0 API bills) |
| Privacy | Cloud-based (Anthropic servers) | 100% local (Never leaves disk) |
| Model quality | Frontier reasoning (Claude 3.5 Sonnet) | Good (Hardware & parameter-dependent) |
| Speed | Consistent server latency | Varies by GPU/RAM bandwidth |
Local models are strong for day-to-day tasks: writing functions, refactoring code, explaining errors, and integrating public endpoints like those in our 320,000 free APIs developer directory. Where they fall short is reasoning through genuinely complex architecture decisions or handling very long context windows.
Think of this as a free, always-available coding companion for everyday work, with cloud Claude as the option you reach for on harder problems.
A practical starting point
If you have 16GB of RAM and want the most reliable experience right now, pull deepseek-coder-v2:16b. It runs on most modern laptops, handles the majority of coding tasks cleanly, and gives you a real sense of what this setup can do before you commit to a larger model.
For anyone working on a codebase where privacy matters—a client project, an internal business tool, or a bespoke publishing platform like the ones we design at Builder Hustle Studio—this setup makes sense even beyond the cost angle. Your code stays strictly local, air-gapped from cloud data retention policies.
A quick reality check: your setup is only as safe as the machine you run it on. If your laptop is already updated, backed up, and free of anything suspicious, running a local coding stack like this doesn't introduce some new exotic risk. The tools live on your hardware, your code stays on your disk, and nothing leaves your device unless you deliberately connect it to a remote service. Treat it the same way you treat your editor and terminal: keep your system secure, and it will quietly do its job in the background.