Run a coding agent on local models with Ollama or LM Studio
A practical guide to running an AI coding agent on local models: which model sizes work for which jobs, how to route between local and hosted, and what still breaks.
- Local models
- Privacy
- Ollama
What “local” actually buys you
Open source is often sold as the privacy property, but it is only half of it. An open source agent that calls a hosted model still sends your prompt, the contents of every file it reads, and its tool output to that provider's API. Auditing the agent tells you what the agent does with that data. It says nothing about the model provider.
Running the model on your own hardware is what closes the loop. With Ollama or LM Studio bound to localhost, the request never leaves the machine, which matters when:
- Your code cannot leave your infrastructure. Not a policy promise from a vendor, not a zero-retention setting with an expiry, a network fact you can verify with
lsof. - You are on a plane, on a train, or air-gapped. The agent keeps working. This is a bigger deal than it sounds.
- Marginal cost is zero. Every extra local turn is free, so the tool that wakes on every failed test is affordable.
- No rate limits. This is what makes an unattended loop possible at all.
Setup: two commands, no API key
# 1. install and start Ollama (default host http://localhost:11434)
ollama serve
# 2. pull a coding-capable model
ollama pull qwen2.5-coder:14b
# 3. run the agent, no API key of any kind
cd SENTINEL-CLI && npm install && npm link
sentinelSentinel auto-discovers models from a local runtime on startup, so there is nothing to configure. LM Studio works the same way. If your runtime is on a different port, set OLLAMA_HOST or LMSTUDIO_HOST.
export OLLAMA_HOST=http://localhost:11434
# confirm what the agent actually sees
sentinel --version
sentinel ask "list the entry points in src/agent"Model sizing: match the model to the job
The mistake is picking one model and expecting it to do everything. Local models differ enormously by size, and the difference shows up in a specific place: how many tool calls in a row they can complete without losing the thread.
| Class | Good at | Weak at | Context |
|---|---|---|---|
| 7–8B | Search, explain, review, one-file edits | Multi-step changes, novel abstractions | 8–32k |
| 13–14B | Most refactors, test writing, migrations | Long autonomous runs without checkpoints | 32k |
| 30–34B | Complex debugging, architectural edits | Raw throughput; needs the VRAM | 32–128k |
| 70B+ | Reasoning-heavy work | Everything about latency and memory | 32–128k |
The column people underrate is context. Coding agents live on how much of the repo they can hold at once, and local models generally ship with shorter context windows than hosted frontier models. Past roughly 30B parameters, the wall you hit is your context budget rather than your VRAM. Which is why compaction, read-only code maps and hard output caps exist as features rather than as niceties.
Tool calling is the skill that matters
Benchmark charts measure prose quality. Agents need something narrower and stranger: emitting a syntactically valid tool call, with the right arguments, in the right order, for twenty consecutive steps without inventing a parameter. That is a different axis, and a coder-tuned model at 14B can beat a general model twice its size.
Two practical consequences. First, prefer a model explicitly tuned for tool use over a general chat model. Second, give it fewer, wider tools, every additional tool is another chance to pick the wrong one.
The answer most people land on: hybrid routing
Pure-local and pure-hosted are both wrong for real work. The useful setup is a default local model for the mechanical majority, and an explicit escape hatch to a hosted model for the one hard problem.
| Task | Route | Why |
|---|---|---|
| Find the caller of X | Local | High volume, zero risk, mechanical |
| Explain this module | Local | Read-and-summarise is a local model strength |
| Review a diff | Local | Pattern matching over structure, not novel reasoning |
| Fix one failing test | Local or small hosted | Bounded, has a verifier |
| Refactor across 6 files | Hosted frontier | Long-horizon planning is where small models fail |
| Diagnose a production incident | Hosted frontier | The cost of being wrong is not symmetric |
This is cheap to do because the model is a runtime choice, not a build-time decision. In Sentinel you switch with /model inside the TUI, or per call with --model on the CLI. Set a project budget and a standing loop will not quietly spend your afternoon on a hosted model.
# local by default, with a ceiling that survives the process
sentinel budget --usd 5 --deadline 4h
# mechanical, on hardware you own
sentinel ask "which files import the deprecated logger?"
# hard, on a frontier model, still inside the budget
sentinel ask -m claude-sonnet-4 "why does this deadlock only under load?"What still breaks, and what to do about it
The model stops emitting tool calls
Symptom: the agent answers in prose instead of acting. Fix: switch to a tool-tuned model, and reduce the number of tools in the active mode.
Context exhaustion mid-task
Symptom: the agent starts re-reading files it already read, or contradicts itself. Fix: keep compaction on, and use read-only code maps instead of whole-file reads to orient.
Silent truncation on big files
Symptom: the fix is applied to the top half of a file. Fix: prefer exact-match edits over whole-file writes, and lean on the diff tool to preview before applying.
Hallucinated paths
Symptom: it claims to have edited a file that does not exist. Fix: this is what the sandbox is for, an out-of-root path is rejected rather than created, so the failure is loud.
On the local-routing claim
Frequently asked questions
Can an AI coding agent run fully offline?
Yes, if the model is local. With Ollama or LM Studio running on the same machine, the agent, the model and the tools all execute locally and no request leaves your network. The agent binary being open source is not sufficient on its own, a local model is what actually makes the run private, because the model call is where your source code would otherwise be sent.
What is the best local model size for coding?
For reading, searching, explaining and reviewing, a 7B to 8B model is genuinely usable. For multi-file edits and long-horizon reasoning, expect to want 20B and up, or a quantised 30B-class model if you have the VRAM. Beyond roughly 30B parameters the practical limit stops being the model and becomes your context window: local models generally have shorter context windows, and coding agents live or die on how much of the repo they can hold at once.
Are local models cheaper than hosted APIs?
Per token, no. A local token costs electricity and hardware you already own. Per engineering hour, often yes, if the tasks you route locally are the high-volume mechanical ones. Local inference is free at the margin, which makes it a good fit for a standing loop that wakes on every failed test, and a poor fit for the one hard reasoning problem you need to get right.
Written by Kunj Shah
Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.