10 min read

Cut your LLM bill in half: cost control for coding agents

Where coding agent spend actually goes, and the seven changes that reduce it most: free-tier defaults, hard context caps, persisted budgets, and a verifier.

  • Cost
  • Performance
  • Operations

Where the money actually goes

Teams audit their model choice first, which is the wrong end of the problem. Here is the order of magnitude of what a typical agent turn spends tokens on:

  • Tool output re-read every turn. A 40k-character grep result stays in context for the rest of the loop and is re-billed on each call. This is usually the largest single line item.
  • Cumulative file contents. Ten files read at 5k characters each is 50k characters of context, resent on every subsequent turn.
  • Failed attempts. A non-converging loop is the most expensive failure mode there is, because each retry carries the full history.
  • The answer itself. Genuinely the smallest part, and the part everyone optimises.

Once you accept that ordering, most “use a cheaper model” advice is a rounding error, and the real wins are all about how much context you admit and how many turns you take.

Seven changes, ranked by leverage

1. Route by task, not by preference

Most teams run one model for everything because choosing per call felt like overhead. It is not: a free-tier model handles listing, searching, explaining, formatting and mechanical edits, which is the large majority of calls in a working session.

terminal
# free default, on every new machine
sentinel

# frontier, when the task is actually hard
sentinel ask -m claude-sonnet-4 "why does this deadlock under load?"

The deeper version of this is a standing loop that wakes on every failed test. That job is perfect for a cheap model and terrible for a frontier one, because it runs unattended and often.

2. Truncate tool output hard, and do it in one place

A cap enforced at the tool boundary is worth more than any prompt instruction asking the model to be brief, because the model cannot see what it was never sent. Sentinel truncates tool results at 20k characters, and 30k in SWE mode, before they reach context. If you are writing your own agent, this is a five-line change with a disproportionate effect.

The corollary matters more: truncate from the right end, not the left. The first few kilobytes of a grep result carry the matches; the tail is usually the same file repeated. Truncating the head is how you get an agent that confidently edits the wrong part of a file.

3. Orient with a code map, not a file dump

Reading six files to answer “where is the retry logic” costs six files of context, every turn, forever. A symbol-level overview (functions, classes and exports per file) is a fraction of the size and is usually enough to pick the one file worth reading. This is the cheapest quality-per-token win available to an agent.

4. Compact on a threshold, not at the edge

Let context run to the model’s limit and you are paying full price for the last turn before a failure. Compact at a fixed fraction (40k in Sentinel’s case), and the expensive tail never happens.

Watch for the compaction loop bug

A compaction that re-triggers itself on the message it just produced will loop until the iteration cap, burning the whole budget. If your spend jumps on long sessions specifically, check this before anything else.

5. Cap iterations and make them mean something

An unbounded loop is a budget you did not set. A cap is necessary but not sufficient, a loop that hits its cap having accomplished nothing has wasted the maximum. Pair the cap with a progress signal and a backoff: an agent that has not written anything and has not met the goal has not moved, and should wait rather than retry immediately.

bash
# 60 iterations, not infinity
sentinel swe

# a ceiling that outlives the process
sentinel budget --usd 25 --deadline 2h --condition "npm test exits 0"
sentinel budget   # active  ████░░░░░░  $12.40 of $25.00 (50%) · 1h 59m left

6. Give the agent a verifier instead of a conversation

Every turn spent asking “is this right yet?” is a turn you pay for. A failing test is a free, deterministic, unambiguous verifier, and it is why the reproduce-first SWE workflow exists: reproduce, fix, verify, regress. The agent stops negotiating with itself.

The same idea generalises. Any place you can replace “ask the model whether this is right” with a command that exits 0 or non-zero is a place you have deleted a turn from the budget.

7. Persist the budget where the team can see it

A per-turn cost cap is a safety rail. A per-project ceiling is an operating control, because it answers the only question a lead asks: what has this cost so far this week. Sentinel stores the ceiling and an append-only spend log in the project, so a run that crashes mid-write loses one row rather than the history.

Two details worth copying. Spend recorded before the budget existed should not count against it, or you cannot adopt a budget mid-engagement. And a turn finishing in the same millisecond the budget was created should count, a ceiling must fail toward charging you, not away.

Measure it or it does not happen

Cost work fails silently because nobody can attribute it. Three things make it stick:

  1. Print cost per turn, always. Not in a debug flag. If the number is one keystroke away, people make better decisions; if it needs a dashboard, they do not.
  2. Attribute to a run id, not a session. Sessions blur across days. Runs do not.
  3. Default the cheap model to be free. Cost control that depends on everyone remembering is not a control. If the safe choice is also the zero-cost choice, behaviour follows.
terminal
# every run ends with the receipt
sentinel ask "add a regression test for the retry path"
#   in 4,182 · out 611 · $0.0000 · openai/gpt-oss-20b · 3 turns · 11.4s

Frequently asked questions

Why is my coding agent so expensive?

Almost never because of the answers. It is the context: every file read, every grep result and every test log becomes input tokens on the next call, and input is re-billed on every subsequent turn of the loop. A single unbounded grep output can cost more than every message in the conversation combined. The other common cause is a loop that has stopped converging, an agent retrying the same failing edit burns a full turn's context each time.

What is the single highest-leverage cost change?

Defaulting to a free or near-free model. Teams routinely run a frontier model for every turn including the mechanical ones (listing files, grepping, formatting) when a small free model does the same job. Routing by task is the difference between a bill you watch and a bill you find out about on the invoice.

Should I set a hard spending limit?

Yes, at the project level rather than the session level. A per-turn cap cannot answer the only question a team lead asks, which is what has this cost so far this week. A persisted ceiling that every run honours, and that fails toward charging you rather than away, is what makes agent spend auditable.

Written by Kunj Shah

Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.