How to evaluate a coding agent you are building
Capability gates, FAIL_TO_PASS fixtures, deterministic graders and reproducible task evals, how to measure an agent without fooling yourself.
- Evaluation
- Testing
- SWE
Why bother, when it obviously works
Because “it obviously works” is how agents get worse without anyone noticing. The loop is full of small changes that each look safe: a prompt tweak, a new tool, a summarisation step, a higher iteration cap. Every one of them can silently halve your success rate on the hard cases while the demos keep looking great.
The uncomfortable version: an agent improvement is a hypothesis, and without an eval it is an anecdote. The best return on a small team’s time is not a bigger benchmark, it is a fast, boring, deterministic one that runs in under a minute and never needs a model key.
Layer one: capability gates, no model required
These test the harness. Can the agent loop actually call a tool, respect a mode, sandbox a path, parse a test result? They should be fast, hermetic and run in CI on every commit, because they catch the class of bug that masquerades as intelligence.
sentinel bench
# 15/15 checks: roundtrips, mode refusals, sandbox, atomic edits,
# structured test parsing, patch application, undo, FAIL_TO_PASS gatesThe checks worth having, in rough priority order:
- Mode refusals. PLAN mode must refuse a write tool. If it does not, everything else you believe about your permissions is fiction.
- Sandbox escape. Traversal, absolute paths outside the root, and a symlink pointing out of the tree must all be rejected, see the permission model.
- Atomic batch edits. A batch that fails halfway must leave nothing applied, and must not leave a stale read in the next operation.
- Undo across turns. Write in turn one, undo in turn two, assert the original content. An implementation that only undoes within the current turn passes a naive test and fails a real one.
- Structured test parsing. Feed it jest, pytest, mocha and TAP output. If the parser silently returns “no failures” on a format it does not recognise, your eval is measuring nothing.
The last one is the dangerous one
Layer two: task evals that can be trusted
Now the agent has to actually solve something. Three artefacts per task, and all three are required.
The fixture must fail first
A pristine copy of the repo state, plus a test that fails on it. This is the FAIL_TO_PASS gate, and it is the single highest-value check in the whole system, because it eliminates the two false positives that make agent benchmarks meaningless:
- The task was already solved. Nothing to fix, agent changes nothing, reported as a pass.
- The test never exercised the bug. Agent breaks something unrelated, the untouched test still passes.
The reference solution must pass
Someone has to write the fix. It does not need to be the best fix, it needs to exist, so you know the task is solvable and your grader is satisfiable. A task where the reference solution fails is a broken task, and shipping it teaches your agent the wrong lesson.
The grader must be mechanical
A Node script that exits 0 or non-zero. No model in the loop, no rubric, no “does this look like a reasonable fix”. If a model grades the model, you have built a vibe check, and vibe checks drift silently in the direction that flatters you.
npm run eval:check # validate fixtures + graders, CI-gated
node evals/run.mjs --agent --model gpt-6-luna # real agent runs + reportThe loop that keeps it honest
- Reproduce before you touch anything. Run the failing test first. If you cannot reproduce it, stop, the agent must not guess, and neither should the task. This is the discipline the SWE workflow encodes, and it is also the discipline your eval tasks require.
- Localise to the smallest scope. A symbol-level code map first, then grep, then read the test. The test is the specification.
- Make the smallest edit that addresses the root cause.Never rewrite a file, and never edit a test to make it pass. A grader that allows the second thing is not a grader.
- Verify, then check for regressions. The repro plus the related passing suite. A regression is an undo and a retry, not a patch on top.
Reporting numbers without lying
This is where self-published agent evals usually go wrong, and the fix is simple: publish the context or publish nothing.
- State the harness version. Results are not comparable across harness revisions, and the number moves.
- State the exact model id and the retry count. A pass rate with three retries is a different measurement from one with none.
- Separate capability from reasoning. “Our tool layer passes 15/15 offline gates” is a true, useful, narrow claim. It is not “our agent solves 42% of SWE-bench”.
- Keep the failing tasks. A suite that only reports wins is a marketing asset. Report the ratio and the distribution.
The rule
Frequently asked questions
How do you benchmark a coding agent you built yourself?
In two separate layers, because they fail for different reasons. Capability gates test whether the harness can do things at all, a round trip, a refused write in the wrong mode, a traversal blocked, an undo that restores. Task evals test whether the agent solves problems. Keep them separate, because a capability failure masquerading as a reasoning failure will send you optimising a prompt when the bug is in your tool result parser.
What is a FAIL_TO_PASS test?
A test that must fail before the fix and pass after it. It is the core primitive of a trustworthy agent eval, because it removes the two classic false positives: a task that was already solved, and a test that never actually exercised the bug. Every eval task should ship one, alongside a fixture that reproduces the failure and a grader that decides pass or fail mechanically.
Should I report SWE-bench scores for my own agent?
Only with the official harness version, the exact model id, the retry count and the Docker configuration attached. Self-awarded percentages without that context are marketing, not measurement, the numbers move with the harness, the model and the retry budget, and readers have no way to tell which. Report the harness, the model and the retries, or report nothing.
Written by Kunj Shah
Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.