Docs · Agent
SWE workflow
Disciplined bug fixing: reproduce the failure before touching source, localize to the smallest scope, make the smallest edit, verify, then check for regressions. Available as sentinel swe with a 60-iteration budget and verbatim tool history.
The five phases
- Reproduce: run the failing tests with
runTestsfirst. If you cannot reproduce, stop, do not guess. - Localize:
codeMapfor orientation, then grep + read. Read the test file first; it is the specification. - Fix: smallest edit that addresses the root cause. Never rewrite files, never edit tests to make them pass.
- Verify: re-run the repro and failing tests. All must pass.
- Regress: run the related passing suite. On regression,
undoLastChangeand retry.
Offline capability gates
sentinel bench
# 15/15 checks: roundtrips, mode refusals, sandbox, atomic edits,
# structured test parsing, patch application, undo, FAIL_TO_PASS gatesTask evals
npm run eval:check # validate fixtures + graders (CI-gated)
node evals/run.mjs --agent --model gpt-6-luna # real agent runs + reportEvery task ships a pristine fixture (must fail), a reference solution (must pass), and a Node grader, so results are reproducible on Linux, macOS, and Windows.
Honest benchmarking
Bench gates measure tool capability, not model reasoning. Real SWE-bench % scores require the official Docker harness plus a model key, report them with harness version, model id, and retries, never as self-awarded numbers.