Docs · Agent

SWE workflow

Disciplined bug fixing: reproduce the failure before touching source, localize to the smallest scope, make the smallest edit, verify, then check for regressions. Available as sentinel swe with a 60-iteration budget and verbatim tool history.

The five phases

  1. Reproduce: run the failing tests with runTests first. If you cannot reproduce, stop, do not guess.
  2. Localize: codeMap for orientation, then grep + read. Read the test file first; it is the specification.
  3. Fix: smallest edit that addresses the root cause. Never rewrite files, never edit tests to make them pass.
  4. Verify: re-run the repro and failing tests. All must pass.
  5. Regress: run the related passing suite. On regression, undoLastChange and retry.

Offline capability gates

bash
sentinel bench
# 15/15 checks: roundtrips, mode refusals, sandbox, atomic edits,
# structured test parsing, patch application, undo, FAIL_TO_PASS gates

Task evals

bash
npm run eval:check              # validate fixtures + graders (CI-gated)
node evals/run.mjs --agent --model gpt-6-luna   # real agent runs + report

Every task ships a pristine fixture (must fail), a reference solution (must pass), and a Node grader, so results are reproducible on Linux, macOS, and Windows.

Honest benchmarking

Bench gates measure tool capability, not model reasoning. Real SWE-bench % scores require the official Docker harness plus a model key, report them with harness version, model id, and retries, never as self-awarded numbers.

Harness guide →