Part 12 of 12 · Build a Cursor-style AI coding agent in your terminal

12 min read

The cautious tech lead asks three questions. Answer those.

How to answer a sceptical tech lead about running a terminal coding agent on a production repo, the questions they actually ask, and what you can honestly offer instead of a demo.

  • Adoption
  • Architecture
  • Engineering Management

Why you are having this conversation

Someone on your team is already using one of these. They are not using it because they read a blog post; they are using it because it answered a question about an unfamiliar service in thirty seconds that would otherwise have cost them an afternoon of reading. That usage is going to appear in your repo whether or not you approve it, and “prohibit it” is not an outcome available to you.

So the real decision is not whether your organisation adopts a terminal coding agent. It is whether that adoption is governed or accidental. Those are very different positions to be in eighteen months from now, and choosing between them is cheap this month.

The honest framing for the opening

“I want to show you something and then argue against it, because there are three reasons not to adopt this and I think two of them are good ones. If you disagree with my ranking, that’s the useful conversation.”

Question one: what can it actually do to my repository?

This is not a soft question, so do not answer it with reassurance. Answer it with the mode table and let them interrogate it.

src/shared/schemas/mode.js
export function isToolAllowedInMode(toolName, mode) {
  if (mode === Mode.BUILD || mode === Mode.SWE) return true;

  if (mode === Mode.PLAN || mode === Mode.REVIEW || mode === Mode.SCAN) {
    return isReadOnlyTool(toolName) || toolName === 'diffFile';
  }

  if (mode === Mode.FIX) {
    // FIX mode: read + write tools, but no shell
    return toolName !== 'bash' && toolName !== 'runTests';
  }

  return true;
}

Then hand them the one command that makes it checkable, because the answer to “can I trust this” is a query they can run themselves:

terminal
$ node -e "
  const { isToolAllowedInMode } = await import('./src/shared/schemas/mode.js');
  const tools = ['readFile','grep','writeFile','editFile','bash','runTests'];
  console.log('tool'.padEnd(12), ['PLAN','FIX','BUILD'].map(m=>m.padEnd(6)).join(''));
  for (const t of tools) {
    console.log(t.padEnd(12),
      ['PLAN','FIX','BUILD'].map(m => (isToolAllowedInMode(t,m)?'yes':'no ').padEnd(6)).join(''));
  }
"
tool         PLAN   FIX    BUILD
readFile     yes    yes    yes
grep         yes    yes    yes
writeFile    no     yes    yes
editFile     no     yes    yes
bash         no     no     yes
runTests     no     no     yes

Thirty lines of source, printed on demand, no documentation required. That is the whole answer to question one, and it lands better than any policy document because it is verifiable in the room.

Question two: what stops it running something harmful?

Three layers, and the order matters as much as the existence. Say the order out loud, because “there are checks” is not an answer while “the mode gate runs first, and it is the only one the model cannot influence” is.

  1. Mode gate. The only control the model has no influence over, because the loop dispatches the tool, not the model.
  2. Risk ledger. Grades the shape of each shell command against what this repository has already approved. A session grant for git commit never authorises git push --force.
  3. Blast-radius gate. Challenges the first write to a migration, CI workflow, lockfile, infrastructure or auth path once per turn, demanding a justifying file:line and an exact rollback.
  4. The interactive prompt, for anything not already proven safe here.

And the four commands they will actually want:

terminal
# what has this repo learned?
sentinel risk --list

# how does this specific command grade, right now?
sentinel risk "npm publish"
sentinel risk "git push --force origin main"

# what does the agent think it has already proven?
sentinel budget

Question three: what does it cost, and who pays?

The question that decides budgets, and the one every non-technical stakeholder asks first. The answer is that cost is a first-class, persisted, capped quantity — not a surprise on an invoice.

terminal
$ sentinel budget --usd 25 --deadline 2h --condition "npm test exits 0"
Engagement budget set in .sentinel/budget.json:
  budget     $25.00
  deadline   2026-10-16T18:00:00.000Z (1h 59m left)
  condition  npm test exits 0
  since      2026-10-16T16:00:00.000Z

$ sentinel budget
active  ████████░░░░░░░░░░░░  $6.12 of $25.00 (24%), 1h 58m left
  budget     $25.00
  remaining  $18.88
  spent      $6.1184 over 7 turn(s) this engagement
  lifetime   $41.02 over 23 turn(s) (.sentinel/spend.jsonl)

Two things to point out, because both are what turns a wary lead into a sponsor. First, the ceiling is enforced, not advisory: the loop stops hard when the budget is exhausted, and setting it once gates every later run. Second, the default model is a free tier, so most usage costs nothing at all — which is a fact you can demonstrate in thirty seconds rather than argue for.

Now the part that wins: argue against it

Give the three reasons not to adopt. Rank them honestly. This is the move that separates a proposal a lead trusts from one they tolerate, because a lead who finds the objection you buried is now checking all your other claims.

Three reasons not to adopt, ranked honestly
ReasonVerdictWhat to say
The model can be confidently wrongReal, and not solved by anything on this listThe controls reduce blast radius; they do not improve judgement. This is why the permission model is code and not a prompt. But do not oversell it
It will train people to accept unreviewed writesReal, and a team-level risk rather than a tool riskThe interactive prompt is per-tool, and the gate is per-path. The mitigation is reviewing diffs as a habit, which no tool can do for you
A tool with shell access on a production repo is a supply-chain riskLegitimate, and the one to resolve before deployingOffer FIX mode for unattended work (edits with no shell), and run the pilot on a repository with a working restore procedure

Rank your objections by how much they actually matter

If your first objection is the vendor and your third is that the model hallucinates, you have told the lead your priorities are the wrong way round. Most people rank them in reverse, because the first one is easier to say out loud.

The proposal that gets accepted

Not “let everyone try it”. A pilot with a success criterion and a kill switch, narrow enough to be obviously reversible.

the ask
Repository:   one service, restore-from-backup verified in the last 30 days
Mode:        PLAN only (read-only). No writes, so the failure mode is
             a wasted afternoon, not an incident.
Owner:       one named engineer, who is also the person who reports back
Duration:    two weeks
Measured:    (a) unfamiliar-codebase questions answered without help
             (b) hours spent on those questions before, from git history
             (c) zero writes, trivially true, and stated anyway so it is checked

Kill switch: delete .sentinel/ and remove the binary. No server, no database,
             no data to migrate out.

The sentence that does the most work

“There is no server and no database, so if this goes badly the cleanup is rm -rf ~/.sentinel.” It reframes the whole decision from “can we afford to be wrong” to “how quickly can we be wrong”, which is a much easier question to answer yes to.

Who should not adopt this

Being useful means disqualifying people, and this is where credibility is actually earned. If any of these are true, say so and stop:

  • Your security team requires centrally reviewed, logged agent access. This tool has no team dashboard, no SSO and no audit log. That is a platform product and you want one. Claiming otherwise is how you lose the room.
  • Nobody on the team scripts, and the whole team lives inside one editor. You will lose time before you gain any.
  • An API key on a laptop is a policy violation. That decision was made before you installed anything. A local model does not change it, and should not be offered as if it does.
  • The team wants it to work unattended on unreviewed changes, tonight. Not this. Not yet, and probably not this tool.

None of those make the tool bad. They make it a different product for a different team, and saying so out loud is what makes the rest of your argument credible.

The close

Three things, in this order:

  1. Read the mode table with them. Not the README, not the blog post. Thirty lines they can interrogate.
  2. Show the refusals. Ask it to fix something in read-only mode and let them watch it explain what it would do and decline. A tool you have watched say no is trusted differently from a tool you have been told says no.
  3. Name the pilot, not the rollout. One repository, one owner, read-only, two weeks, three measurements.

That is the whole argument, and it does not require you to claim the model is reliable. It requires you to show that every action passes through a boundary the team can read, that the cost is capped, and that turning it off is one command. Those are claims a sceptical engineer can verify in an afternoon, which is exactly what you want when you are asking someone to change how they work.

Frequently asked questions

What if the tech lead's objection is fundamentally about trust in vendors?

Then it is not an objection you should argue away, because it may be correct for your context, regulated data, a contractual no-training requirement, a procurement process that takes six months. The answer that works is to show what the agent looks like with a local model: point it at Ollama or LM Studio and no request leaves the machine. That does not satisfy a vendor objection about the model, but it satisfies the operational one, which is usually the real blocker, and it lets the conversation move to tooling where you have actual leverage.

How do I pitch this to someone who has been burned by an AI tool before?

Do not pitch it as AI. Pitch it as a permission system, because that is what they actually need to govern, and what they have no equivalent for today. The interesting claim is not 'the model is smart'. It is 'every action this tool takes passes through an allowlist I can print and read'. That framing sidesteps the hype objection entirely, and it happens to be the part of the system that is hardest to build well.

What if they want a pilot?

Good, agree immediately and make it concrete, because a pilot is the only thing that settles this argument. The mistake is proposing a time-boxed trial without a success criterion, which just delays the decision. Propose instead: one repository, one named owner, PLAN mode only for two weeks, with three things measured, questions answered without help, time saved on unfamiliar-codebase tasks, and zero writes outside the pilot repository. A pilot that cannot fail is not a pilot.

The comparison against AI IDEs and hosted agents is on the compare page, and the mode reference is in docs/modes.

Written by Kunj Shah

Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.