13 min read

Guardrails for AI coding agents: a design that survives week two

How to design AI coding agent guardrails engineers keep switched on: layered controls, once-per-path gating, and the friction traps that get them disabled.

  • Security
  • Architecture
  • Guardrails

The problem is friction, not safety

Every guardrail in an agent is a tax on the human. One extra keystroke here, one lost minute of context there, one turn where the agent stops and asks instead of shipping. The tax is defensible on a migration. It is indefensible on a typo fix, and a design that charges the same tax everywhere gets switched off.

So the question that actually matters is not “does this block dangerous writes?”, it always does, on the tenth attempt. The question is “will it still be enabled next month?” That reframes guardrail design from a security exercise into a product exercise, and it changes the answers.

Two corollaries worth stating up front, because they invert the usual instinct:

  • Fewer controls, better controls. Six layered checks that all hold beat fifteen that one developer worked around.
  • A block that is wrong is a bug, not a false positive. Nobody turns off a control that has never once been annoying. They turn off the one that blocked them four times for nothing.

The threat model these controls answer is set out in designing file permissions for an AI coding agent. This post is about the part that decides whether any of it gets used.

Four layers, and only the last one is interesting

Put every control in one place and the order becomes obvious, because three of the four are table stakes and one is the whole design problem.

Layer 1: expectations, not enforcement

The system prompt tells the agent to stay inside the project. This is documentation, not a control, and it is worth keeping for the same reason any interface has a label: it aligns the model with the boundary you are about to enforce, so the gate is a formality rather than a fight. Never count it as a layer.

Layer 2: modes as allowlists

A review mode that cannot write is worth more than a sandbox you have to trust. Use allowlists so the failure mode is closed: a tool you forgot to think of is refused rather than permitted. This layer is easy to build, which is exactly why it is easy to get wrong quietly, assert it in tests, because nothing else will tell you it regressed.

Layer 3: canonicalised paths, refused secrets

Resolve every path to its real location before comparing, and refuse secrets and catastrophic commands in every mode including the read-only ones. Nothing about this is subtle once you have read the threat model post, and it does not generate friction, so there is no adoption cost to worry about.

Layer 4: the expensive-path gate

This is the one that gets skipped, and the one that carries the actual risk. Layers 1 to 3 are uniform: everything outside the root is refused, everything inside is fair game. But the cost of being wrong is wildly non-uniform across the paths inside the root, and a uniform sandbox cannot express that.

where a wrong write actually costs something
db/migrate/**, *.sql          a migration is rarely undone by reverting it
.github/workflows/**           this gates every merge
prisma/schema.*, *schema.json changing a schema changes everything under it
package-lock.json, yarn.lock   a lockfile edit is invisible in review
src/auth/**, src/billing/**    the code you cannot roll back
Dockerfile, docker-compose*    build and deploy definitions
infra/, terraform/, k8s/      infrastructure definition
.sentinel/config.yaml          the project's own permission config

A typo in a comment costs nothing. A migration that quietly drops a column costs a restore, and it is the kind of restore that happens at 2 a.m. with a customer waiting. That asymmetry (one keystroke versus one incident) is the entire justification for a separate control, and it is why the trade should be explicitly toward over-asking on these globs. A utility file inside an auth directory should still get challenged. Nobody minds being asked once about a file that turned out to be harmless; everybody minds the alternative.

The once-per-path rule

This is the single most important implementation detail, and getting it wrong in either direction produces the two failure modes people actually see.

Gate granularity compared
GranularityBehaviourWhy it fails or holds
Once per writeAsks every timeFires on every edit in a migration. Becomes noise in a single turn, and the override becomes habit.
Once per path per turnAsks, then opens the path for the rest of the turnOne cost per risky decision. Correct agent behaviour stays the path of least resistance.
Once per path per repoOpens the path forever after one approvalA migration approved in March edits a different migration in July with no one watching.

The mechanism is deliberately boring: hold a set of challenged paths on the run, populate it when the gate refuses, test membership on every write, and clear the set at the turn boundary.

the whole gate
const CHALLENGED = new Set(); // per run, cleared at the turn boundary

function isExpensive(relPath) {
  return EXPENSIVE.some((rx) => rx.test(relPath));
}

function gateWrite(relPath, justification) {
  if (!isExpensive(relPath)) return { ok: true };

  // already paid for this path in this turn
  if (CHALLENGED.has(relPath)) return { ok: true };

  // the agent supplied both required fields, so record and proceed
  if (justification?.fileLine && justification?.rollback) {
    CHALLENGED.add(relPath);
    return { ok: true, recorded: justification };
  }

  return { ok: false, reason: BLOCK_MESSAGE };
}

Three properties of that shape are deliberate and worth naming:

  • It is scoped to the run, not the repo. Gating the same path on every turn forever trains people to pre-empt the gate by saying “yes, yes” without reading. A once-per-turn cost is paid once.
  • It opens on a complete answer, not on a confirmation. “yes” is not a justification. Requiring a specific file and a specific undo converts a rubber stamp into a two-second act of reasoning, and sometimes into the moment the agent realises it is about to do the wrong thing.
  • The answer is recorded. Once the file:line and the rollback are in the trajectory, the next person (or the next session) can see why the change was made instead of reconstructing it.

Design the refusal message, not just the refusal

A denial that says “blocked: not allowed” is a dead end. The agent cannot proceed, cannot satisfy the requirement, and will either retry blindly or give up. A refusal is an interface, and it has to carry three things: what happened, why this path is treated differently, and precisely what would let it through.

gate output
◆ blocked  db/migrate/0042_add_index.sql
  a migration is rarely undone by reverting it
  required: justifying file:line · exact rollback
✓ opened   for the rest of this turn

The middle line is doing more work than it looks. It explains the category of risk (irreversibility) rather than asserting a policy, which is what makes it survive contact with a model that has never seen this tool before. “Required: justifying file:line · exact rollback” is equally important: it is a contract the agent can satisfy on its next turn without guessing.

Test the message, not just the block

A gate that blocks correctly but explains badly produces retry loops, and a retry loop looks like a model failure in every log you will read. If a run is burning turns on a refusal, suspect the wording before the model.

Make it learnable, but keep memory separate

Some friction should disappear over time. If an engineer approves the same CI workflow change every week, asking forever is just noise. But approval memory and the safety gate are different concerns and should not be the same mechanism:

Gate memory compared with the risk ledger
MechanismScopeQuestion it asks
GatePer runThis specific change, right now
Risk ledgerPer repoThis kind of command, going forward

The rule that keeps this from becoming a backdoor: the gate is per-turn and asks about this specific change; the ledger is per-repo and asks about this kind of command. Approving git commit -m "a" must never authorise npm publish, and no amount of prior approval should let an agent modify a migration without saying why this time. If a repo-level “trusted path” list grows over time, the gate has been laundered into a suggestion, and that is the moment to delete the feature.

Six ways to build a guardrail that gets switched off

  1. One global toggle. “Allow this agent to edit files” covers a typo fix and a migration identically, so it is either useless or reckless. Granularity is what makes a control tolerable.
  2. A denylist of tools. A tool you forgot to list is permitted. Allowlists fail closed; denylists fail into your incident.
  3. Prompt-only enforcement. The most common and the most damaging, because it looks like a control in a demo and is absent in production.
  4. Blocking every write. Maximum safety, zero adoption, and the fastest route to an agent nobody trusts with anything real.
  5. Placeholder-normalising flags. Collapsing --force into a generic flag means git push --force matches an approved git push. Normalise values, never risk.
  6. Fail open on error. If an approval ledger is missing, grade everything novel as “ask”. An approving ledger that does not exist is not consent.

Prove they hold

Guardrails are the part of an agent most likely to regress silently, because they only fire in situations you are not looking at. They belong in the capability gate suite, next to the tool tests, and they need assertions on the failure path, not just the happy path.

the four assertions that matter
it("refuses a write outside the root", ...)          // traversal + absolute path
it("refuses a symlink that points out of the tree", ...)     // the bug that survives review
it("refuses a PLAN-mode write tool", ...)                    // allowlist fails closed
it("asks, when the approval ledger is absent", ...)          // fail closed, not open
it("asks once per path per turn, then opens it", ...)        // the friction contract
it("still asks on turn two for a new migration", ...)        // memory is not a bypass

That last one is the one teams skip, and it is the one that tells you whether the gate is a control or a formality. The full harness, including the reproduce-first workflow these assertions belong to, is in the SWE workflow docs and in how to evaluate a coding agent.

Measure the thing that actually predicts failure

“Number of blocked writes” is a vanity metric, a healthy project with dangerous habits can post a high number, and a broken gate can post zero. The signals worth watching are these:

  • Disable rate. The only metric that is unambiguously bad. If anyone turns a guardrail off, the design failed, not the user.
  • Retries after a block. The agent asking again without supplying what was requested means the refusal message is broken, not the agent.
  • Justifications that are vague. A file:line pointing at a test file, or a rollback of “revert the commit”, is a rubber stamp with extra steps.
  • Near misses. Times the gate stopped a change that was genuinely about to be wrong. This is the number that justifies the whole system, and it is the one to bring to a code review.

Frequently asked questions

What are guardrails in an AI coding agent?

Code-enforced limits on what the agent may do, independent of what the system prompt says. In practice there are four kinds: permission modes that map to tool allowlists, path sandboxing that keeps reads and writes inside the project, refusals for secret files and catastrophic commands, and higher friction on paths where a mistake is expensive, migrations, CI workflows, lockfiles, auth, billing, infrastructure. Only the last kind is usually built, and it is the one that decides whether the agent is trusted with a real repository.

Why do agent guardrails get disabled?

Because they are calibrated for correctness instead of for adoption. A guardrail that fires on every write trains people to reach for the override within days, and once the override is muscle memory the control is decorative. The fix is to over-ask only where the cost of being wrong is asymmetric, and to ask once per path per turn rather than once per write, so the correct agent behaviour is the path of least resistance.

Should an AI agent be allowed to edit migrations and CI workflows?

Yes, with a gate. Refusing outright is the wrong control, because it makes the agent useless for exactly the tasks people most want help with. The gate refuses the first write to a risky path in a turn and requires a justifying file:line plus an exact rollback; after that the path is open for the rest of the turn. The agent still does the work, and the justification gets recorded where a reviewer can see it.

Written by Kunj Shah

Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.