12 min readupdated 30 Sept 2026

Designing file permissions for an AI coding agent

A threat model and eight-layer defence for coding agent file access: path sandboxing, symlink escapes, secret refusal, risk grading and the blast-radius gate.

  • Security
  • Architecture
  • Sandboxing

Start with the threat model, not the feature list

“Sandbox the agent” is not a requirement, it is a candidate solution. Write down what you are actually defending against first. For a coding agent running with your developer’s credentials on a laptop, the realistic set is short and specific:

  • Read escape. The agent reads ~/.ssh, .env or an API key file and puts the contents into a prompt that leaves the machine.
  • Write escape. A path traversal or symlink puts a write outside the project, into a directory that happens to be mounted.
  • Destructive command. Not a bug in the agent, a correct execution of rm -rf on a path the model mis-resolved.
  • Unreviewed high-blast-radius change. A migration, a CI workflow, a lockfile or an auth file edited in a way that looks fine in a diff and is catastrophic in production.
  • Unattended spend. Not a security hole exactly, but a standing loop with no ceiling is a way to spend money quietly.

Notice what is absent: “the model turns hostile”. Prompt injection through repository content is real, but it is an amplifier of these five, not a separate category. If the controls below hold, an injected instruction reads no more than the operator’s own.

Layer 1: the prompt is not a control

Almost every agent ships a system prompt that says something like “never write outside the project directory”. Treat that as documentation. A model is a prediction function over a context window; it can be wrong, it can be talked out of a rule by content it just read, and on a long context the rule is competing for attention with a thousand lines of code.

The correct mental model: the prompt sets expectations, the code enforces boundaries. Every layer below is code. If a control only exists as text, it is not a control.

Layer 2: read-only modes that are actually read-only

A mode that is meant to review code should not be able to write code, and the way to guarantee that is an allowlist, not a denylist. Denylists fail open: a tool you forgot to list is permitted. Allowlists fail closed: a tool you did not think about is refused.

mode → allowlist
PLAN    read, list, glob, grep, codeMap, web, diff, todos, skills
REVIEW  read, list, glob, grep, codeMap, web, diff, todos, skills
BUILD   everything
SCAN    read-only + a test runner
FIX     build tools minus the shell
SWE     full loop, 60-iteration budget, verbatim tool history

The tell for a real implementation is whether the check happens in one place, above the tool, rather than being re-implemented per tool. Enforcement that lives in each handler is enforcement that one handler will eventually forget.

Layer 3: canonicalise, then compare

This is the layer most sandboxes get wrong, and the bug is always the same shape: path.startsWith(root) evaluated on the string the model sent.

the naive version
function isInside(root, path) {
  return path.startsWith(root);   // defeated by ../ and by symlinks
}
the version that holds
import { realpath } from "node:fs/promises";
import path from "node:path";

async function isInside(root, candidate) {
  // 1. make the path absolute against the project root
  const abs = path.resolve(root, candidate);

  // 2. canonicalise BOTH sides. This is what defeats symlink escapes
  const [realRoot, realTarget] = await Promise.all([
    realpath(root),
    realpath(abs).catch(() => abs), // unresolvable => treat as the raw abs path
  ]);

  // 3. compare with a separator so /repo-evil cannot pass as /repo
  return realTarget === realRoot || realTarget.startsWith(realRoot + path.sep);
}

Two details carry the weight. You must canonicalise the root too, because the project itself is often reached through a symlink, /tmp to /private/tmp on macOS is the classic case, and comparing a canonical path against a symlinked one refuses everything. And the comparison needs the trailing separator, or /repo-backup passes a /repo prefix check.

Symlink escape deserves a worked example, because it is the bug that survives code review. A project contains docs/link → ../../etc. Every read the model attempts as docs/link/passwd is textually inside the root, so a string check passes it, and the content of a system file enters the context and then the model call. Canonicalisation is what makes realTarget come back as /etc/passwd and get refused.

Unresolvable is not permitted

When a path cannot be canonicalised, refuse the operation and say why. Treating an error as a pass is how "fail open" gets into a sandbox.

Layer 4: refuse secrets and the obvious footguns

Path sandboxing answers “where may it write”. It does not answer “which of the legal paths should be off limits in every mode”. Two categories need a hard deny:

  • Secret material: .env, *.pem, *.key, id_rsa, .npmrc, cloud credential files. A read refusal is the point. The model does not need the value, and anything it reads can end up in a provider request.
  • Inescapable commands: rm -rf /, fork bombs, mkfs, shutdown, dd to a device. These are refused before execution, in every mode, including the ones that allow shell.

Note the asymmetry: refusing a secret read is safe and cheap, because nothing legitimate needs the value. Refusing a command is a judgement call, so keep the list deliberately short and unmistakable rather than clever, a denylist that blocks rm is a denylist that gets disabled.

Layer 5: grade by command shape, not by command

“Allow bash for this session” is one grant covering git status and npm publish alike. Either uselessly strict or uselessly loose. The fix is to record the shape of the command, per repository, and grade each new one.

bash
$ sentinel risk
green   git status                     approved shape
green   git commit -m <msg>            approved shape
yellow  npm run bench                  new, not destructive
red     git push --force               always asked

Three properties make the grading trustworthy:

  1. Keep flags, drop values. git commit -m "a" and -m "b" are one shape, so the second does not ask. But --force is not a value, it is a different verb, and it is never collapsed into a placeholder.
  2. The subcommand is part of the verb. Approving git commit must never authorise git push. Same for npm run vs npm publish, docker build vs docker push.
  3. A missing ledger means nothing is approved. A corrupt or absent risk file must grade everything novel as yellow. An approving ledger that does not exist is not consent.

Layer 6: a blast-radius gate that asks once

Sandbox rules are uniform: everything outside the root is refused. But the cost of being wrong is wildly non-uniform. A typo in a comment costs nothing; a migration that quietly drops a column costs a restore. So add a second, orthogonal check on the paths where being wrong is expensive.

Paths challenged by the blast-radius gate and why
PathWhy it is challenged
db/migrate/**, *.sqlA migration is rarely undone by reverting it
.github/workflows/**This gates every merge
prisma/schema.*, *schema.jsonChanging a schema changes everything under it
package-lock.json, yarn.lockA lockfile edit is invisible in review
src/auth/**, src/billing/**, **/rbac*The code you cannot roll back
Dockerfile, docker-compose*, MakefileBuild and deploy definitions
infra/, terraform/, k8s/Infrastructure definition
.sentinel/config.yamlThe project's own permission config

The gate blocks once per path per turn. The first write to db/migrate/ is refused with the requirement spelled out, and the agent must supply a justifying file:line and an exact rollback before the write lands. After that the path is open for the rest of the turn.

gate output
◆ blocked  db/migrate/0042_add_index.sql
  a migration is rarely undone by reverting it
  required: justifying file:line · exact rollback
✓ opened   for the rest of this turn

This design is deliberate and it is the difference between a gate that survives week two and one that gets switched off on day one. A gate that blocks forever trains people to disable it. A gate that asks once and records the answer builds a habit. The trade is explicitly toward over-asking: a utility file in an auth directory still gets challenged, because a false positive costs one prompt and a missed billing change does not come back. The full design (refusal wording, memory boundaries, the metrics that predict failure) is in guardrails for AI coding agents.

Layers 7 and 8: reversibility, and cost of unattended loops

Checkpoints. Every write is recorded, and undo works across turns, not just within one. The turn that breaks something is rarely the turn that looks like it did, so per-turn undo is not enough.

Budgets. A standing loop that wakes on every failed test is a presence, not a cron job, and it is a way to spend money quietly unless the ceiling is checked before every wakeup. Pair it with backoff (double the wait on a tick that made no progress, stop after five) and record failures rather than treating them as fatal.

bash
sentinel watch "keep the sync green" -t "command:npm test" -t git
sentinel budget --usd 25 --deadline 2h

The review checklist

If you are auditing an agent, or building one, these are the questions. Each has a yes/no answer that takes a minute to verify in the source.

  1. Is enforcement in code, or only in the system prompt?
  2. Does the sandbox canonicalise both the root and the target?
  3. Does it compare with a path separator, or is /repo-backup inside /repo?
  4. Are modes allowlists, and do they fail closed?
  5. Are secret reads and catastrophic commands refused in every mode?
  6. Does command approval key on shape, with the subcommand included?
  7. Does a missing approval ledger mean “ask” rather than “allow”?
  8. Is the expensive-path gate once-per-path, not once-per-write?
  9. Does undo cross turn boundaries?
  10. Does an unattended loop check its budget before every wakeup?

Frequently asked questions

How do you stop an AI coding agent from reading files outside the project?

Resolve every path to its real location before comparing it, then compare against the project root, string prefix checks are bypassable. `root.includes(path)` is defeated by `../secrets` and by a symlink pointing outside the tree. Canonicalise first, reject on the resolved path, and treat an unreadable or unresolvable path as a refusal rather than as a pass. Sentinel does this at the tool boundary so it applies to every mode, including the ones you think are read-only.

Is a system prompt enough to stop an agent from doing damage?

No, and treating it as a control is the most common security mistake in agent tooling. A prompt is a suggestion to a model that may be wrong, distracted, or talking to itself in a long context; it is not an enforcement boundary. Only the code between the model's request and the filesystem is enforcement. Prompts are useful for informing behaviour, and useless as the only thing standing between a hallucinated path and your production config.

Should an agent ask before every file write?

No. A gate that fires on every write trains people to disable it, usually within a day. Ask once per risky path per turn, then open the path for the rest of that turn, and record the justification. The cost of a false positive is one keystroke; the cost of a real mistake on a migration is not recoverable, so the design should err heavily toward over-asking on the paths that are expensive to get wrong.

Written by Kunj Shah

Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.