Designing file permissions for an AI coding agent
A threat model and eight-layer defence for coding agent file access: path sandboxing, symlink escapes, secret refusal, risk grading and the blast-radius gate.
- Security
- Architecture
- Sandboxing
Start with the threat model, not the feature list
“Sandbox the agent” is not a requirement, it is a candidate solution. Write down what you are actually defending against first. For a coding agent running with your developer’s credentials on a laptop, the realistic set is short and specific:
- Read escape. The agent reads
~/.ssh,.envor an API key file and puts the contents into a prompt that leaves the machine. - Write escape. A path traversal or symlink puts a write outside the project, into a directory that happens to be mounted.
- Destructive command. Not a bug in the agent, a correct execution of
rm -rfon a path the model mis-resolved. - Unreviewed high-blast-radius change. A migration, a CI workflow, a lockfile or an auth file edited in a way that looks fine in a diff and is catastrophic in production.
- Unattended spend. Not a security hole exactly, but a standing loop with no ceiling is a way to spend money quietly.
Notice what is absent: “the model turns hostile”. Prompt injection through repository content is real, but it is an amplifier of these five, not a separate category. If the controls below hold, an injected instruction reads no more than the operator’s own.
Layer 1: the prompt is not a control
Almost every agent ships a system prompt that says something like “never write outside the project directory”. Treat that as documentation. A model is a prediction function over a context window; it can be wrong, it can be talked out of a rule by content it just read, and on a long context the rule is competing for attention with a thousand lines of code.
The correct mental model: the prompt sets expectations, the code enforces boundaries. Every layer below is code. If a control only exists as text, it is not a control.
Layer 2: read-only modes that are actually read-only
A mode that is meant to review code should not be able to write code, and the way to guarantee that is an allowlist, not a denylist. Denylists fail open: a tool you forgot to list is permitted. Allowlists fail closed: a tool you did not think about is refused.
PLAN read, list, glob, grep, codeMap, web, diff, todos, skills
REVIEW read, list, glob, grep, codeMap, web, diff, todos, skills
BUILD everything
SCAN read-only + a test runner
FIX build tools minus the shell
SWE full loop, 60-iteration budget, verbatim tool historyThe tell for a real implementation is whether the check happens in one place, above the tool, rather than being re-implemented per tool. Enforcement that lives in each handler is enforcement that one handler will eventually forget.
Layer 3: canonicalise, then compare
This is the layer most sandboxes get wrong, and the bug is always the same shape: path.startsWith(root) evaluated on the string the model sent.
function isInside(root, path) {
return path.startsWith(root); // defeated by ../ and by symlinks
}import { realpath } from "node:fs/promises";
import path from "node:path";
async function isInside(root, candidate) {
// 1. make the path absolute against the project root
const abs = path.resolve(root, candidate);
// 2. canonicalise BOTH sides. This is what defeats symlink escapes
const [realRoot, realTarget] = await Promise.all([
realpath(root),
realpath(abs).catch(() => abs), // unresolvable => treat as the raw abs path
]);
// 3. compare with a separator so /repo-evil cannot pass as /repo
return realTarget === realRoot || realTarget.startsWith(realRoot + path.sep);
}Two details carry the weight. You must canonicalise the root too, because the project itself is often reached through a symlink, /tmp to /private/tmp on macOS is the classic case, and comparing a canonical path against a symlinked one refuses everything. And the comparison needs the trailing separator, or /repo-backup passes a /repo prefix check.
Symlink escape deserves a worked example, because it is the bug that survives code review. A project contains docs/link → ../../etc. Every read the model attempts as docs/link/passwd is textually inside the root, so a string check passes it, and the content of a system file enters the context and then the model call. Canonicalisation is what makes realTarget come back as /etc/passwd and get refused.
Unresolvable is not permitted
Layer 4: refuse secrets and the obvious footguns
Path sandboxing answers “where may it write”. It does not answer “which of the legal paths should be off limits in every mode”. Two categories need a hard deny:
- Secret material:
.env,*.pem,*.key,id_rsa,.npmrc, cloud credential files. A read refusal is the point. The model does not need the value, and anything it reads can end up in a provider request. - Inescapable commands:
rm -rf /, fork bombs,mkfs,shutdown,ddto a device. These are refused before execution, in every mode, including the ones that allow shell.
Note the asymmetry: refusing a secret read is safe and cheap, because nothing legitimate needs the value. Refusing a command is a judgement call, so keep the list deliberately short and unmistakable rather than clever, a denylist that blocks rm is a denylist that gets disabled.
Layer 5: grade by command shape, not by command
“Allow bash for this session” is one grant covering git status and npm publish alike. Either uselessly strict or uselessly loose. The fix is to record the shape of the command, per repository, and grade each new one.
$ sentinel risk
green git status approved shape
green git commit -m <msg> approved shape
yellow npm run bench new, not destructive
red git push --force always askedThree properties make the grading trustworthy:
- Keep flags, drop values.
git commit -m "a"and-m "b"are one shape, so the second does not ask. But--forceis not a value, it is a different verb, and it is never collapsed into a placeholder. - The subcommand is part of the verb. Approving
git commitmust never authorisegit push. Same fornpm runvsnpm publish,docker buildvsdocker push. - A missing ledger means nothing is approved. A corrupt or absent risk file must grade everything novel as
yellow. An approving ledger that does not exist is not consent.
Layer 6: a blast-radius gate that asks once
Sandbox rules are uniform: everything outside the root is refused. But the cost of being wrong is wildly non-uniform. A typo in a comment costs nothing; a migration that quietly drops a column costs a restore. So add a second, orthogonal check on the paths where being wrong is expensive.
| Path | Why it is challenged |
|---|---|
| db/migrate/**, *.sql | A migration is rarely undone by reverting it |
| .github/workflows/** | This gates every merge |
| prisma/schema.*, *schema.json | Changing a schema changes everything under it |
| package-lock.json, yarn.lock | A lockfile edit is invisible in review |
| src/auth/**, src/billing/**, **/rbac* | The code you cannot roll back |
| Dockerfile, docker-compose*, Makefile | Build and deploy definitions |
| infra/, terraform/, k8s/ | Infrastructure definition |
| .sentinel/config.yaml | The project's own permission config |
The gate blocks once per path per turn. The first write to db/migrate/ is refused with the requirement spelled out, and the agent must supply a justifying file:line and an exact rollback before the write lands. After that the path is open for the rest of the turn.
◆ blocked db/migrate/0042_add_index.sql
a migration is rarely undone by reverting it
required: justifying file:line · exact rollback
✓ opened for the rest of this turnThis design is deliberate and it is the difference between a gate that survives week two and one that gets switched off on day one. A gate that blocks forever trains people to disable it. A gate that asks once and records the answer builds a habit. The trade is explicitly toward over-asking: a utility file in an auth directory still gets challenged, because a false positive costs one prompt and a missed billing change does not come back. The full design (refusal wording, memory boundaries, the metrics that predict failure) is in guardrails for AI coding agents.
Layers 7 and 8: reversibility, and cost of unattended loops
Checkpoints. Every write is recorded, and undo works across turns, not just within one. The turn that breaks something is rarely the turn that looks like it did, so per-turn undo is not enough.
Budgets. A standing loop that wakes on every failed test is a presence, not a cron job, and it is a way to spend money quietly unless the ceiling is checked before every wakeup. Pair it with backoff (double the wait on a tick that made no progress, stop after five) and record failures rather than treating them as fatal.
sentinel watch "keep the sync green" -t "command:npm test" -t git
sentinel budget --usd 25 --deadline 2hThe review checklist
If you are auditing an agent, or building one, these are the questions. Each has a yes/no answer that takes a minute to verify in the source.
- Is enforcement in code, or only in the system prompt?
- Does the sandbox canonicalise both the root and the target?
- Does it compare with a path separator, or is
/repo-backupinside/repo? - Are modes allowlists, and do they fail closed?
- Are secret reads and catastrophic commands refused in every mode?
- Does command approval key on shape, with the subcommand included?
- Does a missing approval ledger mean “ask” rather than “allow”?
- Is the expensive-path gate once-per-path, not once-per-write?
- Does undo cross turn boundaries?
- Does an unattended loop check its budget before every wakeup?
Frequently asked questions
How do you stop an AI coding agent from reading files outside the project?
Resolve every path to its real location before comparing it, then compare against the project root, string prefix checks are bypassable. `root.includes(path)` is defeated by `../secrets` and by a symlink pointing outside the tree. Canonicalise first, reject on the resolved path, and treat an unreadable or unresolvable path as a refusal rather than as a pass. Sentinel does this at the tool boundary so it applies to every mode, including the ones you think are read-only.
Is a system prompt enough to stop an agent from doing damage?
No, and treating it as a control is the most common security mistake in agent tooling. A prompt is a suggestion to a model that may be wrong, distracted, or talking to itself in a long context; it is not an enforcement boundary. Only the code between the model's request and the filesystem is enforcement. Prompts are useful for informing behaviour, and useless as the only thing standing between a hallucinated path and your production config.
Should an agent ask before every file write?
No. A gate that fires on every write trains people to disable it, usually within a day. Ask once per risky path per turn, then open the path for the rest of that turn, and record the justification. The cost of a false positive is one keystroke; the cost of a real mistake on a migration is not recoverable, so the design should err heavily toward over-asking on the paths that are expensive to get wrong.
Written by Kunj Shah
Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.
Read next
- 30 Sept 2026 · 13 minGuardrails for AI coding agents: a design that survives week two
- 2 Sept 2026 · 9 minThe 7 best open source AI coding agents for the terminal
- 22 Sept 2026 · 10 minCut your LLM bill in half: cost control for coding agents
- 8 Sept 2026 · 8 minTurn any CLI coding agent into an MCP server