Part 11 of 12 · Build a Cursor-style AI coding agent in your terminal
Quiz: can you spot the three agents that would leak?
Twelve questions on building a terminal coding agent, with answers, because the wrong options are more instructive than the right ones.
- Quiz
- Tutorial
- Security
1. Where should the permission check for a tool call live?
The model requests writeFile in PLAN mode. Where does the refusal happen?
- In the system prompt, phrased firmly.
- In the tool handler, before dispatch.
- After the write, by rolling it back.
- In the provider layer, by filtering the response.
B. The loop checks the mode before calling the tool. A prompt is advice that competes for attention with the task; a rollback is a write that already happened. See part 8.
2. A user approves `git commit` for the session. Later the agent runs `git push --force`. What should happen?
- It runs. Bash is allowed.
- It asks, force-push is a different command shape.
- It runs, the session grant covers all git.
- It is denied permanently. Git is dangerous.
B. A session grant covers a shape. This only works if git commit and git push reduce to different shapes, which takes deliberate work. See part 9.
3. Which SSE detail, if missed, silently breaks your cost accounting?
- Using
TextDecoderwithout the streaming flag. - Flushing the buffer after the stream ends.
- Skipping `[DONE]` frames.
- Not setting `signal` on the fetch.
B. Servers routinely close without a trailing blank line, so the final frame (carrying finish_reason and usage) is still in your accumulator. (A is a real bug too: you get replacement characters mid-word. It breaks output rather than billing.) See part 5.
4. Your agent CLI writes a spinner. CI runs it and the build log fills with spinner frames. What is the correct fix?
- Detect CI and slow the spinner down.
- Send the spinner to stderr and leave stdout clean.
- Gate the spinner on `process.stdout.isTTY`.
- Disable animation whenever `--json` is passed.
C. B is necessary but not sufficient, CI captures both streams, so stderr frames land in the log too. D handles one code path and misses the twenty other ways CI runs your tool. See part 3.
5. The risk ledger file is corrupted by a merge conflict. What happens?
- It falls back to prompting on everything, correct.
- It falls back to allowing everything for the session.
- It refuses all shell commands until fixed.
- It regenerates an empty ledger and logs a warning.
A. The module fails closed: unreadable is treated as empty, and empty means every novel command is asked. B is the trap, a truncated write silently disables the whole mechanism, and disables it quietly. See part 9.
6. Why record tool-call arguments in history when trimming for a request budget?
- For debugging, so a trace is complete.
- They count toward request size, one big writeFile can blow the limit.
- Because the provider requires them alongside tool results.
- To let the model see what it already wrote.
B. A 40,000-character file body stays in the request forever after, so trimming only tool results is not enough. Keep the call id when truncating, or the result stops linking back and the provider rejects the ordering. See part 5.
7. The user presses Ctrl-C mid-turn. Which cleanup happens automatically?
- Nothing, you need an explicit handler per resource.
- The generator's `finally` blocks run.
- The HTTP request is aborted.
- All three, automatically.
B. Breaking a for await loop calls the generator’s .return(), which runs finally. C still needs an AbortSignal threaded into the provider call, the socket does not know about your loop. See part 7.
8. The blast-radius gate blocks the first write to a migration. Why not every write?
- Subsequent writes to the same path are provably part of the same change.
- It would be too slow for large migrations.
- The gate has a per-turn cap of one block.
- Because migrations should only ever be written once.
A, with the reason being adoption rather than logic. A gate that fires repeatedly trains people to disable it, and a disabled gate is worse than none because you stop looking for it. C is not true. The state is per path, not per turn. See part 10.
9. Which mode lets an agent edit files but not execute commands?
- PLAN
- FIX
- REVIEW
- SCAN
B. This is the most useful mode most permission designs skip. It is the difference between “let it help” and “let it help and verify”, and the reason teams end up either refusing writes entirely or granting shell access. See part 8.
10. Your pre-flight reports `warn` on Windows because of the PATH separator. What should `doctor` do?
- Exit 1, so the problem is visible in CI.
- Exit 0 and print the warning.
- Suppress warnings on Windows.
- Exit 0 but suppress the warning too.
B. Only a genuine failure should be non-zero. If warnings fail the run on any platform, people wrap it in || true and then it never runs at all — which is worse than not shipping the check. (This one is a real bug I shipped and fixed in this repo.) See part 4.
11. You want to store an API key. Where?
- In
.sentinel.yaml, committed. - In the environment, with the config file holding only non-secret settings.
- In the config file at 0600, gitignored.
- In a prompt template, so it is easy to copy.
B. The reason is not just git hygiene: it is that a pre-flight, a diagnostic and a bug report can all mention the variable name safely. Sentinel’s doctor prints groq (GROQ_API_KEY) and never the value, and there is a test that fails if a canary key appears anywhere in the output. See part 4.
12. The agent loop needs to serve a TUI, a `--json` pipeline, and MCP. What is the right shape?
- Three loops, one per output format, sharing the tool layer.
- One loop that prints, plus a mode flag.
- One loop that yields typed events; three consumers.
- An HTTP server with three routes.
C. B cannot produce JSON without a parallel code path, which is how the two paths drift. D adds an auth story and a port for something running on your machine. See part 7.
How did you do?
| Score | What it means |
|---|---|
| 11–12 | You are ready to build this. The next thing that will teach you is a real task on a real repository, not more reading |
| 8–10 | Sound grasp of the shape. Re-read parts 8–10; the misses cluster there for almost everyone |
| 5–7 | You have the vocabulary without the failure modes. Build parts 1–4, which need no model access, and the rest will follow |
| 0–4 | Not a problem. These are subtle by design. Start at part 3, since the output discipline is the cheapest habit to build |
The three that catch experienced engineers out
Questions 5, 8 and 10. All three are about failure modes that never announce themselves: a mechanism that silently disabled, a control that trains its own removal, and a health check that fails so routinely nobody runs it. The pattern is consistent — the dangerous bugs in agent tooling are the ones that work.
What part 12 adds
The last part is not code. It is the conversation you will actually have: what you say when a cautious tech lead asks whether the team should run an agent with shell access on the production repository, and what you can honestly offer instead of a demo.
Frequently asked questions
How should I use this quiz?
Answer before reading the explanations, and treat a wrong answer as more informative than a right one. The wrong options are all things that look reasonable, ship, and pass a code review. Which is precisely why they get written. If you get all twelve right, the next thing worth doing is writing the parts you disagreed with, because an agent you can defend is an agent your team will adopt.
Is there a score that means 'ready to build this'?
Not really. What matters is whether you can explain why the wrong options are wrong without referring to the article. Questions 4, 5 and 8 are the ones that separate 'has read about guard rails' from 'has had to debug one'. If those three are shaky, read parts 8 through 10 again rather than pressing on.
Why are there so many questions about refusals?
Because an agent's value is judged almost entirely on the turns where it says no. A model that answers every question competently is easy to build; a model that is trustworthy on the tenth run, when it is tired and you are watching a migration, is the entire product. Every control in parts 8 through 10 exists to make one specific refusal correct.
Prefer to go back? The full curriculum lists all twelve in order.
Written by Kunj Shah
Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.