Part 1 of 12 · Build a Cursor-style AI coding agent in your terminal
What a Cursor-style terminal agent actually has to do
The shortest honest brief for building a terminal coding agent: find the command, stream the turn, gate the writes, print the cost, and fail in a way a human can read.
- Tutorial
- Architecture
- Agent
The brief, in one screen
Here is the whole product. Not a roadmap — the actual obligations, each of which either works or the tool is unusable.
- A command. One name, subcommands, a
--helpthat is accurate, and a non-zero exit code on failure. A CLI that lies about its own flags is worse than no CLI. - A loop that streams. Tokens arrive incrementally. A turn that prints nothing for 40 seconds reads as a crash, and you debug the wrong thing.
- Tools. Read, list, glob, grep, write, edit, run. The model can only discuss files it has been shown, and a grep tool shows it thousands.
- A gate. Read-only by default. The refusal is code, not a sentence in the system prompt — a prompt is advice, and advice is advisory.
- Accounting. Tokens and dollars per turn, printed. A tool whose cost you cannot see is a tool whose cost you cannot control.
What this course builds
This course builds a working terminal agent and grounds every part in Sentinel, the MIT-licensed agent in this repository. That is a deliberate constraint. Every command, file path and module name in the twelve parts is one you can run and grep, so when a part disagrees with the code you have found a bug rather than a version difference.
# the finished thing, in four commands
git clone https://github.com/KunjShah95/SENTINEL-CLI.git
cd SENTINEL-CLI && npm install && npm link
export GROQ_API_KEY=gsk_... # a free tier is enough for every part
sentinel doctor # pre-flight: runtime, keys, tool layer
sentinel ask "what does this repo do?"The one file that matters
Everything else is plumbing around a loop. Sentinel’s is 1,100 lines and yields typed events rather than printing directly, which is the single most important decision in the codebase: it lets the TUI and the CLI drive the identical agent without either one knowing about the other.
// The contract, in full. Everything the UI can observe:
export async function* runAgentTurn({ history, mode, model, goal }) {
yield* streamModel(history); // -> { event: 'text', data: { delta } }
// -> { event: 'tool_call', data: { toolName, input } }
const decision = gate(toolName, mode); // <- the control that matters
if (decision === 'deny') {
yield { event: 'tool_result', data: { denied: true, why: decision.reason } };
continue; // never reaches the tool
}
yield* execute(toolName, input); // sandboxed to the project root
yield { event: 'finish', data: { usage, costUsd } };
// -> { event: 'error', data: { message } }
}Three properties fall out of that shape. The gate is called in the loop rather than in the prompt, so it cannot be talked past. Tools are executed in-process, so there is no HTTP boundary between the model and the filesystem — which means no auth, no port, and an attack surface you can read in one sitting. And the loop yields, so a terminal UI, a --json pipeline and an MCP server are all just different consumers of the same generator.
Permission modes are the product
The cheapest useful design decision in this space is also the one most projects skip: define a small number of named modes, each an explicit allowlist, and refuse anything outside it. The number of modes does not matter much; the fact that they are enumerated in code does.
export function isToolAllowedInMode(toolName, mode) {
if (mode === Mode.BUILD || mode === Mode.SWE) return true;
if (mode === Mode.PLAN || mode === Mode.REVIEW || mode === Mode.SCAN) {
return isReadOnlyTool(toolName) || toolName === 'diffFile';
}
if (mode === Mode.FIX) {
// Read and write, but no shell: the agent can fix code, not run code.
return toolName !== 'bash' && toolName !== 'runTests';
}
return true;
}FIX is the interesting one. It grants writes while withholding the shell, which is what makes “let it edit but do not let it execute” a real mode rather than a promise. See part 8.
The stack, and why each piece earns its place
| Piece | Why it is here | What it replaces |
|---|---|---|
| Commander | Subcommands, flag parsing, generated --help from the same declarations that parse the args | Hand-rolled argv branching, which is where usage bugs live |
| Chalk | Colour that respects NO_COLOR and non-TTY pipes | Raw escape codes sprinkled through output code |
| Ink + React | A TUI that re-renders without fighting the terminal | readline plus a repaint loop you will debug for a week |
| A streaming client | One fetch path for OpenAI-compatible, Anthropic and Gemini wire formats | Three provider SDKs and three retry behaviours |
| node:test | Tests with zero dependencies, so the repo stays installable offline | A test framework as a supply-chain dependency |
| Nothing: no DB, no server | Sessions are JSON files you can read and delete | A datastore, a migration, and an auth system to run locally |
The dependency that is deliberately absent
There is no server component, no database and no account. That is not minimalism for its own sake — it is what makes git clone && npm install && sentinel a complete setup. Every dependency you add to a locally-run tool is one more thing that can break on someone else’s machine, and one more supply-chain question your team has to answer.
And it should speak MCP in both directions
Not as a feature, as a hedge. A local agent that can expose itself over the Model Context Protocol can be driven by Claude Desktop, Cursor or any other MCP client; one that can consume MCP servers can reach tools you did not write. Both directions cost a few hundred lines and remove the argument for switching tools later. Part 12 is the whole argument in one file.
How to read the twelve parts
Each part ends with something runnable, and several deliberately show a refusal. Type the commands instead of trusting the output blocks: the failure modes are the lesson, and the gate is much easier to believe once you have watched it deny you.
- Parts 2–4 — the shell: a command, a banner, and
doctor. No model access needed. - Parts 5–7 — the agent: a streamed turn, live tracking of what it is doing, and async generators as the interface.
- Parts 8–10 — the parts that decide whether you trust it: permission modes, the risk ledger, and cost you can cap.
- Parts 11–12 — a quiz to find the gaps, and the pitch for a cautious tech lead.
Part 9 mentions pnpm for global installs because that is a genuinely better workflow for a CLI you use daily. This repository itself uses npm, so where the two differ you will see npm — the commands are otherwise identical.
Frequently asked questions
What makes a CLI coding agent different from a chat wrapper?
Three things, and all three are failure modes rather than features. It must reach files the prompt never mentions, so it needs tools rather than text in and text out. It must stream, because a 40-second silent turn is indistinguishable from a hang. And it must be able to refuse, because the model will eventually ask to rewrite a file it should not touch. A CLI without the third property is a script that can delete your work.
Do I need the Claude Agent SDK to build one of these?
No, and this course deliberately does not require it. You need a streaming completion API and a tool-calling loop; those are about 200 lines. The Claude Agent SDK is a good default because it ships file and shell tools, permission modes and hooks you would otherwise write yourself, but it is an implementation choice, not the definition of the product. Part 5 shows the hand-rolled loop so the choice stays yours.
How long does it take to go from empty directory to working agent?
Parts 1 to 4 (argument parsing, a banner, pre-flight checks and a read-only chat loop) are a couple of hours and no model access beyond one API key. That first version already answers questions about your codebase. Everything after that is about letting it write, and about being able to prove what it did.
Should I build this or just install an existing agent?
Install one if you want the capability, build one if you want to understand the permission surface. The reason the second reason matters is that every guard rail in a finished agent is a few hundred lines of ordinary code, and you cannot audit a control you have not seen the shape of. This course is the cheapest way to get that shape into your head.
New here? The comparison page covers where a local agent fits against an AI IDE, and the quickstart gets you running in two minutes.
Written by Kunj Shah
Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.