Part 2 of 12 · Build a Cursor-style AI coding agent in your terminal

11 min read

Building the command surface with Node.js and Commander

How to build a real CLI command surface with Commander: subcommands, variadic arguments, flag descriptions, exit codes, and the small conventions that make a CLI worth typing twice.

  • Tutorial
  • Node.js
  • CLI

Ten minutes to a working skeleton

This part needs no model access and no API key. By the end you have a command that parses correctly, lies to you nowhere, and exits with codes a shell can act on.

terminal
mkdir owl && cd owl
npm init -y
npm install commander
node -e "const p=require('./package.json');p.type='module';p.bin={owl:'bin/owl.js'};require('fs').writeFileSync('package.json',JSON.stringify(p,null,2))"
mkdir bin src/agent

"type": "module" is not optional decoration. Everything in this course is ESM, and mixing require with import in one project is a class of error you do not want to meet while debugging an agent loop.

The program object

Start at the top, because the first three decisions are the ones you will not revisit cheaply: the name, whether the version comes from one place, and which command runs with no arguments.

bin/owl.js
#!/usr/bin/env node
/**
 * owl, a terminal coding agent.
 *
 *   owl                 interactive chat (the default command)
 *   owl ask "..."       one-shot question, streamed answer
 *   owl doctor          pre-flight checks
 *   owl -V, --version
 */
import { Command } from 'commander';
import { readFileSync } from 'node:fs';
import { fileURLToPath } from 'node:url';
import path from 'node:path';

const root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');

// One source of truth for the version. Reading package.json beats a second
// literal that drifts on the first release you cut from a tag.
const VERSION = (() => {
  try {
    return JSON.parse(readFileSync(path.join(root, 'package.json'), 'utf8')).version || '0.0.0';
  } catch {
    return '0.0.0';
  }
})();

const program = new Command();

program
  .name('owl')
  .description('An AI coding agent for the terminal')
  .version(VERSION, '-V, --version', 'output the version number');

// isDefault means "run this when no subcommand is given", which is how you get
// `owl` to start a chat while `owl --help` still works.
program
  .command('chat', { isDefault: true })
  .description('Start an interactive chat')
  .action(async () => {
    const { runChat } = await import('../src/agent/chat.js');
    await runChat();
  });

Dynamic imports are not premature cleverness here

Each action loads its own module. For a CLI this buys real startup time — owl doctor does not pay to parse the agent loop — and it keeps a crash in one command from taking down the others, because the failing module never got imported by the others.

The one command that matters: variadic arguments

Users will type quoted and unquoted prompts interchangeably. If you declare <question> singular, then owl ask why is this test failing loses everything after the first word. Variadic plus a join is the fix, and it takes one line.

bin/owl.js
program
  .command('ask [question...]')
  .description('Ask a question and stream the answer (read-only by default)')
  .option('-m, --model <id>', 'Model id (defaults to the cheap one)')
  .option('-b, --build', 'Allow file edits and shell commands')
  .option('-y, --yes', 'Auto-approve tools; destructive commands are still denied')
  .option('-q, --quiet-cost', 'Do not print the token/cost summary')
  .option('--budget <usd>', 'Stop the turn once it has cost this much USD')
  .action(async (questionParts, options) => {
    const question = (questionParts || []).join(' ').trim();
    if (!question) {
      // Usage errors go to stderr and exit 1: a pipeline should not consume
      // this, and a shell `||` should fire.
      console.error('Usage: owl ask "your question"');
      process.exit(1);
    }

    // PLAN by default. Read-only is not a setting the user has to remember.
    const mode = options.build ? 'BUILD' : 'PLAN';

    const { runAgentTurn } = await import('../src/agent/loop.js');
    for await (const ev of runAgentTurn({ question, mode, model: options.model })) {
      if (ev.event === 'text') process.stdout.write(ev.data.delta);
      else if (ev.event === 'tool_call') process.stderr.write(`\x1b[2m→ ${ev.data.toolName}\x1b[0m\n`);
      else if (ev.event === 'finish') {
        process.stderr.write(`\x1b[2m${ev.data.usage.inputTokens} in / ${ev.data.usage.outputTokens} out\x1b[0m\n`);
      } else if (ev.event === 'error') {
        console.error(`\x1b[31m${ev.data.message}\x1b[0m`);
        process.exitCode = 1;
      }
    }
  });

Note the separation: delta goes to stdout, and everything else goes to stderr. That is not pedantry. It means this works:

terminal
# pipe the answer somewhere, still see progress and cost
owl ask "what does src/agent/loop.js do?" > answer.md

# fail the script if the turn errored
owl ask -b "fix the typo" || echo "the agent did not finish"

Exit codes are the API

This is the part tutorials skip and users depend on. Three rules, applied consistently:

Exit code conventions for a coding agent CLI
CodeMeaningWhen
0The tool did what you askedAnswer printed, goal verified, all pre-flight checks passed
1It did notProvider error, a failed check, a write the mode forbids
2You asked for something impossibleNo question given, mutually exclusive flags
124It ran out of timeReserved to match timeout(1), so a caller can distinguish it from a real failure

There is a subtlety worth copying: process.exitCode = 1 rather than process.exit(1) inside the loop. process.exit() kills the process immediately, which truncates buffered stdout — you lose the last few hundred characters of the answer on exactly the turns you most wanted to read.

Flag hygiene

Four conventions that cost nothing and make a CLI feel considered.

Short flags for what you type, long flags for scripts

-m for model because you will type it dozens of times; --budget spelled out because it will appear in a CI file someone else reads. Never make the short form the only form.

requiredOption fails before your handler runs

bin/owl.js
program
  .command('verify [task...]')
  .requiredOption('-c, --check <cmd>', 'Command that must exit 0 for the work to be accepted')
  .action(async (taskParts, options) => {
    // Commander has already exited 1 with a usage message if --check is missing.
  });

Build flags in a function, never a shared constant

bin/owl.js
// Wrong: the same Option instance attached to two commands leaks state.
const BUDGET = new Option('--budget <usd>', 'cap the spend');

// Right: fresh options per command.
function budgetOption() {
  return new Option('--budget <usd>', 'cap the spend for this run').default('1.00');
}

Describe the flags that change what is allowed

--yes is the flag that needs the most care in its help text, because it is the one that removes a gate. Write what it does not: “With --build: auto-approve tools including the shell; destructive commands are still denied.” A user reading only --help should not be able to misunderstand it.

Check it works

terminal
# the version is read from package.json, so it cannot drift
owl -V

# --help is generated from the declarations above
owl --help
owl ask --help

# a usage error must be non-zero and must not pollute stdout
owl ask; echo "exit=$?"

That last one is the test. If it prints exit=0, your exit handling is wrong and every script that wraps this tool is silently lying.

What part 3 adds

You have a correct command surface that tells you nothing about whether it works. Part 3 makes it look like a tool people keep open — a banner, a spinner, colours that disappear when you pipe — which is unglamorous and disproportionately effective, because a CLI that looks broken gets abandoned before anyone discovers its features.

Then part 4 adds doctor, which is the command that turns “why is this not working?” from a guess into a list.

Frequently asked questions

Why use Commander instead of parsing process.argv myself?

Because the help text and the parser must not disagree. With Commander, the declaration that makes `--budget <usd>` required to have a value is the same declaration that generates the `--help` line, so the two cannot drift. Hand-rolled argv parsing works right up until someone adds a flag to one place and forgets the other, which is the single most common way a CLI becomes untrustworthy.

Should every action be its own subcommand?

One line of test: if a user could plausibly want two of them in one invocation, they are one command. `sentinel risk 'git push'` and `sentinel risk 'npm publish'` are the same command with a different argument, not two commands. Splitting by argument rather than by verb produces a command list nobody can hold in their head.

How should a CLI signal failure?

Exit non-zero, print the reason to stderr, and keep stdout clean. That contract is what lets `sentinel ask ... || echo failed` and `sentinel doctor && sentinel ask ...` work at all. It is also what makes the tool scriptable into someone else's CI without a wrapper, which is the difference between a demo and a tool.

Where do I put shared flag definitions?

In a function, not a constant. Commander options carry state, once you attach them to a command, reusing the same Option instance across commands causes subtle leakage. A factory that returns fresh options per command costs three lines and removes the entire class of bug.

Prefer to read the finished version? The development guide covers lint, typecheck and the release check that gates a publish.

Written by Kunj Shah

Sentinel is an open source AI coding agent for the terminal, MIT licensed, no servers, no telemetry. Read the source or install it.