A fast, minimal terminal coding agent (the canarycode command). It runs on the Bun runtime with an OpenTUI TUI and flexible model support. The core stays small. You get the quality-of-life features that matter: plan mode, auto mode, thinking modes, web search, sub-agents, custom agents, MCP, skills, hooks, an AI permission engine, and project-context files.
One engine drives two front-ends. A headless -p print mode handles scripting. An interactive TUI handles live work. Both run the same agent loop.
CanaryCode CLI is built to run on CanaryLLM — our hosted multi-provider model gateway. It ships baked in as the canaryllm provider, so all you need is a key.
Email contact@canarycoders.es to request one. Every CanaryLLM key starts with
clk…. Plans, pricing, and signup: https://canaryllm.canarycoders.es/
Set your key and you're ready — see CanaryLLM under Configuration for the env-var and manual setup.
CanaryLLM is the recommended path, but it's not the only one: canarycode works with any OpenAI- or Anthropic-compatible gateway, and you can bring or build your own provider — see Custom providers.
A single self-contained binary — no Bun or Node required:
curl -fsSL https://raw.githubusercontent.com/CanaryCoders/CanaryCodeCli/main/install.sh | shIt downloads the binary for your platform from the latest GitHub Release, verifies its checksum, and installs it to ~/.local/bin (override with CANARYCODE_INSTALL_DIR, or pin a version with CANARYCODE_VERSION=v0.1.0). Update later with canarycode update.
Run without installing:
nix run github:CanaryCoders/CanaryCodeCliOr add it to a Home Manager config (flake input named canarycode):
{
inputs.canarycode.url = "github:CanaryCoders/CanaryCodeCli";
# in your home-manager configuration:
imports = [ canarycode.homeManagerModules.default ];
programs.canarycode = {
enable = true;
# optional — renders a read-only ~/.canarycode/config.json:
# settings = { model = "opus"; thinking = "off"; };
};
}Nix installs are immutable, so self-update is disabled automatically — bump the flake input to upgrade.
You need Bun.
bun install
bun add -g . # installs the `canarycode` binary globallyRun it from the repo without installing:
bun run src/index.ts -p "say hi"
# or
bun run devcanarycode needs an API key before it can talk to a model. Pick one:
# Recommended — CanaryLLM (see #canaryllm above)
export CANARYLLM_API_KEY=clk...
canarycode --model <a-canary-model-id>
# Alternative — raw Anthropic
export ANTHROPIC_API_KEY=sk-ant-...
canarycodeEither can also be set persistently in ~/.canarycode/config.json — by hand, or with /config set providers.<name>.apiKey ... inside the TUI — and ${VAR}-style values there interpolate from the environment at load time. See CanaryLLM and Configuration for details.
Binary installs (via install.sh) self-update. On startup canarycode checks GitHub Releases at most once a day in the background; when a newer version is available the TUI banner shows a notice. Apply it with:
canarycode update # or /update inside the TUIAuto-update never mutates anything on its own — it only checks and notifies. Disable the check entirely with autoUpdate.enabled: false in ~/.canarycode/config.json or by setting CANARYCODE_DISABLE_UPDATE=1. Source checkouts (update with git) and Nix installs never self-update.
Release notes live in CHANGELOG.md and are published with each GitHub release. After an update lands, the next launch shows a one-line what's-new notice; read the full notes anytime with:
canarycode changelog # notes for the version you're running
canarycode changelog 0.1.0 # or any released version — /changelog inside the TUIcanarycode -p "<prompt>" # headless print mode: streams to stdout, then exits
canarycode # interactive TUI in the current directory| Flag | Description |
|---|---|
-p, --print <s> |
Run a single prompt headless |
--model <id> |
Override the configured model |
--think [level] |
Extended thinking: off, think, think-hard, or ultrathink (a keyword in the prompt triggers this too) |
--plan |
Planning mode: investigate read-only, emit a structured plan, then stop |
--auto, --yolo |
Autonomous mode: run to completion with no confirmations (capped at autoMaxTurns, default 25) |
--no-tools |
Disable tools for read-only quick Q&A |
--resume [id] |
Continue a saved session: bare --resume opens an interactive picker, an id resumes directly. Opens the TUI when interactive; runs headless with -p |
-c, --continue |
Continue the most recent session directly (no picker) |
--json |
Stream structured JSON events (JSONL) on stdout, one event per line, for scripting |
--no-color |
Force raw markdown to stdout even on a TTY (also honors NO_COLOR) |
-h, --help |
Show help |
--version |
Show version |
Stdin folds into the prompt as context:
git diff | canarycode -p "write a commit message"Resume a conversation:
canarycode --resume # pick a recent session to reopen in the TUI
canarycode --resume 1a2b3c4d # reopen a session by id (prefix is fine)
canarycode --continue # reopen the most recent session
canarycode -c -p "and now?" # continue the most recent session headlessInside the TUI, /resume lists recent sessions and /resume <id> switches to one in place (transcript, model, mode, and thinking level all restore).
Pipe structured events into a script:
canarycode -p "list the files then read package.json" --json | jq -c 'select(.type=="tool_start")'Each line is a JSON AgentEvent: text, thinking, tool_start {id,name,input}, tool_end {id,name,isError,result,diff?}, usage, compaction, done {reason}, or {type:"error",message}. JSON mode suppresses human formatting. Startup notes stay on stderr.
- Plan mode. Runs read-only and allows
read_file,list_dir,grep, andweb_search. It blocks write, edit, and bash. It emits a structured plan with steps, files to touch, and risks. - Auto mode. Runs autonomous multi-turn execution with no per-step input, bounded by
autoMaxTurns.Escaborts. - Thinking modes. They map to Anthropic extended-thinking budgets:
off,thinkat 4k,think-hardat 10k,ultrathinkat 32k. Non-Anthropic providers degrade gracefully. Setting a level with/thinkwrites it to~/.canarycode/config.jsonunder thethinkingkey, so it becomes the default on the next launch.--thinkoverrides it for one run without changing the saved default. - Web search. A read-only
web_searchtool works with no key by default through a free DuckDuckGo backend. You can configure Brave or Tavily backends underwebSearch. - Sub-agents. A
spawn_agenttool delegates focused work to a child agent with its own fresh context, bounded bymaxConcurrentandmaxDepth. - Claude Code subscription models. Run Anthropic models on your Claude Code Pro/Max/org subscription through the official Claude Code Agent SDK — no API key. Run
claude login, then pickclaude:opus,claude:sonnet,claude:haiku, orclaude:fable. The SDK drives its own agent loop with full tool execution; your skills, MCP, and project instructions are forwarded in, and even ifANTHROPIC_API_KEYis set this path uses the subscription, not the API. See Claude Code under Configuration. - OpenCode Zen. Use opencode's model gateway inside canarycode: sign in once with
opencode auth login(or setOPENCODE_API_KEY) and/login-opencodemakes the whole Zen catalog (Claude, GPT, Qwen, Kimi, GLM, the free stealth models, …) available to/model. - Extension toggles.
/extensionslists every toggleable feature —canaryllm,claude-code,codex,opencode,websearch,skills,agents,mcp,hooks— and/extensions enable|disable <name>flips one, persisted to~/.canarycode/config.jsonunderextensions.<name>. A disabled extension contributes nothing: no tools, no prompt text, no startup work. - Custom agents. You define personas as files in
~/.canarycode/agents/and./.canarycode/agents/. Frontmatter setsname,description, an optionalmodel, and an optionaltoolsallowlist. The body is the system prompt. Each name and description loads into the prompt.spawn_agentdispatches to one byagentname and applies its persona, model, and tool restrictions. - AI permission engine. Opt in with
permission.mode: "ai". A separate cheap model classifies each gated mutating tool call as safe or unsafe before it runs. Safe calls run silently. Unsafe calls escalate to the human y/n/a box in the TUI with the reason, or block the call when headless. Auto and--yoloskip it. - Hooks. Shell commands fire on lifecycle events:
PreToolUse,PostToolUse, andStop. A regex on the tool name matches them. A non-zeroPreToolUseexit blocks the call and its output becomes the reason the model sees. The rest observe only. Hooks run in every mode. - MCP. Connect stdio and SSE MCP servers from config. Their tools merge in namespaced as
mcp__<server>__<tool>. - Skills. Progressive-disclosure capabilities live in
~/.canarycode/skills/and./.canarycode/skills/. Only the name and description load into the prompt. The agent reads bodies on demand throughread_skill. - Tasks. The agent tracks its own todo list through
update_tasks. The TUI renders it live as a panel. Headless prints a compact checklist to stderr. The list is read-only and ephemeral. - Ask user. The
ask_usertool lets the agent ask you multiple-choice questions and wait for your answer. The TUI shows the choices. Headless auto-picks each question's recommended option. - Project context. Walking up to the repo root, it prepends the first file it finds in this order:
CANARYCODE.md, thenAGENTS.md, thenCLAUDE.md./initscaffolds a starterCANARYCODE.md. - Diff preview. Every
write_fileandedit_fileshows a unified diff of the change. It colorizes on a TTY. The TUI collapses it to a+N -Msummary you can expand. - Markdown rendering. Assistant text renders as styled terminal output with bold, headings, lists, and code fences. Under a pipe,
--json, orNO_COLORit stays raw. - Command visibility. The exact
bashcommand and every tool call shows before it runs, with no truncation. - Background shells.
bashcan start long-lived commands withrun_in_background: true. They stay alive across agent turns, retain combined stdout/stderr, expose incremental logs and exit status throughbash_output, and can be stopped withbash_kill. Remaining processes are terminated when the session closes.
read_file, write_file, edit_file, list_dir, bash, bash_output, bash_kill, grep, web_search, spawn_agent, read_skill, update_tasks, ask_user, plus any MCP tools. Each tool carries a readOnly flag. Plan mode filters on it.
Config lives at ~/.canarycode/config.json. ${VAR} references interpolate from the environment. Sessions store at ~/.canarycode/sessions.db in SQLite.
Model resolution order: --model <id>, then config model, then the first available model. /model lists and switches at runtime in the TUI.
In the TUI, /config inspects and edits ~/.canarycode/config.json without leaving the session:
/config # show the effective config
/config get model # show raw and effective values for one path
/config set model sonnet
/config set thinking think-hard
/config unset thinking
set values are parsed as JSON when possible, so booleans, numbers, arrays, and objects can be written directly. Strings can be bare words or JSON strings. Display output redacts secret-looking fields such as apiKey; the raw file still preserves literal ${VAR} references instead of writing interpolated secret values.
Common examples:
/config set webSearch.provider brave
/config set webSearch.apiKey "${BRAVE_API_KEY}"
/config set extensions.opencode false
/config set mcpServers.fs {"command":"mcp-server-filesystem","args":["/path"]}
/config set providers.mycorp {"api":"openai-compat","baseUrl":"https://llm.mycorp.internal/v1","apiKey":"${MYCORP_KEY}","models":[{"id":"company-default"}]}
/config set permission.mode ai
/config set permission.model haiku
/config set permission.scope writes
/config set confirm writes
/config set ui.nerdFont true
/config set autoMaxTurns 40
/config set checkpointEvery 8
/config set compactAtTokens 120000
/config set maxConcurrent 4
/config set maxDepth 2
/config set models.reasoning opus
/config set models.coding sonnet
/config set models.subagent haiku
/config set models.permission haiku
/config set hooks.PreToolUse [{"matcher":"bash|write_file|edit_file","hooks":[{"type":"command","command":"./scripts/guard.sh","timeout":10}]}]
/config set autoUpdate.enabled false
/config reload # reload config from disk
/config reload mcp # reload config and reconnect MCP servers
The TUI defaults to ASCII-safe icons. If your terminal uses a Nerd Font, enable richer glyphs with:
{ "ui": { "nerdFont": true } }canarycode supports custom OpenAI-compatible and Anthropic-compatible gateways:
{
"model": "company-default",
"providers": {
"anthropic": {
"api": "anthropic",
"apiKey": "${ANTHROPIC_API_KEY}"
},
"mycorp": {
"api": "openai-compat",
"baseUrl": "https://llm.mycorp.internal/v1",
"apiKey": "${MYCORP_KEY}",
"models": [
{ "id": "company-default" },
{ "id": "company-fast" }
]
}
}
}CanaryLLM ships baked in as the canaryllm provider and stays inert until you supply your key (every key starts with clk… — request one at contact@canarycoders.es). Once the key is present, canarycode discovers the gateway's chat models from the unauthenticated GET /api/public/models and makes them selectable through --model <id> and /model.
Supply the key by environment variable:
export CANARYLLM_API_KEY=clk...
canarycode --model <a-canary-model-id> -p "say hi"…or set it manually on the baked-in provider with /config set (persists to ~/.canarycode/config.json under providers.canaryllm.apiKey):
/config set providers.canaryllm.apiKey clk...
Display output redacts the key, and a literal ${CANARYLLM_API_KEY} reference is preserved in the file rather than written as an interpolated secret.
The preset is an openai-compat provider pinned to https://canaryllm.canarycoders.es/v1. The spec's servers list is localhost-only, so the base URL is hard-set. To use the Anthropic-compatible path instead, declare your own provider against /v1/messages with "api": "anthropic".
canarycode ships a baked-in claude provider backed by the official @anthropic-ai/claude-agent-sdk, so Claude Code Pro/Max or org users can run Anthropic models on their Claude Code subscription — no Anthropic API key required:
claude login
canarycode --model claude:sonnet -p "say hi"The preset exposes opus, sonnet, haiku, and fable, and authenticates through Claude Code itself (the claude login credentials under ~/.claude).
Subscription, not API billing. The claude provider always uses your subscription. Even if ANTHROPIC_API_KEY is set in your environment (for the separate anthropic provider), it is stripped from the SDK's environment so this path never silently bills the API. The two are deliberately distinct billing routes — qualify the model with a provider prefix to pick one explicitly when a name is otherwise ambiguous:
claude:opus— Anthropic model on your Claude Code subscriptionanthropic:opus— the same model family on your API key (ANTHROPIC_API_KEY)
Full tool execution. This provider hands the whole agent loop to the SDK rather than using it as a plain model transport. The SDK's native tools (Read/Write/Edit/Bash/Grep/Glob and its built-in web search) do the file, shell, and search work, while canarycode keeps ownership of skills, MCP servers, and your project instructions (CANARYCODE.md/AGENTS.md/CLAUDE.md) — these are forwarded into the loop so the model uses them alongside the SDK's tools. Your permission gates, plan mode, and PreToolUse/PostToolUse hooks still apply, routed through the SDK's permission callback.
Because the SDK owns the loop, a few behaviors differ from the other providers: earlier conversation turns are replayed to the SDK as text rather than a native session, queued input runs as the next turn rather than being injected mid-turn, and context compaction plus the interactive "keep going?" checkpoint are handled by the SDK within a turn (bounded by autoMaxTurns) rather than by canarycode.
Web search works with no API key. It defaults to a free DuckDuckGo backend that scrapes the no-JS results page and is subject to DuckDuckGo rate limits. For higher reliability and volume, point it at a keyed provider:
{
"webSearch": { "provider": "brave", "apiKey": "${BRAVE_API_KEY}" }
}provider accepts "duckduckgo" (default, keyless), "brave", or "tavily". When DuckDuckGo returns a challenge page from a rate limit, retry shortly or switch to a keyed provider.
Zen is opencode's OpenAI-compatible model gateway (https://opencode.ai/zen/v1). canarycode ships a baked-in preset that stays inert until credentials appear; nothing is written to config.json:
- Sign in once:
opencode auth login→ pick "opencode" (or setOPENCODE_API_KEYin your environment). canarycode login-opencode(or/login-opencodein the TUI) picks the key up from opencode's credential store (~/.local/share/opencode/auth.json) and lists the Zen catalog; switch with/model <id>.
The model list mirrors opencode's local models.dev cache (with a static fallback), so it tracks Zen's catalog without a network call. logout-opencode hides the models for the session; the credentials themselves belong to opencode (opencode auth logout removes them). Note: canarycode only reads Zen API keys from opencode's store — it never touches the Anthropic/OpenAI subscription OAuth tokens opencode may also hold, since refreshing those from a second client would invalidate opencode's own sign-in.
Bare /extensions opens an interactive checkbox picker: ↑/↓ move, Space flips a checkbox, Enter applies every change at once (one session reassembly), Esc cancels. /extensions enable|disable <name> flips one directly. Either way a change persists as extensions.<name> in ~/.canarycode/config.json and applies immediately.
A disabled extension contributes nothing — enforced by the registry kernel, not by the extension itself:
- no commands (absent from
/help, autocomplete, slash dispatch, andcanarycode <subcommand>) - no startup work and no provider preset (provider presets are folded into config for enabled extensions only)
- no session pieces (no tools, no system-prompt section, no tool hooks)
Core plumbing (the six core tools, ask_user, update_tasks, the permission engine) is not toggleable.
Drop .ts or .js modules into ~/.canarycode/extensions/ (user-global, implicitly trusted) or ./.canarycode/extensions/ (project-level). Project extensions require approval on first load: the TUI prompts before first paint, distinguishing a first-ever approval from a file that changed since the last one; headless skips unapproved files with a note. Approvals are tracked by content hash in ~/.canarycode/trusted-extensions.json.
An extension's name is its filename stem — any name field inside the module is overridden. This means the extensions.<name>: false config toggle is decidable before the file is even imported; a disabled file is never executed, but an inert stub keeps it listed in /extensions so it can be re-enabled. Name collisions with built-ins or with a user-global file are skipped with a note. A broken or invalid module is always skipped with a note; a user extension can never crash canarycode. Changed files take effect on the next launch (Bun module cache).
A minimal example:
// ~/.canarycode/extensions/greet.ts
export default {
description: "demo extension",
commands: [
{
name: "greet",
description: "say hello from a user extension",
run: async (ctx) => ctx.note("hello from the greet extension!"),
},
],
};User extensions have the same capabilities as built-ins: provider presets, startup discovery, login-style commands (canarycode <name> and /<name>), and full session extensions (tools, system-prompt section, pre/post tool hooks) via session().
{
"mcpServers": {
"fs": { "command": "mcp-server-filesystem", "args": ["/path"] },
"remote": { "url": "https://mcp.example.com/sse" }
}
}Token cost: while a server is connected, every MCP tool's name, description, and input schema is sent with each request — popular servers add thousands of tokens per turn. For tools you use occasionally, a skill or a plain CLI invoked via
bashis cheaper: the agent reads the docs only when it actually needs them.
Drop a markdown file in ~/.canarycode/agents/<name>.md for a global agent or ./.canarycode/agents/<name>.md for a project agent. The project file overrides the global one. The frontmatter configures the agent. The body is the agent's system prompt:
---
name: test-writer
description: Writes thorough unit tests for a given module.
model: haiku # optional, defaults to the parent's model
tools: read_file, grep, write_file # optional, defaults to the full inherited set
---
You are a meticulous test engineer. Given a module, write comprehensive tests…The model delegates to it through spawn_agent with { "agent": "test-writer", "task": "…" }.
{ "permission": { "mode": "ai", "model": "haiku", "scope": "writes" } }mode:"off"(default, falls back to the deterministicconfirmgate) or"ai".model: the checker model id (any configured model, defaults to a cheap one).scope:"bash"or"writes"(bash plus write_file plus edit_file).failClosed:false(default) ortrue— when the checker errors or can't run, treat the call as unsafe instead: headless blocks it, the TUI escalates to the human confirm box.
The checker runs one-shot and tool-free. It fails open by default, so a network or parse error counts as safe and a flaky checker degrades to running the tool rather than wedging the agent; set failClosed to flip that. Auto and --yolo skip the checker entirely — including the failClosed policy.
{
"hooks": {
"PreToolUse": [
{
"matcher": "bash|write_file|edit_file",
"hooks": [
{ "type": "command", "command": "./scripts/guard.sh", "timeout": 10 }
]
}
],
"PostToolUse": [
{
"matcher": ".*",
"hooks": [{ "type": "command", "command": "./scripts/log.sh" }]
}
],
"Stop": [{ "hooks": [{ "type": "command", "command": "echo done" }] }]
}
}Hooks follow Claude Code's matcher-group shape: event names map to arrays of groups, each group has an optional matcher and a hooks array of command hooks. timeout is seconds in this Claude-style shape. The older canarycode shorthand still works for compatibility ({ "matcher": "bash", "command": "...", "timeout": 10000 }, timeout in milliseconds).
Supported events are PreToolUse, PostToolUse, UserPromptSubmit, SessionStart, Stop, SubagentStop, and SessionEnd. matcher is a regex on the canarycode tool name for tool events; omitting it matches all tools. Commands run through bash -c and receive Claude-style JSON on stdin with fields such as session_id, cwd, hook_event_name, tool_name, tool_input, and tool_response; legacy aliases (event, tool, input, result, isError) are also included. PreToolUse can block with Claude-style JSON stdout (permissionDecision: "deny") or by exiting non-zero; the reason is sent back to the model. Other events are observational. Hooks run in every mode, including auto, and Stop/SessionEnd also run on TUI quit so external state trackers can observe shutdown.
canarycode runs unsandboxed by design. By default there is no automatic gating. Tool calls run as you invoke them. --auto and --yolo skip any pausing and run to completion. Run it in a directory you trust. Prefer --plan or --no-tools for untrusted work, and review what auto mode does. Esc in the TUI and Ctrl+C in headless abort an in-flight run.
For tighter control, three opt-in gates compose in a single pre-tool pipeline: PreToolUse hooks, then the AI permission check, then the human confirm. Hooks give you deterministic policy. The AI permission engine gives you model-judged safety. The confirm gate below gives you a human checkpoint. Auto and --yolo bypass the AI and human gates, and hooks still run.
For a lighter check than the AI engine, the TUI supports an opt-in confirm gate through the confirm config key. It applies only when permission.mode is "off", since the AI engine supersedes it:
{ "confirm": "writes" }"off"(default) runs everything."bash"pauses and shows the full command before eachbashrun."writes"pauses beforebash,write_file, andedit_file, showing the command or a diff preview, with[y]es · [n]o · [a]lways (this session).
Auto mode and --yolo bypass the gate. Plan mode never reaches mutating tools. Headless is non-interactive, so it ignores confirm. Use --auto to grant write and bash access unattended, or permission.mode: "ai" to gate it automatically.
/model, /think, /plan, /auto, /normal, /config, /extensions, /login-codex, /logout-codex, /login-opencode, /logout-opencode, /clear, /compact, /resume, /cost, /copy, /copy-last, /init, /update, /changelog, /help, /exit.
/config supports /config (show effective config), /config get <path>, /config set <path> <value>, /config unset <path>, /config reload, and /config reload mcp (reload config and reconnect MCP servers).
bun run typecheck # tsc --noEmit
bun run lint # biome lint ./src
bun run format # biome format --write ./src
bun run check # biome check ./src && tsc --noEmit
bun test # run the unit testscanarycode uses the Bun runtime and ESNext modules with no build step. .ts runs directly. The core targets under ~2000 LOC; keep new dependencies minimal and intentional.
canarycode follows a small-core design (inspired by Pi): the core is the agent loop (src/agent.ts), the provider layer (src/provider.ts), the core tool set (src/tools.ts, frozen), and session storage. Everything else — web search, skills, sub-agents, MCP, hooks, the AI permission engine, the task list — is an Extension (src/extension.ts): a named bundle covering two lifecycles in one interface.
Extension interface (src/extension.ts): { name, description, defaultEnabled?, providerPresets?(), startup?(config, "fast"|"live"), commands?, session?() }. The outer lifecycle (providerPresets, startup, commands) runs once per process. The per-session lifecycle is produced by the session() factory, which returns a fresh SessionExtension for each agent assembly; this keeps session state (MCP connections, tool instances) from leaking across sessions.
src/extensions/registry.ts is the single config-aware authority. It holds the built-in extension list and the user-loaded list together. Whether an extension is enabled (via extensions.<name> in config, toggled by /extensions) is decided here and nowhere else — extensions never self-check their own toggle. A disabled extension contributes nothing: no commands, no startup work, no provider presets, no session pieces.
To add a built-in feature: write a module in src/extensions/<name>.ts exporting an Extension and add it to BUILTIN_EXTENSIONS in registry.ts. To remove one: delete its file and its entry. Extensions that are disabled or not configured contribute nothing — no tokens, no startup work, no prompt text.
This project is licensed under the PolyForm Noncommercial License 1.0.0. Noncommercial use and forks are permitted under the terms of LICENSE; commercial use requires a separate commercial license from Canary Coders.