Your agent writes a spec. Instead of a wall of markdown in the terminal, it opens a page you can skim in a minute — and argue with, live.
npx skills add seyonv/red-penHere is how a coding agent hands you a 334-line spec today:
Spec written and committed to
docs/superpowers/specs/2026-08-07-the-source-in-two-registers-design.md.Self-review caught three things a fixture provider would otherwise have hidden, now written into Phase 1 explicitly: the on-add toggle ships defaulting to "ask",
FixtureComposertakes a configurable delay and charges zero, and⌘⌥3gets checked againstHotkeyCollisionbefore it's bound.Please review it and tell me if you want changes before I write out the implementation plan.
You now have three bad options: read all 334 lines, skim and hope, or type "looks good" and find out later.
| What's wrong | Why it matters |
|---|---|
| The review is buried in the message | The useful part is the middle paragraph — three judgment calls you might reverse. It arrives as prose and scrolls away. |
| There's no skim path | To find the part you disagree with, you read the whole thing. |
| You can't point at anything | Reacting to §4.2 means typing a message that describes where §4.2 is. |
It doesn't just render your markdown prettier. Before opening anything, the agent reads its own document and writes an annotation layer. That layer is the entire point:
1. A one-line summary under every section. Nine of them tell you the whole spec. This is the skim path.
2. The open decisions, pulled to the top. Every place the agent picked one of two defensible options, written as a real question with its recommendation and three buttons. That buried paragraph becomes the first thing you see.
3. An honest lede on what's actually risky, so you know where to spend attention.
The rule that makes it work: if there are no real open decisions, the block doesn't render at all. A page that always shows three decisions teaches you to ignore them.
Read the decision cards. Yes, ship it on two, Discuss on one. Each answer
collapses the card to a single line and ticks the counter down. Skim the section
summaries in the rail. Approve. Done.
Select any sentence in the document. A pill appears — click it or press c.
A thread opens in the margin with your selection quoted above the box. Type,
press ⌘↵.
The agent answers in about three seconds — in the thread, not the terminal. It says plainly whether it's changing the spec or defending the choice.
Two buttons sit under the reply: Change it and Fine as is.
Your button is what gets recorded — never the agent's reply. Its prose is reasoning; your click is the decision. A persuasive answer can't count itself as your agreement.
Once installed, it fires on its own. You don't type anything.
| When | What happens |
|---|---|
| Your agent finishes a spec or design doc | Page opens instead of the "please review it" message |
| Your agent finishes an implementation plan | Same |
| Your agent is about to exit plan mode | Page opens first; it exits with the revised plan |
| You ask for it | /red-pen docs/specs/my-thing.md on any markdown file |
When it opens, your terminal gets three lines and nothing else:
Spec written → docs/superpowers/specs/2026-08-08-….md
Review open → http://localhost:7654 (3 decisions need you)
or just tell me here — say "skip reviews" and I'll stop opening these.
Just say so, in any words:
"skip reviews" · "auto-approve everything" · "don't stop for me on this one"
That holds for the rest of the session. "let me review this one" turns it back on.
There's no config file and no size threshold, deliberately: that third line carries the off-switch every single time, so it's right there at the moment you're annoyed by it.
| Key | Does |
|---|---|
c |
Comment on the current selection |
j / k |
Next / previous section |
⌘↵ |
Send a comment or reply |
esc |
Cancel the comment you're writing |
a |
Approve |
Request changes sends the agent off to revise. Round two opens with a change
strip where the decisions used to be: every point you raised, what was done
about it, and a link straight to the changed text. Points the agent declined
are listed too, with the reason.
You never re-read a document hunting for your own feedback.
A Python server owns all review state and is the only writer. The page and the agent both talk to it over HTTP, which removes the write race between them.
page ──▶ POST /comment /decision /verdict /submit
◀── GET /state (long-poll)
agent ──▶ POST /reply /close
◀── GET /events?since=N (long-poll)
server ──▶ review.json (mirrored on every change)
This is the part worth stealing. A naive implementation polls in a loop and burns a turn every few seconds — a twenty-minute review would cost hundreds of calls. Instead the server holds the connection open:
curl -s --max-time 300 "localhost:7654/events?since=3"That request blocks for up to five minutes and returns the instant you do something. One call buys five minutes of attentive waiting for effectively nothing.
A twenty-minute review with three comments costs about seven calls.
After four consecutive empty polls the agent stops watching and tells you so — it never goes quiet without saying it has.
Anchors are quote-based — prefix / exact / suffix — not character offsets,
which shatter the moment a spec is edited. Threads carry into round two. And
when a revision genuinely destroys the quoted text, the thread says
text changed rather than silently drifting onto unrelated prose.
Port busy, no browser, malformed annotations — every failure falls back to the plain terminal message and the work continues. A review tool that can block a spec from shipping is worse than no review tool.
Two more guarantees:
- Nothing is written into your repo. Review state lives in
~/.claude/red-pen/<project>/<spec>/. No.gitignoreto edit, no way to commit review chatter by accident. - The page is completely self-contained. No CDN, no webfont, no network at runtime. If the server dies, the file still opens from disk, read-only.
Python 3.10+. That's the whole list. No pip installs, no npm installs, every asset vendored into the repo.
Node 20+ only if you want to run the JS tests.
python3 -m unittest discover -s tests # 122 tests
node --test "tests/*.test.mjs" # 12 testsMIT


