Skip to content

feat: revive Sentinel — upgrade all dependencies, add every needed agent, agent parity - #113

Open
chaqchase wants to merge 133 commits into
mainfrom
revive/2026-10
Open

chaqchase wants to merge 133 commits into
mainfrom
revive/2026-10

Conversation

@chaqchase

@chaqchase chaqchase commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Summary

Draft PR that tracks the Sentinel revival. Each phase lands here as one or more commits, and every commit passes the gate before the next one starts. The gate checks frozen install, typecheck, the full test suite, format, tripwires and next build.

Goals:

  • Update every package to its latest version: Node 24, AI SDK 7, Next 16.4, React 19.3, Electron 44, zod 4, TS 6, and the agent SDKs.
  • Keep every existing engine working.
  • Bring every agent up to parity:
    • unified status
    • install and update
    • version advisories
    • in-app auth
    • usage limits
    • model manifest
    • skills and slash commands
    • multiple instances per agent

Phases

  • P0 Harness: runner, CI coverage, gates and tripwires
  • P1 Node 24 + bun
  • P2 Low-risk dependency sweep
  • P3 zod 4 · P4 TypeScript 6
  • P5 Next 16.4 + React 19.3 + HeroUI 3.2
  • P6 Electron 44 + electron-builder + better-sqlite3 13
  • P7 AI SDK 7
  • P8 Built-in model catalog refresh
  • P9 Agent SDK upgrades: Claude 0.3, Copilot 1.0, OpenCode 1.18, Codex 0.160 (+ protocol fixtures for P12/P13)
  • P10 Engine driver contract + instance model
  • P11 Platform services and UI: manifest, install/update, auth, usage, slash commands
  • P12 Shared ACP layer + Cursor port
  • P13 New engines: Grok, Antigravity, ACP Registry, Pi, OpenCode v2
  • P14 Cross-engine extras: MCP forwarding, deprecation sweep
  • P15 CI · P16 Docs · P17 Release gate

Notes for reviewers

  • AI SDK 7 / @ai-sdk/mcp 2: OAuth MCP servers authorized under MCP 1.x have to be re-authorized once. The stored tokens carry no authorization-server pin.
  • OpenAI/Azure function tools: these stay at AI SDK 7's default strict: false, pinned by a test. AI SDK 6 left the field out, and strict: true rejects schema keywords Sentinel uses.
  • xAI: @ai-sdk/xai 5 is Responses-only. Sentinel now sends store: false to keep the old no-retention behaviour.
  • Electron 44 needs macOS 13+. The mac update feed now has minimumSystemVersion, so macOS 12 installs are not offered an update they can't run.
  • Built-in model catalog (P8):
    • Every provider now leads with its current flagship model.
    • Retired model IDs are mapped to the provider's replacement in threads, automations and defaults. A model ID you still have enabled (for example as a custom model) is kept.
    • Claude 4.6+ and Gemini 3.x now show reasoning summaries.
  • Copilot (P9): the SDK's own runtime is now bundled, so no separate Copilot CLI install is needed. The dmg grows from about 163 MB to about 213 MB. COPILOT_CLI_PATH still overrides the bundled runtime.
  • Claude (P9):
    • Agent SDK 0.3 uses your own claude binary; its native CLIs are not shipped.
    • TodoWrite was replaced by the Task tools, which get their own renderers.
    • Effort and adaptive thinking are sent to the agent.
    • AskUserQuestion answers now go through the SDK's updatedInput.
    • Commit messages run with no tools instead of skipping permissions.
  • Codex (P9):
    • Updated to the 0.160 app-server protocol: experimentalApi, item/tool/requestUserInput, permissions and elicitation requests.
    • thread/rollback was replaced by thread/revert.
    • Token usage now reads the total/last shape.
    • Windows codex.cmd support.
  • OpenCode (P9):
    • SDK 1.18, with readiness checks that poll the server instead of matching its log line.
    • The compaction flow now completes instead of ending the run.
    • Reasoning shows as its own part.
    • Approvals from subagents are no longer dropped.
    • Minimum supported OpenCode version is 1.0.224.
  • Cursor (P9 hotfix): session/cancel is sent as a notification, so Stop can no longer hang.
  • Holds: TypeScript 6.0.3 (TS 7 has no compiler API yet), drizzle 0.45.x (1.0 is still an RC), officeparser 6.1.1 (v8's pdfjs-dist 6 hangs under bun).

Test plan

  • P0: G1 gate passes (183/183 test files, 1150 tests; the old runner stopped at the first failure)
  • P1–P7: G0 at each commit; G2 (unsigned mac package + bundle audit + packaged native smoke: Electron 44.6 / Node 24.21, better-sqlite3 13, sqlite-vec 0.1.9, undici 7) after merging P3–P7
  • P8/P9: G0 per lane, then G2 on the merged tip (unsigned mac package + bundle audit: Copilot runtime target-only, no Claude native CLIs, undici 7, SQLite natives)
  • Each later phase: G1 gate; packaging gate G2 at P17
  • Protocol fixtures for every agent (mock ACP agent, mock Pi RPC, Codex/Claude/OpenCode/Copilot replays)
  • Isolated UI and probe check on a dev server

🤖 Generated with Claude Code

chaqchase and others added 30 commits October 7, 2026 09:17
- run-tests.mjs now runs every test file (it used to exit on the first
  failure, hiding ~500 tests), prints a failure summary, supports
  --jobs/--filter/--bail, includes scripts/ tests, and gives each file
  its own throwaway SENTINEL_STATE_PATH/DB/MEDIA so runs never touch a
  developer's real ~/.sentinel data
- make run_task and shell tests independent of the host `node` binary
- extend the bun:test type stub (test, spyOn, beforeAll, skipIf, ...)
- exclude .claude and worktrees from tsc
- CI: trigger on scripts/, desktop/, package.json, bun.lock and config
  changes, add a build job, run tests in parallel
- add scripts/verify/{gate,tripwires}.mjs for revival phase gates

Gate G1: typecheck ok, 183/183 test files (1150 tests) pass, format ok,
next build ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Node 21.7.3 is end-of-life and below the minimum of the upgrade targets
(ai 7, @ai-sdk/*, Electron 44, better-sqlite3 13, Copilot SDK 1.0 and Pi
all require Node >= 22). Electron 44 embeds Node 24, which is also what
the packaged Next server runs on, so pin the 24 LTS line everywhere:
.nvmrc, .node-version, engines, @types/node, and the CI setup action
(now reads .nvmrc).

Gate G0: typecheck ok, 183/183 test files pass, format ok, tripwires ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bun 1.4 is the current release line. The full suite (183 files, 1150
tests, 260 mock.module call sites) passes unchanged on 1.4.2, so pin it
for packageManager, engines and the CI setup action.

Gate G0: typecheck ok, 183/183 test files pass on bun 1.4.2, format ok,
tripwires ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tRPC 11.19, TanStack Query 5.104, drizzle-orm 0.45.3 / drizzle-kit
0.31.11 (staying on 0.x; 1.0 is still a release candidate), sqlite-vec
0.1.9 (first non-alpha), Tailwind 4.3.3, tiptap 3.31, lucide-react 1.52,
hugeicons, @pierre/diffs 1.5, notion 5.27, mongodb 7.7, mysql2, pg,
react-hook-form 7.89, resumable-stream, turndown, electron-updater 6.8.9,
prettier 3.9.9, lefthook 2.1.17, postcss, wait-on.

Gate G0: typecheck ok, 183/183 test files pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Prettier 3.9 collapses short union type annotations onto one line.
Formatting only, no behaviour change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- googleapis 183: build OAuth clients with google.auth.OAuth2 instead of
  importing the undeclared google-auth-library, which now resolves to a
  different major than the one googleapis uses
- @linear/sdk 97, @slack/web-api 8 (fetch transport; Sentinel uses no
  axios-specific options or ErrorCode checks), motion 14 (no breaking
  changes for React), shiki 4 (only removed misspelled aliases, unused)
- commitlint 21, lint-staged 17, concurrently 10 (verified the commit-msg
  hook and `concurrently -k` still behave)

Gate G0: typecheck ok, 183/183 test files pass, format ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- remove tailwind-variants, highlight.js and rehype-highlight (no
  imports anywhere) and the two Prettier plugins, which never loaded
  because the repo has no Prettier config listing them
- officeparser 6.1.1 (latest 6.x). Holding below 8: officeparser 8 pulls
  pdfjs-dist 6, which hangs under bun (the test runtime) on PDF parsing
  even though it works on Node. Revisit when bun or pdfjs fixes that.
- add PDF and ODT regression tests for the officeparser extraction path
  (the document parsers are loaded through runtime specifiers, so they
  were easy to mistake for unused dependencies)

Gate G0: typecheck ok, 183/183 test files pass, format ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zod 4.6.5 with @t3-oss/env-nextjs 0.13.11 and @hookform/resolvers 5.9.1.
This also satisfies claude-agent-sdk's zod ^4 peer, and the lockfile now
resolves a single zod copy (copilot-sdk's nested zod 4.3.6 is deduped
onto the root), so instanceof ZodError checks see one class.

Code changes for the v4 breaks:
- z.record() takes an explicit key schema at all 19 one-argument sites
  (integration tools, providers router, thread message JSON schema)
- appearance font size: invalid_type_error/required_error -> error fn,
  same messages as before
- tRPC errorFormatter uses z.flattenError; tRPC 11 still calls zod's
  parseAsync, so input errors keep a ZodError cause and the zodError
  payload keeps its { formErrors, fieldErrors } shape
- threadCreate threadId: z.guid() keeps zod 3's 8-4-4-4-12 check (v4
  uuid() is RFC-strict); the only producer is crypto.randomUUID()
- ZodTypeAny -> z.ZodType

ai 6.0.116 -> 6.0.301 (and @ai-sdk/react 3.0.304 so ai stays single):
provider-utils before 4.0.47 forced additionalProperties:false onto
every zod 4 object schema, so record inputs (Mongo queries, Airtable
fields, Notion properties: 23 tool fields) reached models as objects
that accept no keys. ai 6.0.265 is the first release with the fix; the
rest of the AI SDK stays where it is for P7.

Checked and unchanged: no .default().optional() or .partial() over
defaults, no superRefine ctx.path, .default() values already match the
output types, env.js still validates and SKIP_ENV_VALIDATION still
skips. Remaining zod 3 idioms that v4 still accepts (.url(), .email(),
ZodIssueCode.custom, superRefine) are left as they are.

Tool JSON schemas vs zod 3 (now snapshotted): .int() adds safe-integer
min/max, records add propertyNames, the email field adds a pattern,
computer_action's discriminated union is oneOf instead of anyOf, and
.optional().nullable() drops the old anyOf/not:{} wrapper. Sentinel
sends no strict tool schemas, and the Google converter keeps oneOf and
drops the other new keywords.

Tests: tool input JSON-schema snapshot over all 48 built-in and 159
integration tools (fails on objects that accept no properties), tRPC
zodError shape, env.js validation and skip, appearance messages,
thread/workspace schemas, engine-state round trip, settings defaults.
Tripwires: invalid_type_error/required_error, one-argument z.record,
ZodTypeAny.

Gate G0: install, typecheck ok, 188/188 test files pass, format ok,
tripwires clean (4 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
typescript ~6.0.3, not 7: TS 7 has no compiler API yet, which the Next
TS plugin and Next's tsconfig handling still need.

tsconfig.json for the TS 6 defaults and deprecations:
- drop "baseUrl" (deprecated in 6, an error in 7); "paths" are already
  relative to the tsconfig, and Next, bun and esbuild all resolve them
  without it
- "types": ["node"], since TS 6 no longer loads every @types package
  by default; everything else Sentinel uses comes in through imports
- include .next/dev/types/**/*.ts, where Next 16 dev writes route types
- keep the "next" language-service plugin

noUncheckedSideEffectImports is on by default in TS 6 and flagged
`import "@/styles/globals.css"`. next-env.d.ts declares *.css but is
generated and gitignored, so src/types/css.d.ts declares it the same way
Next 16 does (`declare module "*.css" {}`); side-effect imports of code
modules stay checked.

TS 6 also reported an always-nullish `?? null` in
resolveOpenCodeTraitValueForThreadMode; the redundant fallback is gone,
behaviour is unchanged.

Tripwire: "baseUrl" in tsconfig*.json.

Gate G0: install, typecheck ok, 188/188 test files pass, format ok,
tripwires clean (5 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ai 6.0.301 bump brought @ai-sdk/provider-utils 4.0.57. On Node, its
default download fetch (generateVideo URL results, prompt asset and
transcription downloads) loads undici via
module.createRequire(<caller file>)("undici") to pin connections to
validated DNS results. Once webpack bundles and minifies that code into
.next/server/chunks, nft cannot see the require, so `output:
"standalone"` left undici out: a G1 build had no
.next/standalone/node_modules/undici, and resolving it from chunk 119 in
the packaged layout failed with MODULE_NOT_FOUND. That would reject every
default download in the desktop server (ELECTRON_RUN_AS_NODE makes
isNodeRuntime() true). bun skips this path, so the tests never saw it.

- scripts/desktop/untraced-server-packages.mjs lists such packages
  (undici)
- next.config.js adds them to outputFileTracingIncludes for every route
- the bundle audit fails if the packaged server lacks one
- a test checks the trace include and that the top-level copy, which the
  bundled chunks resolve, satisfies every installed dependent's range
  (the three provider-utils 4.0.57 copies want ^6.28.0; root is 6.29.0)

A narrower ai bump would not avoid this: undici arrived in provider-utils
4.0.45 and the record additionalProperties fix in 4.0.47, and AI SDK 7's
provider-utils 5 depends on undici ^7.28.0.

Verified with a production build (scripts/next-build.mjs):
.next/standalone/node_modules/undici now exists and survives
prepare-production + prune-server. From a copy outside the repo, the
chunk-relative require resolves server/node_modules/undici and builds an
Agent.

Behaviour note: since provider-utils 4.0.47 these downloads also reject
hostnames that resolve to private, loopback or link-local addresses.
Literal localhost and private-IP URLs were already rejected in 4.0.19.
Sentinel sets no global dispatcher or proxy, so the dedicated Agent
bypasses nothing.

Gate G0: install, typecheck ok, 189/189 test files pass, format ok,
tripwires clean (5 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two zod 4 behaviour changes that the P3 commit let through:

- computer_action: zod 4 emits JSON Schema oneOf for discriminated
  unions (they are always exclusive), where zod 3 sent anyOf. The Google
  provider forwards oneOf unchanged, but Gemini's function-declaration
  Schema (an OpenAPI 3.0 subset; @google/genai's Schema type) has anyOf
  and no oneOf. Gemini requests that activate the computer tools would
  likely be rejected. The tool input now uses z.union over the same
  action options. Parsed data is identical for valid input, and the
  discriminated union stays in automation-types for everything else. The
  tool schema test now flags any oneOf/allOf, and the snapshot changes
  only that keyword. Trade-off: a malformed action gets the union's
  per-branch issues instead of the single discriminated-union issue.
- codexWriteConfig: zod 4 rejects an absent key for z.unknown() at
  runtime ("expected nonoptional"); zod 3 accepted it as undefined.
  value is now z.unknown().optional(), which restores both the runtime
  check and zod 3's inferred `value?: unknown`. An audit of z.unknown,
  z.any, z.undefined, z.void, z.custom and undefined unions found no
  other production object key. The test mock now keeps input schemas so
  the router test can check this.

The one-argument z.record tripwire comment now says it matches single
lines with at most two nested parenthesis levels and that typecheck is
the main guard: a multi-line call fails with TS2554, checkJs included.

Gate G0: install, typecheck ok, 189/189 test files pass, format ok,
tripwires clean (5 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Moves ai 6.0.301 -> 7.0.130 and every @ai-sdk package to its v7 major:
provider 4.0.24, react 4.0.133 (still not imported), mcp 2.0.69, openai
4.0.86, anthropic 4.0.74, google 4.0.90, google-vertex 5.0.104,
amazon-bedrock 5.0.108, azure 4.0.94, xai 5.0.17, groq 4.0.57, cohere
4.0.57, mistral 4.0.59, deepseek 3.0.61, moonshotai 3.0.65,
black-forest-labs 2.0.57, bytedance 2.0.59, fal 3.0.57, klingai 4.0.58,
replicate 3.0.57, plus @openrouter/ai-sdk-provider 3.1.0. Adds
@ai-sdk/provider-utils 5.0.56 (the version ai pins) for the type-only
ProviderOptions import: ai does not export it, and SharedV3ProviderOptions
is tied to provider spec V3. undici moves 6.29 -> 7.30 for
provider-utils 5; it is still loaded through createRequire, so the
standalone-server include from the zod/TS lane still applies.

Codemods (@ai-sdk/codemod 4.0.3 v7, each run through jscodeshift so the
result was visible): kept isStepCount, system -> instructions (title,
router, tool repair), onStepEnd/onEnd, transcribe and createGoogle.
Reverted the false positives: theme `system` keys renamed to
`instructions` (appearance page, workspace sidebar), Sentinel's own
metadata.usage.reasoningTokens and the Codex protocol usage rewritten to
outputTokenDetails, a doubled outputTokenDetails in runtime/reasoning.ts,
and `context: experimental_context` in prepareStep.

Silent v7 changes handled:
- Thread agent: runtimeContext replaces experimental_context in
  prepareCall/prepareStep, and repairToolCall replaces
  experimental_repairToolCall. prepareStep now returns the full
  instructions from every branch: v7 carries returned instructions into
  later steps, so a validation or task directive would otherwise stick
  to every later step. Without a reroute the default branch reproduces
  prepareCall's instructions exactly; after a reroute it describes the
  rerouted tool set instead of the stale initial prompt.
- allowSystemInMessages: true on the agent. Context compaction feeds its
  summary back as a synthetic role "system" message built server-side,
  which v7 rejects by default.
- createAgentUIStream is replaced by createThreadAgentUIStream
  (agent/ui-stream.ts): same flow, but the transcript is only validated
  structurally, as in ai 6.0.116. The agent builds its tools in
  prepareCall, so agent.tools is empty there. Since ai 6.0.301 the stock
  helper throws "No tool schema found for tool part edit" when the
  transcript holds an approval-responded tool part (approving a tool call
  failed), and v7 also replaces every earlier static tool output with
  "Tool output omitted because the tool is no longer available."
- createUIMessageStream gets an explicit onError; v7 redacts errors to
  "An error occurred." by default. The agent stream already passed one.
- Reasoning tokens come from totalUsage.outputTokenDetails.reasoningTokens
  (the top-level field is gone).
- generateObject -> generateText with Output.object (memory autosave,
  commit messages).
- validateUIMessages drops Sentinel's approval.decision and
  approval.response (6.x did too); messages/ui.ts copies them back after
  validation. Dynamic tool titles now survive v7's schema (6.0.301
  dropped them).
- Ollama uses the OpenAI Chat Completions model (.chat) through
  createProviderLanguageModel; languageModel() on an OpenAI provider is
  the Responses API.
- The OpenRouter providerOptions key is now "openrouter" (3.x ignores
  "openai"). No OpenRouter catalog model has a reasoning config yet, so
  nothing was sent under either key; P8 must use OpenRouter's
  `reasoning: { effort }` payload when it adds one.
- MCP 2: the HTTP transport sets redirect: "follow" (2.x default
  "error"), and both transports set protocolVersionDiscovery: false
  (2.x probes server/discover before initialize), keeping the 1.x
  handshake for user-configured servers. The OAuth provider reports its
  client info as dynamically registered, so invalid_client still leads
  to re-registration as in 1.x.
- createVertex -> createGoogleVertex.

Decisions and holds:
- OpenAI Responses function tools keep @ai-sdk/openai 4's explicit
  strict: false. 3.x omitted strict, which let the Responses API apply
  its own strict normalisation (optional tool parameters arrived as
  empty strings, vercel/ai#11869).
- OpenAI reasoningSummary default ("detailed" once an effort is set):
  nothing to change, Sentinel already sends reasoningSummary "detailed"
  with every effort.
- Tool-level needsApproval stays (still honoured, now under test); the
  move to toolApproval is P14.
- xAI 5 is Responses-only and createXai().languageModel() maps to it.
  The catalog's grok-3, grok-4 and grok-4-fast ids were retired by xAI
  on 2026-05-15 and redirect to grok-4.3; the catalog refresh is P8.
- streamText now runs tools after the model call finishes instead of
  mid-stream; no test depended on the old timing.
- Not marked breaking: release-please has bump-minor-pre-major false,
  so `!` would cut 1.0.0.

Tests: agent/ui-stream.test.ts drives a real ToolLoopAgent with
MockLanguageModelV4: an approved tool call runs and earlier tool outputs
reach the model, needsApproval still yields an approval request,
provider error text is not redacted, the compaction system message is
accepted, and usage/reasoning tokens map into thread metadata. The agent
test checks the per-step instruction reset and allowSystemInMessages,
messages/ui.test.ts the approval fields and titles, the factory test
Ollama chat and OpenAI/xAI Responses, the MCP test the redirect and
discovery options, and the orchestrator test both onError hooks. Mocks
use the v7 names (isStepCount, onStepEnd, onEnd, embeddingModel). The
tool JSON-schema snapshot is unchanged.

Tripwires (P7): stepCountIs, experimental_context,
experimental_repairToolCall, SharedV3ProviderOptions, generateObject,
experimental_transcribe, onStepFinish, createGoogleGenerativeAI.

Gate G0: frozen install ok, typecheck ok, 191/191 test files pass,
format ok, tripwires clean (13 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
extractTaskState read step.toolResults[].result, the AI SDK 4 shape.
Since AI SDK 5 a tool result carries the tool's return value in .output,
so the task map was always empty: the allTasksResolved stop condition
never fired and buildStepProgressAddon always returned "". It now reads
.output, so the agent stops once every tracked task is completed or
blocked, and steps with open tasks get the "Step Progress" directive.

Behaviour to watch: the loop now stops right after the step whose
manage_task call resolves the last task, as the stop condition intended,
so a run can end without a separate closing message after that call.

The tool router's evidence builder had the same dead branch ("result" in
toolResult ? toolResult.result : toolResult); it now reads .output
directly, which is what the old fallback always resolved to.

A test drives prepareStep and the stop condition with AI SDK 7 tool
result objects and checks that the AI SDK 4 shape no longer counts; it
fails on the previous commit. Tripwire P7-ai-v4-tool-result-shape
flags `toolResult(s)` followed by `.result` and `result.result`.

Gate G0: frozen install ok, typecheck ok, 191/191 test files pass,
format ok, tripwires clean (14 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The previous commit made the allTasksResolved stop condition work, but it
only saw manage_task results from the current agent.stream call. Each
approval resume or follow-up is a new call, so a continuation run that
updated one task stopped as soon as that task was terminal, even with other
plan tasks still pending.

- The orchestrator passes the plan's task statuses from run start as the
  new planTasks call option; prepareCall keeps them and this run's
  manage_task results are applied on top.
- Task state only counts once this run has changed a task, so leftover or
  already finished plan tasks never stop or steer an unrelated run.
- The stop fires one step after the last open task resolves. That step lets
  the model report back (a text answer ends the loop by itself), and a tool
  call there ends the run. Opening a new task in it keeps the run going.
- The Step Progress addon drops the step number, so the instructions only
  change with the task counts and no longer break prompt-prefix caching on
  every step.

Tests: loop.test.ts drives the real ToolLoopAgent from createThreadAgent
with MockLanguageModelV4 (routing, tools and instructions stubbed). It
checks that a step directive reaches exactly one model call (AI SDK 7
carry-forward), that the run continues while earlier plan tasks are open,
that resolved plan tasks do not stop a run, and the reporting step. Two of
these fail on the previous commit. Unit tests cover the stop condition and
progress text, and a thread-chat test checks planTasks is passed.

Gate G0: frozen install ok, typecheck ok, 192/192 test files pass, format
ok, tripwires clean (14 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ictness

xAI: @ai-sdk/xai 5 only has the Responses API, whose `store` option
defaults to true. xAI then stores prompts and responses for later
retrieval, and Zero Data Retention teams get an API error. Sentinel's
earlier Chat Completions path stored nothing and Sentinel never reads
stored responses back. createProviderLanguageModel now wraps xAI models in
defaultSettingsMiddleware with providerOptions.xai.store = false, so every
caller gets it (the agent, routing, titles, memory, commit messages). With
store off the provider also requests reasoning.encrypted_content, so
reasoning still round-trips. A per-call xai option still wins.

OpenAI/Azure strict tools (decision, no runtime change): AI SDK 7 sends
Responses function tools with `strict: false` unless a tool opts in. AI SDK 6
omitted the field. The Responses API then tried strict mode when a schema
looked compatible and normalised it, which made optional parameters arrive
as "" (vercel/ai#11869, fixed by vercel/ai#15889). That omitted-field
behaviour cannot be reproduced from the client, and sending strict: true
has no non-strict fallback (an unsupported keyword fails the whole request)
and would also switch on strict tools for Anthropic and others. So Sentinel
keeps strict: false. A factory test pins the value sent for OpenAI and
Azure so a provider bump cannot change it silently.

Tests stub fetch and check the request bodies: xAI sends store false and
the encrypted reasoning include (fails without the wrapper), and OpenAI and
Azure send strict: false. The agent and thread-chat tests that mock 'ai'
by name now also export defaultSettingsMiddleware and wrapLanguageModel.

Gate G0: frozen install ok, typecheck ok, 192/192 test files pass, format
ok, tripwires clean (14 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
AI SDK 7 runs the start/status polling for xAI, Kling AI and ByteDance
video models itself (experimental_generateVideo's `poll`). The providers
now ignore pollTimeoutMs in providerOptions, and ByteDance also returns a
"deprecated setting" warning that generate_video copied into the tool
output shown to the user and the model.

buildVideoProviderOptions becomes buildVideoPollOptions and passes
poll: { timeoutMs: 600_000 } for those three providers. That is the SDK
default today, so timing does not change, but the setting is live again
and the warning is gone. These video models have no doGenerate, so passing
poll does not change which flow the SDK picks. Other providers get no poll
option, as before.

The video test now checks the poll option and that no providerOptions are
sent.

Gate G0: frozen install ok, typecheck ok, 192/192 test files pass, format
ok, tripwires clean (14 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- next 16.4.0, react/react-dom 19.3.0, @types/react(-dom) 19.3.0
- next.config.js: drop the eslint key (removed in Next 16) and set
  compress: false; this is a loopback desktop app and Next's compression
  buffers streamed responses (the chat resume SSE route already forced
  Content-Encoding: none; the planned tRPC SSE subscription needs the same)
- next.config.js: agentRules: false; Next 16 `next dev` otherwise writes a
  managed AGENTS.md into the repo root whenever it detects a coding agent
- dev/dev:desktop: drop --turbo, Turbopack is the Next 16 dev default
- scripts/next-build.mjs: production builds pass --webpack (decision D6:
  Turbopack standalone output with serverExternalPackages can emit aliased
  externals/symlinks that break desktop packaging);
  SENTINEL_NEXT_BUNDLER=turbopack opts into Turbopack
- tripwires: --turbo in package.json, eslint key in next.config

Audit: request APIs are already async, the only parallel slot (@settings)
has default.tsx, no middleware, images, revalidateTag, runtime config,
AMP or process.argv checks; no global smooth scroll-behavior, so no
data-scroll-behavior. React 19.3 StrictMode double-invokes effects during
hydration in dev; the thread session store refcounts subscribers and
aborts the superseded SSE fetch, so only one stream stays open.

Hold: next dev/build rewrite tsconfig.json (mandatory jsx: react-jsx, plus
.next/dev/types in include); tsconfig is owned by the TS 6 lane, so those
edits are left for it.

Gate G0: frozen install ok, typecheck ok, 183/183 test files pass,
format ok, tripwires clean (3 active). next build --webpack and an
isolated next dev (port 3300) verified at the lane tip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- @heroui/react and @heroui/styles 3.2.6; 3.2.1+ no longer bundles
  react-aria, so its peers are now direct dependencies: react-aria 3.52.1,
  react-aria-components 1.21.1, @react-aria/ssr 3.10.1, @react-aria/utils
  3.34.1, @internationalized/date 3.12.4 (single copy of each in bun.lock)
- 3.2.0 makes Switch/Checkbox/Radio `*.Content` the clickable <label> and
  requires the control inside it. All 33 toggles are migrated: control-only
  switches wrap the control in `*.Content`; labelled ones nest the control
  in `*.Content` with font-normal (Content now sets font-medium, which
  descriptions would otherwise inherit); option cards in the user-input and
  plan renderers move their card styling onto `*.Content` so the whole card
  stays clickable, keeping label and description inside it. The plan
  renderer keeps the control offsets HeroUI used to apply.
- globals.css: drop the `--default-hover` overrides. 3.0.1 ignored them
  (hover was derived from `--default`); 3.0.5+ reads them, which would have
  made light-mode hovers lighter than the default background
- globals.css: zero ScrollShadow's new 10px scrollbar gutter in its fade
  mask; Sentinel hides scrollbars, so the gutter showed as an unfaded strip
- workspace sidebar: overlay triggers are inline-block since 3.0.2; keep
  the linked-folder tooltip trigger block-level so the row stays full width
- test: controlled switch/checkbox fields render control, input and label
  inside the clickable content (fails on the old composition)

Audited, no change needed: Text->Typography (unused), tooltip delays (every
tooltip sets `delay`; close delay is React Aria's 500ms as before),
Tabs.ListContainer, Select/ComboBox/ListBox, Modal/Drawer (overlay z-index
now 100000), Toast (Sentinel uses sileo). Accepted upstream restyles:
tooltip padding p-2, radius tokens, accessible soft-foreground palette.

Gate G0: frozen install ok, typecheck ok, 184/184 test files pass,
format ok, tripwires clean (3 active). G1: next build --webpack ok;
standalone keeps better-sqlite3 (with .node) and sqlite-vec. Isolated
next dev on :3300: settings switches render and toggle, label text toggles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Next 16 treats `jsx: react-jsx` as mandatory and rewrites tsconfig.json
on every build or dev run when it is set to `preserve`, which left the
tree dirty. Set it explicitly (`.next/dev/types` is already included).

Gate G0 (rebased on zod 4 + TS 6 + AI SDK 7): typecheck ok, all test
files pass, format ok, tripwires ok; next build (webpack) ok and leaves
tsconfig.json untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…qlite3 13

electron 44.6.0, electron-builder 26.17.0 (v26 tag), better-sqlite3 13.0.3
and @types/better-sqlite3 9.6.0 move together because the server's native
SQLite module has to load under Electron's embedded Node (24.21).

Electron binary (42+ no longer downloads it in postinstall):
- scripts/desktop/electron-binary.mjs runs the package's install-electron
  script when node_modules/electron/dist is missing or from another
  version; exposed as `bun run electron:install`
- dev-launch.mjs and the packaging preflight call it, so
  build.electronDist and the dev launcher always find a binary
- .github/actions/setup-desktop-build runs it, covering desktop-verify
  and publish-release

Main process (desktop/main/index.mjs):
- await clipboard.writeText (returns a Promise in 44)
- console-message listener reads the details object (level is now a
  string; the positional form is deprecated and logs a warning)
- Electron 43 opens dialogs in Downloads when no defaultPath is passed;
  remember the folder of the last pick so the workspace and file
  pickers reopen there, as the OS did before
- checked webview + did-attach-webview, guest debugger attach,
  setWindowOpenHandler, webRequest, permission handlers, media access,
  nativeImage, dock, dialogs and updater net.request against the 41-44
  breaking changes; no other changes needed. The app uses no login
  items, Notification or renderer clipboard.

Packaging:
- drop the ia32 and armv7l targets (no Electron 44 binaries)
- better-sqlite3 13 is N-API with bundled prebuilds: the packaged copy
  keeps only prebuilds/<platform>-<arch>.node and drops binding.gyp,
  build, deps and src instead of rebuilding against Electron headers
  (cross-arch builds now use the bundled prebuild too). The host smoke
  test (select 1 under ELECTRON_RUN_AS_NODE) stays
- rebuild-node-native.mjs only smoke-tests better-sqlite3 and keeps the
  node-pty repair path; node-gyp is now a declared devDependency for it
  (Linux has no node-pty prebuilds)
- audit-bundle asserts the packaged better-sqlite3 has the target
  prebuild and no other prebuilds or build inputs

Docs: README and docs/product install list macOS 13+, 64-bit Windows 10+
and Linux x64/arm64, plus `bun run electron:install`.

Tripwires (P6): ia32/armv7l in scripts/desktop, .github and package.json;
prebuild-install in scripts; ELECTRON_SKIP_BINARY_DOWNLOAD; un-awaited
clipboard.writeText and positional console-message in desktop/.

Holds: node-gyp stays on 12.x (13 needs Node >=24.15; the repo allows
24.14). electron-updater stays 6.8.9 (6.8.10 has the same
builder-util-runtime 9.7.0).

Gate G0: frozen install ok, typecheck ok, 185/185 test files pass,
format ok, tripwires clean (6 active). Unsigned build:desktop:mac and
its bundle audit pass; the packaged binary (ELECTRON_RUN_AS_NODE) loads
better-sqlite3 13 and runs select 1. sqlite-vec does not load from the
packaged server; fixed in the next commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sqlite-vec 0.1.9 locates vec0.<ext> with a dynamic require.resolve of
sqlite-vec-<os>-<arch>, which Next's output tracing does not follow
(0.1.7-alpha.2 joined a path from __dirname, which it did). Since the
sqlite-vec bump the packaged server shipped sqlite-vec without its
platform package, so initVectorDb failed quietly and knowledge vector
search was off in packaged builds (the v0.0.66 release still has
sqlite-vec-darwin-arm64).

- rebuild-standalone-native copies the target's sqlite-vec platform
  package next to sqlite-vec (an error on host builds when it is not
  installed, a warning on cross-target builds) and its Electron-as-Node
  smoke test now also loads sqlite-vec and reads vec_version()
- audit-bundle fails when the packaged server has sqlite-vec but not
  node_modules/sqlite-vec-<os>-<arch>/vec0.<ext>

Gate G0: frozen install ok, typecheck ok, 185/185 test files pass,
format ok, tripwires clean (6 active).
G2 packaging: unsigned `bun run build:desktop:mac` passes (preflight,
host smoke "better-sqlite3 select 1 = 1" and "sqlite-vec v0.1.9",
electron-builder 26.17.0, bundle audit). Packaged import smoke with
ELECTRON_RUN_AS_NODE from Resources/server, using both Sentinel and
"Sentinel Helper (Plugin)": Electron 44.6.0 / Node 24.21.0,
better-sqlite3 13.0.3 select 1, sqlite-vec v0.1.9 loaded, node-pty
spawned a pty. The packaged server.js on port 3399 with an isolated
HOME and state paths served /api/health and DB-backed tRPC queries
(migrations ran, vectors.db created).
Sizes vs the v0.0.66 release: dmg 162.7 MB vs 146.8 MB (+10.9%),
zip 158.9 MB vs 142.4 MB (+11.6%), Resources/server/node_modules
94.8 MiB vs 86.6 MiB (+9.4%).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Electron 44 needs macOS 13 and the packaged Info.plist now says so, but
electron-builder 26.17 writes no minimumSystemVersion into latest-mac.yml.
Installed builds (electron-updater 6.8.3 in v0.0.66) auto-download any
newer feed entry, so a macOS 12 user would be updated to an app that no
longer opens.

- scripts/desktop/update-feed.mjs reads LSMinimumSystemVersion from the
  packaged .app, maps it to the Darwin kernel version electron-updater
  compares with os.release() (macOS 13 = 22.0.0) and writes it into every
  dist/*-mac.yml. Unmapped or minor-version floors fail loudly instead of
  guessing.
- package.mjs stamps the feed right after electron-builder (builds run with
  --publish never; publish-release uploads dist/latest-mac.yml afterwards).
- audit-bundle fails a mac build whose feed is missing the field or
  disagrees with the bundle's floor.
- Tests cover the mapping, idempotent stamping, the parsed feed under
  electron-updater's own js-yaml/semver (Darwin 21.6 blocked, 22.1
  allowed) and the package/audit wiring.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install,
typecheck, 186/186 test files, format check, tripwires clean).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
better-sqlite3 13 dropped its prebuild-install script and sets
`gypfile: false`, but bun still runs an implicit `node-gyp rebuild` for
any trusted package with a binding.gyp. The rebuild compiles nothing (the
binding skips when a prebuild exists), yet node-gyp configure needs
Python, and Visual Studio on Windows. With no Python on PATH, a frozen
`bun install` of the previous commit fails with "install script from
better-sqlite3 exited with 1". 12.x only needed prebuild-install there.

- package.json gets an explicit trustedDependencies list: electron-winstaller,
  esbuild, lefthook, node-pty and sharp, which are the packages whose
  scripts run today under bun's default allowlist. better-sqlite3 is left
  out, and tesseract.js stays blocked as before. An explicit list replaces
  bun's default one, so package-config.test.ts now scans node_modules and
  fails when any installed package with install scripts (or an implicit
  gyp build) is neither trusted nor deliberately untrusted.
- Docs: README, install.md and environment-and-build.md no longer promise
  a better-sqlite3 source build. They say it loads its bundled N-API
  prebuild, document the Linux glibc 2.34 / libstdc++ (GCC 11) floor of
  those prebuilds (GLIBC_2.34 and GLIBCXX_3.4.29 in prebuilds/linux-*.node),
  say build tools are only needed for node-pty (always on Linux), note
  that macOS 12 installs stop getting updates, and add the
  `bun run electron:install` note to environment-and-build.md.
- rebuild-node-native's better-sqlite3 error names the glibc floor on
  Linux. Its comment no longer mentions an install-time source build.

Verified: a full `bun install --frozen-lockfile` with PATH limited to
node, bun, git and /bin (no python3) exits 0, and better-sqlite3 select 1
plus require('node-pty') work. The same install of the parent commit
fails in node-gyp configure.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install,
typecheck, 186/186 test files, format check, tripwires clean).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Electron 43 opens dialogs in Downloads when defaultPath is omitted. The
earlier workaround kept one in-memory folder shared by the workspace and
file pickers. So the first dialog after every launch still opened in
Downloads, and attaching files moved the workspace picker to the
attachment folder. Before Electron 43 the OS reopened the last folder,
even across restarts.

- dialog-paths.mjs now keeps the last folder per dialog kind ("directory"
  and "files") in userData/dialog-paths.json, using the existing
  scripts/desktop/state.mjs helpers, which are already packaged. Until the
  first pick, or once the remembered folder no longer exists, it falls back
  to the home folder rather than Downloads. Unreadable state is ignored,
  and a failed save only logs a warning.
- The store is created in registerIpc (after app ready) so userData is
  final. PICK_DIRECTORY and PICK_FILES each use their own kind.
- Tests cover the fallback, separate kinds, restoring after a restart,
  cancelled dialogs, deleted folders and corrupt state.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install,
typecheck, 186/186 test files, format check, tripwires clean).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pwires

The P6 tripwires only matched the exact old lines: an un-awaited
clipboard.writeText( and a `(_event, level, message` parameter list.
desktop/main is plain .mjs that typecheck does not cover, so other
Electron 44 breaks would only show up at runtime.

- P6-unawaited-clipboard (desktop/main): has, read, readText, write and
  writeText now return Promises and must be awaited or returned. This
  replaces P6-unawaited-clipboard-write.
- P6-removed-clipboard-api (desktop): availableFormats and the
  HTML, image, RTF, bookmark, buffer and find-text read/write methods.
- P6-renderer-clipboard (desktop/preload): Electron's clipboard
  imported or required in the preload, which Electron 44 removed from
  renderers.
- P6-positional-console-message now matches any console-message listener
  with a second positional parameter, including across line breaks and with
  other parameter names, plus named (e, level, message) style handlers.
- tripwires.mjs gains an opt-in `multiline` flag: the pattern runs on the
  whole file and the hit reports the line the match starts on. The
  per-file scan moves into an exported findTripwireHitsInContent so the
  patterns can be unit-tested. Line-based behaviour is unchanged.
- tripwires.test.ts checks hits and non-hits for each P6 pattern. Run
  against the pre-P6 desktop/main/index.mjs (b88ef14), the widened set
  reports both un-awaited writeText calls and the positional listener;
  the current tree has 0 hits.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install,
typecheck, 187/187 test files, format check, tripwires clean, 8 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The packaged server resolves undici from the top-level node_modules
(provider-utils loads it through an untraced createRequire). After the
Electron 44 lane added node-gyp 12 (undici ^6) bun hoisted undici 6,
which no longer satisfies @ai-sdk/provider-utils 5 (^7.29), so default
downloads in packaged builds would have loaded the wrong major.

- declare undici ^7.30.0 directly so the hoisted copy is the one the
  server needs; build tooling keeps its own nested copy
- scope the trace test to the packages that load undici at runtime
  (UNTRACED_SERVER_PACKAGE_LOADERS) instead of every installed dependent

Gate G0: typecheck ok, all test files pass, format ok, tripwires ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sentinel engine's MODEL_CATALOG stopped at claude-opus-4-6 and still
listed models that providers have shut down. Every provider list now
starts with its current flagship, verified against the installed
@ai-sdk/* model-id unions, the providers' model and deprecation pages,
the AI Gateway model list and OpenRouter's /api/v1/models:

- OpenAI: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol/Luna, GPT-5.6 Sol/Terra/Luna
  ahead of the GPT-5.x line; 922K/272K prompt windows.
- Anthropic: Opus 5.5, Sonnet 5.5, Fable 5.1, Haiku 4.5, then Opus 5,
  Sonnet 5, Fable 5, Opus 4.8/4.7/4.6, Sonnet 4.6 and legacy 4.x; 1M.
- Google and Vertex: Gemini 3.8 Flash, 3.1 Pro preview, 3.7/3.6/3.5
  Flash, 3.5 Flash-Lite, 3 Flash preview, then the 2.5 models.
- xAI Grok 4.7/4.6/4.5/4.3/4.20, DeepSeek V4 Pro and Flash, Kimi K3,
  K2.7 Code and K2.6, Mistral Large/Medium/Small -latest, Codestral,
  Ministral 14B, Cohere Command A Plus, Groq GPT-OSS, Bedrock Claude 5.5
  inference profiles, Gemma 4 / Qwen 3.5 / GPT-OSS on Ollama, and current
  Gateway and OpenRouter ids (dotted versions, spacexai/ and x-ai/).

Retired, renamed or never-valid ids move to RETIRED_MODEL_REPLACEMENTS
with the provider's recommended successor (gpt-5-codex, o1, o3-mini,
o4-mini, gpt-4.1-nano, claude-opus-4-1, Claude 3.x, Gemini 1.5/2.0 and
3 Pro preview, grok-3/4, deepseek-chat/-reasoner, moonshot-v1, Kimi
K2/K2.5, Magistral 2507, Pixtral Large, ...). Threads, automations,
stored defaults (normalizeSelectedModelId) and the composer resolve a
retired id to its successor, unless the user still has that id enabled
(for example as a custom model). Unknown ids keep their previous path.

Reasoning options follow each provider's current values without
switching to AI SDK 7's top-level `reasoning`: Claude 4.6+ send effort
with adaptive, summarized thinking (Opus 5.5 defaults to medium; xhigh
maps to max on Opus/Sonnet 4.6), Gemini 3.x thinking levels with thought
summaries and per-model defaults, GPT-6/5.6 efforts, Codex models low to
xhigh, o3 low to high, GPT-5 sends `minimal` instead of the unsupported
`none`, Grok 4.5+ up to xhigh, DeepSeek V4 thinking plus reasoningEffort,
Kimi K3 low/high/max (max by default), Mistral high/none.

Helper models (titles, tool routing) are catalog entries again: deepseek
gets deepseek-flash instead of an OpenAI fallback, Gateway/OpenRouter
use the provider-prefixed google/gemini-3.5-flash-lite, Bedrock uses the
us.anthropic.claude-haiku-4-5-20251001-v1:0 profile, Ollama llama3.2,
Moonshot kimi-k2.6, OpenAI gpt-6-luna. resolveHelperModel asks them for
the least reasoning they accept (the router previously hard-coded
`minimal`). The Codex fallback list follows the catalog to GPT-6 Astra,
6.1 Sol and Luna (same ids as the P9 Codex lane, which replaces it).

Native attachments now derive from the vision capability of built-in
anthropic/google/vertex/openai models instead of a duplicate allow-list.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 199/199 test files, format, 25 tripwires) on the lane tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The media catalogs pointed at models the installed providers no longer
serve: @ai-sdk/google 4 and @ai-sdk/google-vertex 5 throw for any image
model id that does not start with `gemini-` (so both Imagen entries were
dead), OpenAI shut DALL-E 2/3 down on 2026-05-12 and shuts gpt-image-1
down on 2026-10-23, and Google shut Veo 2/3.0 down on 2026-06-30.

Images: GPT Image 2.5 Sunburst/Flare and GPT Image 2 (with reference
image edits, which @ai-sdk/openai supports for them) ahead of the
expiring GPT Image 1.5/1 Mini; Nano Banana 2.1 and Gemini 3 Pro Image
for Google, Nano Banana 2.1 and Gemini 2.5 Flash Image for Vertex (no
mask editing: Gemini image models reject masks); Grok Imagine Image 2.0;
FLUX 3 (reference images, no seed). Video: Veo 3.1 GA ids first on
Vertex, the Veo 3.1 previews on AI Studio, Grok Imagine Video 1.5.

A stored image or video model that a provider retired now resolves to
its replacement (RETIRED_IMAGE_MODEL_REPLACEMENTS /
RETIRED_VIDEO_MODEL_REPLACEMENTS) instead of leaving the provider with
no valid model; custom ids keep their previous handling.

Transcription: OpenAI deprecated whisper-1 and the GPT-4o transcribe
models (shutdown 2027-02-26); gpt-transcribe, on the same
/audio/transcriptions endpoint, is the new default for users who never
picked a model. Explicit choices are kept and still listed.

Embeddings: add gemini-embedding-2 (Gemini API and Vertex, 3072 dims)
and Cohere embed-v4.0 (1536 dims) profiles. Existing profile ids and the
text-embedding-3-small default are unchanged, because stored memory is
tied to its profile's vector size.

Tripwire P8-retired-image-catalog-ids keeps Imagen and DALL-E out of the
image catalog.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 199/199 test files, format, 25 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What and why:
- Mistral: mistral-large-latest is Mistral Large 3 (mistral-large-2512, GA,
  256K, no adjustable reasoning) per docs.mistral.ai/models, and
  @ai-sdk/mistral 4.0.59 only sends reasoning_effort for ids on its own
  allow-list (mistral-large-4 but not mistral-large-latest), so the effort
  picker on the default Mistral model did nothing. mistral-large-latest is
  now listed as Large 3 without a reasoning config (262,144 context), and
  Mistral Large 4 (public preview since 2026-10-06) is added as
  mistral-large-4 with the none/high config (524,288 context, the cap the
  AI Gateway and OpenRouter serve; Mistral lists 1M). The GA model stays
  first, so the default for new Mistral threads is unchanged from before
  the refresh.
- Claude preserved thinking: Opus 5.5, Sonnet 5.5 and Fable 5.1 bind
  thinking blocks to the conversation, and accounts created on or after
  2026-08-31 get a 400 when a replayed block's earlier history changed.
  Sentinel's compaction keeps recent turns behind a new summary and the
  tool router changes the active tool set between steps, so these models
  now send thinking.blockBinding.prefixMismatchBehavior "drop_block"
  (@ai-sdk/anthropic adds the thinking-binding-controls-2026-08-01 beta),
  with or without a selected effort. Other Claude models are unchanged.
- xAI: grok-4.5 offers low/medium/high (xAI's reasoning guide says xhigh
  exists on grok-4.6+ and is treated as high on 4.5); grok-4.3 drops xhigh
  (its model page lists none/low/medium/high).
- Retired ids: resolveStoredCompositeModelId only switches to the
  successor when the user has it available, matching
  normalizeSelectedModelId; otherwise the stored id passes through to the
  caller's existing fallback, as before the refresh.
- Azure: gpt-5 is first again. Azure ids are deployment names, and gpt-5
  was the default before the refresh.
- Bedrock: Claude Sonnet 4.5 has no in-region endpoint, so the entry is
  us.anthropic.claude-sonnet-4-5-20250929-v1:0 and the bare id maps to it.

Tests:
- New reasoning-requests.test.ts sends every effort the picker offers for
  every catalog model of anthropic, cohere, deepseek, google, groq,
  mistral, moonshotai, openai and xai through the real provider package
  and asserts the value reaches the request body. Before the Mistral fix
  it failed on mistral-large-latest; every other provider passed. It also
  checks the drop_block request body and beta header for a compacted
  transcript that replays a signed thinking block.
- models.test.ts and context/model.test.ts updated for the above.
- Tripwire P8-bedrock-in-region-claude-ids forbids bare anthropic.claude-*
  catalog ids (zero hits).

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install,
typecheck, all 200 test files, format check, tripwires clean, 26 active).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bump @anthropic-ai/claude-agent-sdk ^0.2.84 -> ^0.3.292 (bundles Claude Code 2.1.292) and add its
new peers @anthropic-ai/sdk ^0.131.0 (peer >=0.93) and @modelcontextprotocol/sdk ^1.32.1 (peer
^1.29); zod 4.6.5 stays the only zod. The old SDK's @img/sharp 0.34 optional deps drop out, so
sharp 0.35.5's own platform packages dedupe to the top level.

Compatibility with the new SDK:
- options.env replaces process.env since 0.2.113: buildClaudeSdkEnv always layers the runtime env
  over process.env (run, probe, and any caller of buildClaudeSdkBaseOptions).
- permissionMode is always explicit (base default "default"); omitted it follows the settings
  defaultMode, possibly "auto", since 0.3.286. chat -> default (+sandbox), full ->
  bypassPermissions + allowDangerouslySkipPermissions, plan -> plan, as before.
- Native binaries are spawned by the SDK itself (stderr tail in exit errors, graceful shutdown).
  Only Node-script CLIs (.js/.mjs/.cjs or a node shebang, e.g. npm shims) get Sentinel's spawner,
  which runs them under process.execPath with ELECTRON_RUN_AS_NODE=1 (GUI launches often have
  no node on PATH). The same launcher verifies `claude --version`. Windows npm .cmd/.bat/.ps1
  shims resolve to bin/claude.exe or cli.js (ported from t3code ClaudeExecutable.ts, MIT).
- The chat-mode sandbox sets failIfUnavailable:false (0.2.91 default flip) so runs keep going
  unsandboxed when the sandbox cannot start; Bash auto-allow then turns off and Bash is prompted
  (checked in the CLI: auto-allow requires isSandboxingEnabled()).
- Status probe: a prompt stream that never yields (no turn can reach the API) and no hooks, MCP
  servers or IDE auto-connect (t3code ClaudeProvider probe options). initializationResult() and
  close() still exist; models/account are read as before.
- AskUserQuestion is answered the documented way: canUseTool resolves
  {behavior:"allow", updatedInput:{...input, answers}} with answers keyed by question text
  (multi-select comma-joined), the card's extra context as annotations notes, and unmatched free
  text as `response`. The synthetic tool_use_result user message is gone; an answer that arrives
  before canUseTool is held and applied when Claude Code asks.
- Task tools replaced TodoWrite (0.3.142) and are off by default on new models (0.3.233/0.3.268);
  native builds swap Grep/Glob for Bash find/grep (0.3.162). allowedTools names all six on top
  of the claude_code preset (an explicit tools list would freeze the tool surface). The mirror
  accumulates Task* results by task id (structured tool_use_result, seeded from the transcript
  when the session resumes) and stores the list on output.tasks; new claude_task* renderers show
  it. claude_todowrite stays for persisted messages.
- Effort is now sent: clamped to the model's supportedEffortLevels from the last probe snapshot,
  passed through for unknown models, none/minimal -> low. xhigh is offered; max waits for the
  P10 ReasoningEffort widening. Default efforts follow t3code's model manifest instead of the
  lowest level (which would now be sent). Thinking is adaptive with summarized display.
- rate_limit_event is recorded per window (and on the run control) for P11's usage limits, and
  logged at debug. A success result with is_error now fails the run with its error text instead
  of showing it as the answer. local_command_output and session_state_changed are unchanged.
- Context windows and fallback models follow the current lineup (Opus 5.5, Sonnet 5.5, Fable 5.1
  as default, Haiku 4.5; aliases via ModelInfo.resolvedModel; [1m] -> 1M; Opus 4.7/4.8 1M).
- Commit messages run the resolved Claude binary with its managed env, not `claude` from PATH.
- Image attachments in types the API rejects are sent as a text note.

Tripwires (zero hits; verified against the pre-change files): the synthetic AskUserQuestion
answer, TodoWrite-only Claude renderers, `env ?? process.env` in the Claude SDK options, and
spawning `claude` from PATH.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install, typecheck, 202 test files,
prettier, 27 tripwires clean). One run hit a load-sensitive Sentinel-engine test
(thread-chat/index.test.ts, unrelated) that passes alone 6/6 and on the rerun.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
chaqchase and others added 30 commits October 8, 2026 02:42
What: an auth panel in every engine instance card
(components/settings/engines/engine-auth-panel.tsx). It lists the methods
the instance offers (api.engines.auth.methods), starts and cancels flows,
and shows the step the server waits on:
- a sign-in page to open (desktop opens it in the system browser) with a
  link to copy;
- a device code to copy and the page to enter it on;
- on desktop, an embedded terminal running the CLI's sign-in
  (engine-auth-terminal.tsx, over terminal.createCommand and the flow's
  launch ticket); its exit is reported back to verify the result;
- in a browser (or when the terminal cannot start), the command to copy
  and a Done button;
- a credentials form (password inputs) whose values go to the server once.
Sign-out asks for confirmation first. The panel polls auth.status every
second while a flow runs, and EngineEventsBridge refetches it on auth
events. When a flow finishes, snapshots, the composer catalog, models and
methods are refetched (the server already re-probed the instance).

engine-instance-card.tsx gains one line (the panel) and
use-engine-snapshots.tsx a small handler for auth events.

Why: critique G1, driver-contract §2.5; the settings page only said
"Login needed".

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (275 test files);
bun run build passed (the new /api/engines/auth/terminal route is in the
route table, no client module pulls server code).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: loggerLink skips engines.auth.* (credentials, device codes, launch
tickets), engines.codex.login (API keys) and engines.instances.create /
update (secret instance variables), through isSensitiveTrpcOperation in
src/trpc/sensitive-operations.ts. Credential answers are also capped at
16 fields by the schema.

Why: in development loggerLink logs every operation with its input and
result, and the desktop app forwards the renderer console to its own
output, so a sign-in would have printed the key the user entered. In
production it logged failed operations, inputs included.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (276 test files).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…R helper

What: platform/auth/browser-helper.ts writes a /bin/sh script (a .cmd on
Windows) into an instance's state directory. With BROWSER pointing at it,
an agent that opens its own sign-in page prints the URL behind a marker on
stderr instead; parseAuthBrowserMarker reads it back (https, or http on
loopback, only) for the driver to show as a browser interaction. The
script never evaluates the URL. The composer's "needs authentication"
message now points to Settings → Engines.

Why: critique G13. Antigravity and ACP agents (P13) need this to sign in
from Sentinel, and the helper must not depend on a Node interpreter in
packaged builds (t3code uses a Node helper).

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t a runtime

What: the flow store catches a failure of its background runner and ends
the flow as failed instead of leaving it waiting until the TTL. The auth
panel offers "Sign out" only when the instance's runtime was found, so an
engine that is not installed shows its install hint rather than a
sign-out that cannot run.

Why: review of the P11 auth lane.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: claude-sdk exports forgetClaudeEngineStatus(instance), which drops
the instance's in-memory status caches (bumping the generation so a
background refresh cannot write them back) and deletes its on-disk
last-known-good snapshot. The Claude auth controller calls it after
`claude auth login`, after storing an API key, and after `claude auth
logout` (whether or not the command or the key removal did the work).

Why: the status probe answers an empty model list or a failure, which is
how a signed-out Claude Code looks, from that snapshot for up to 7 days,
forced refresh or not. The flow store's verifying refresh therefore still
saw "authenticated" with the old account after a sign-out, and the panel
reported "Signed out, but Claude Code still finds credentials", kept Sign
out visible and the composer kept offering Claude. A sign-in that
switched accounts could likewise show the previous account when the probe
timed out.

Tests: a probe without models after forgetting reports auth_unavailable
with no account and the snapshot file is gone; the controller forgets
after sign-in and sign-out.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files,
install, typecheck, format, tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t sign-out

What:
- EngineAuthController gains logoutNotice(instance), returned by
  api.engines.auth.methods as `logoutNotice` and shown in the panel's
  sign-out confirmation. sharedLoginNotice() (drivers/auth/shared.ts)
  covers engines whose login lives in the CLI's usual configuration unless
  the instance has a home of its own: Claude (CLAUDE_CONFIG_DIR), Codex
  (CODEX_HOME), Copilot (COPILOT_HOME). Cursor always warns (one login per
  OS user, shared by every Cursor instance); OpenCode warns unless the
  instance sets XDG_DATA_HOME. Copilot also says that it restarts on the
  instance, which stops chats running on it.
- Copilot sign-out removes only COPILOT_GITHUB_TOKEN, the variable this
  panel stores. GH_TOKEN / GITHUB_TOKEN set on the instance stay, and the
  outcome (or the failure) names them with how to sign out completely.

Why: signing out a default instance ran `claude auth logout`, `agent
logout` or account.logout against the user's global login with only
"Sign out of X?" as warning, and Copilot sign-out silently deleted token
variables a user may have set for MCP servers or gh in agent shells.
The Copilot sign-in/out restart is kept: like any change to an instance's
configuration (which the snapshot service already retires immediately),
the runtime has to restart to read the new login; the confirmation now
says so.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files,
install, typecheck, format, tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- The embedded terminal step remembers the command's exit and, once it
  exited, shows "Check sign-in" next to Cancel, which sends the exit code
  again. FlowStep is exported as EngineAuthFlowStep.
- Electron main ends command terminals (kind "command", the sign-in PTYs)
  when the main frame navigates to another document or its renderer
  process goes away. Shell terminals are left as they were.
- Render tests (renderToStaticMarkup) of the flow steps: embedded terminal
  with Cancel only, command to copy with Done, credentials in empty
  password fields with autocomplete off, device code, verifying.

Why: when respond({type: "terminal"}) failed after the PTY exited
(network error, server restart), the panel sat on an exited terminal with
only Cancel. After a renderer reload nothing can reattach to a command
PTY (its ticket is spent and the session pool is gone), so the login
process kept running until the app quit.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (278 test files,
install, typecheck, format, tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- The flow store's auth events now carry neither the interaction nor the
  message: flowId, instanceId, phase, methodId and purpose only. Clients
  already refetch auth.status (per user) when one arrives.
- toAuthTerminalInvocation passes an absolute cmd.exe for Windows .cmd /
  .bat shims (new resolveWindowsComSpec: ComSpec when absolute, else
  SystemRoot/windir, else C:\Windows, + System32\cmd.exe).
- EngineAuthClientCapabilities documents the rule for terminal methods:
  offered to every client, embedded on desktop and shown to copy in a
  browser (no launch ticket); a driver leaves one out only when it cannot
  work as a copied command.

Why: engines.onEvents streams every event to every subscriber, so driver
error text and messages naming the instance reached other users' streams
while the details stayed per-user behind auth.status. On Windows a
GUI-launched app can lack ComSpec; buildSpawnInvocation then falls back to
a relative "cmd.exe", which Electron main's absolute-path check refuses,
so the embedded sign-in terminal could never start for npm-installed CLIs.
The P13 note in the lane report said browsers should not get terminal
methods, which contradicts the behaviour the scope asks for (a copyable
command in browser mode); the contract now states the actual rule.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (278 test files,
install, typecheck, format, tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- flow-store.test.ts: a terminal interaction from the flow store goes
  through the renderer's toDesktopTerminalCommand and Electron main's
  prepareTerminalCommand, with a fetch that redeems the ticket as POST
  /api/engines/auth/terminal does. Main runs exactly the vended command,
  args (with a space), cwd and non-secret env, refuses the spent ticket,
  and refuses a renderer that changes the args (that ticket is spent too).
  Main's module is imported through a computed URL so tsc does not
  type-check its plain JS.
- copilot-sdk.test.ts: resolveCopilotLoginCli picks the user's own
  copilot on PATH when the instance runs on the bundled runtime, and the
  instance's runtime (as a Node script) when it is a CLI.

Why: both sides of the ticket check had their own hand-written fixtures,
so a drift between what the renderer echoes and what main compares would
only show in the app. resolveCopilotLoginCli had no tests.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (278 test files,
install, typecheck, format, tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The shared ACP engine (P12) replaced engines/cursor-acp; the Cursor auth
controller now finds `agent` with resolveCursorBinary from
acp/agents/cursor, the same lookup the driver uses.

Gate G0: typecheck ok, all test files pass, format ok, tripwires ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- contract/usage-limits: helpers for building, sorting and merging usage
  limits (sparse merge by window id that keeps known reset times,
  "unsupported" stays authoritative, a failed read keeps the last good
  windows), an update schema, and an optional `unavailable.action`
  ("read-keychain") the UI can offer.
- platform/usage/limits-store: the one place limits live per (user,
  instance). Full reads through `driver.usageLimits.read` at most once per
  TTL (5 min; 1 min after a failure; 30 min for accounts that cannot
  report), one at a time, bounded by a timeout that aborts the read; live
  reports merge in between. `reportEngineUsageLimits()` for runtimes.
- platform/usage/enricher: the "usage" snapshot enricher. It seeds limits a
  probe got for free, starts a background read when one is due and the
  instance is usable, and overlays what the store knows. The snapshot
  service takes the store as a dep (wired in getEngineSnapshotService):
  store changes are republished on cached snapshots, reportUsageLimits
  goes through the store, instance changes forget the instance's usage.
  Enrichment input gains `userId`.
- Readers (ported from t3code, MIT): Claude through the SDK's experimental
  usage request on an idle query (looked up at run time, unsupported when
  absent); Codex through the instance's app-server
  account/rateLimits/read (API-key and Bedrock sign-ins skipped); Cursor
  through the dashboard API with CURSOR_AUTH_TOKEN or the CLI auth file;
  OpenCode Go only when an `opencode-go` credential is configured in
  OpenCode, otherwise unsupported without any request.
- Cursor's macOS Keychain login is read with /usr/bin/security only from
  api.engines.usage.readCursorKeychain (an explicit user action); the
  token stays in process memory per instance and is dropped when Cursor
  refuses it. Without it, Cursor on macOS reports "read-keychain".
- Live updates: Claude rate_limit_event and Codex
  account/rateLimits/updated now feed the run instance's usage.
- api.engines.usage.{get, refresh, readCursorKeychain}.
- Catalog: reportsUsageLimits for codex, claude, cursor and opencode.

Why: P11 usage-limits row (t3features "Usage limits"); the contract had
the snapshot field and a passthrough "usage" enricher but nothing filled
them, and the existing Codex rateLimits endpoint had no consumer.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 277 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- Composer: a compact usage chip next to the context ring for the
  selected instance (engines whose driver reports usage limits). It shows
  the fullest plan window and, in its tooltip, every window with its reset
  time. Nothing renders before usage is known or for other engines.
- Settings → Engines: a "Usage limits" section listing usable instances
  that report usage, with bars per window, the time of the last read, a
  Refresh button (api.engines.usage.refresh) and, when Cursor on macOS
  needs it, "Read login from Keychain" behind a confirmation dialog that
  explains the macOS prompt (api.engines.usage.readCursorKeychain).
- useEngineUsageLimits(instanceId) reads api.engines.usage.get; the
  EngineEventsBridge folds snapshot events into those queries
  (syncUsageLimitsFromEvent), so the chip follows live rate-limit updates
  without polling while the event stream is up (5-minute poll otherwise).
- components/engines/usage-limits.ts: pure presentation helpers (tone,
  percent, reset phrasing, which instances to list, notices).

Why: P11 usage-limits UI (composer chip and settings section); the store
and readers landed in the previous commit.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 278 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- lib/ai/chat/engines/slash-commands.ts (client-safe): Codex's
  Sentinel-run commands (compact, review, rollback), normalisation of
  runtime-reported commands, the composer menu built from an instance's
  commands ("native" ones are inserted as `/name ` text for the runtime,
  "sentinel" ones run through a driver action), the long-standing
  per-driver commands as a fallback until an instance reports its own,
  and canRunSentinelSlashCommands for the thread gate.
- Claude: the status probe keeps the initialize response's `commands`
  (skills included) on the status and in its last-known-good snapshot;
  the driver reports them as the instance's slashCommands.
- Codex: the driver reports its three app-server commands once the CLI is
  detected.
- ComposerEngineOption carries `slashCommands` when an instance reports
  any (so the composer catalog refreshes when they change).
- Composer: the slash menu comes from the selected instance's commands,
  drops commands a listed skill already offers, and shows argument
  hints. Execution is a per-driver action map instead of a Codex-only
  branch; thread-screen gates it through canRunSentinelSlashCommands.

Why: P11 "Slash commands per agent" (t3features): commands were a static
Codex-only list plus a hard-coded Claude list that included commands the
SDK session does not run. Copilot, OpenCode and Cursor report none yet
(ACP available_commands_update arrives with P12).

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 279 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Claude skills natively

What:
- lib/skills: discovery, watching and lookups take `globalDirectories`,
  the global folder a source kind uses instead of `<base>/.claude/skills`
  and the like; snapshots are cached per workspace and set of folders.
  Installs and uninstalls take `globalSkillsDirectory` for global scope.
- lib/skills/instance-skills: Codex, Claude and Copilot keep global skills
  in `<home>/skills` (CODEX_HOME, CLAUDE_CONFIG_DIR, COPILOT_HOME, via
  platform/instance-homes). resolveSkillInstanceContext picks the selected
  instance for its driver and every other driver's default instance.
- api.skills: list, get, install and uninstall accept an optional
  `instanceId` and use those folders (installCustom and the registry use
  the default instances). Without an instance home the calls are exactly
  as before, so ~/.claude/skills, ~/.copilot/skills and skillsBasePath
  keep working.
- Composer: lists the selected instance's skills when it is not its
  driver's default instance.
- Claude: a `$skill` chip of a Claude skill becomes Claude Code's own
  `/skill args` invocation, sent as the message's last text block, with
  earlier mentions rewritten inline (planClaudeSkillDispatch, ported from
  t3code ClaudeSkillDispatch.ts, MIT). `.agents` skills stay prose.

Why: critique G11 and P10c deferred finding 4: Claude and Copilot skills
ignored an instance's CLAUDE_CONFIG_DIR or COPILOT_HOME (only Codex
followed its default instance's home), and Claude treated `$skill` as
prose.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 281 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: the usage enricher clears the store's limits for (user, instance)
and publishes no usage when the probe reports the instance
unauthenticated. The store gains clear(userId, instanceId), which also
drops a read still running for it.

Why: without it, the bars of the account that signed out stayed on the
snapshot (the store overlay and the previous-snapshot fallback both kept
them) until the next successful read after signing in again, possibly
with another account.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 281 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- contract/usage-limits: a runtime's live windows now replace an
  `unsupported` read, and resolveEngineUsageLimitsAfterRead takes
  `liveSinceLastRead`: an unsupported read keeps windows a runtime
  reported since the previous read instead of wiping them.
- limits-store: tracks live reports per entry and passes that to the
  merge; the TTL now counts from when a read starts, and a read is due
  10% before its TTL (USAGE_LIMITS_DUE_SLACK_RATIO).
- Claude reader: an SDK without the usage request is remembered per
  query factory, so no Claude Code process is started for it again;
  `rate_limits_available: true` with no `rate_limits` (usage endpoint
  failed) is now probeFailed, not unsupported; a subscription login that
  cannot be read on demand says its usage shows during runs.

Why (review findings):
- "unsupported" was authoritative, so a missing experimental SDK method
  or a token without profile scope silenced every live
  rate_limit_event for that instance, and the 30-minute re-read kept it
  that way. An account without plan limits never streams windows, so
  letting first-hand windows win costs nothing for API-key logins.
- Background reads start when the 5-minute refresh tick's probe ends and
  the TTL counted from when the read ended, so the next tick landed a
  few seconds short and reads ran about every 10 minutes.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 281 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hange

What:
- cursor-keychain: tokens are keyed by (user, instance);
  readCursorKeychainToken/getCursorKeychainToken take the user id and
  forgetCursorKeychainToken({instanceId?, userId?}) drops one user's,
  an instance's or all logins.
- Driver contract: usageLimits.read receives `userId` (the usage store
  passes it) and an optional usageLimits.forget(instanceId, userId?).
- The usage enricher calls forget for the account that signed out; the
  Cursor driver's invalidate (run by handleInstanceChange when an
  instance is reconfigured or removed) drops that instance's logins.
- api.engines.usage.readCursorKeychain passes the signed-in user.
- New routers/engines/usage.test.ts: the procedures hand the session
  user to the service, map EngineUsageError and Keychain errors to tRPC
  codes, and the Keychain is read only from its own procedure.

Why (review findings): the token sat in a process-global map keyed by
instance id only, so another session user with the same default
"cursor" instance reused a login they never approved, and a signed-out
or reconfigured instance kept reading usage with the old login.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 282 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- skill-dispatch: getClaudeSkillRoots({cwd, env}) lists the folders the
  CLI loads skills from (<cwd>/.claude/skills and <CLAUDE_CONFIG_DIR or
  ~/.claude>/skills); getClaudeDispatchSkillNames(context, roots) keeps
  a Claude skill chip only when its folder sits directly in one of them.
  Chips without a folder stay prose.
- run.ts computes the roots from the run's cwd and the env the CLI is
  started with (options.env, instance home included).

Why (review finding): dispatch trusted the composer's target/sourceKind.
With skillsBasePath set, Sentinel lists <base>/.claude/skills as Claude
skills that Claude Code never loads, so `/name` reached a CLI that does
not know it and the prompt could become an unknown-command reply. Such
chips now keep the pre-dispatch behaviour (prose plus the
<referenced-skills> prefix).

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 282 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- usage-limits: describeUsageCheckedAt (`Read at 14:32` today,
  `Read Oct 5, 14:32` before) and isUsageLimitsStale (instance not
  usable, or read more than 15 minutes ago).
- Settings → Engines → Usage limits: the label carries the day when the
  read is not from today (full time on hover), stale windows are dimmed
  with a warning-toned label, and a not-ready instance says the bars are
  its last known usage.
- The composer chip's tooltip shows when the usage was read.
- New engine-usage-limits.test.tsx renders the section (static markup):
  the Keychain offer reads nothing on render, carried-over windows are
  marked, and accounts that never report stay hidden.

Why (review finding): an instance that is installed but not usable
keeps the usage of its last snapshot (status.json, up to 7 days old),
and the row showed it as plain bars with a time of day only, which read
as current.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 283 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- slash-commands: getRunnableComposerSlashCommands(commands, {actions,
  canExecute}) keeps inserted ("native") commands always and
  Sentinel-run ones only when the composer has an action for them and
  can run it on this thread.
- chat-composer/index.tsx uses it for the menu instead of an inline
  filter (same result: the editor already hid execute commands without a
  handler).
- Test: Claude-reported commands are inserted with their hints on any
  thread; Codex compact/review/rollback are hidden without a Codex
  thread or an action.

Why (review finding): the generalized slash menu was only covered at the
helper level; the filtering that decides what a thread offers lived in
the component.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 283 test files, format, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hers

Snapshot enrichment inputs now carry the instance owner's userId (usage
limits are per user); the manifest and maintenance enricher tests build
their inputs with it.

Gate G0: typecheck ok, all test files pass, format ok, tripwires ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: when a Claude or Codex thread has no native session to resume (its
instance's continuation key changed, for example the home directory was
edited, or the stored session state is gone), the first prompt of the new
session now carries the thread's prior transcript, wrapped in a
<conversation_history> block ahead of the new message. Copilot already
did this; its bootstrap prompt moved unchanged into the shared
runtime/history-replay.ts, which Claude and Codex now use too.

- The history comes from the stored active transcript (truncated at the
  checkpoint anchor), never from client-sent messages: everything before
  the edited message for an edit, else the transcript without the message
  being sent.
- Claude and Codex keep the newest messages within 200k characters and
  say how many earlier ones were left out. Copilot's prompt is unchanged
  (golden test).
- A resumable session is still resumed without any replay, and a plan/chat
  mode switch on a resumable session still starts fresh without replay, as
  before.
- On Codex app-servers without collaboration mode, the replay precedes the
  plan-mode preamble.

Why: P10c deferred finding 7 (driver-contract §3 "null on mismatch ⇒ fresh
session + replay"). It must land before Settings exposes editing an
instance's home, which changes its continuation key.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 267 test files, prettier, 59 tripwires). The new Claude and
Codex tests fail with the replay disabled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: Settings → Engines lists each driver's instances under a driver
heading and can now add, edit, enable/disable, remove or reset them, and
edit their custom models (api.engines.instances, P10).

- Add instance (implemented external drivers): name, accent colour
  (presets or any #rrggbb), binary path, home directory for drivers that
  have one (CODEX_HOME, CLAUDE_CONFIG_DIR, COPILOT_HOME) and environment
  variables. The server derives the id from the name.
- Edit sends only what changed; config keeps the keys the dialog does not
  edit (launchArgs, …). Changing the home warns that threads on the
  instance start new native sessions with their conversation sent along
  (the replay landed in the previous commit).
- Environment editor: per-variable secret toggle. Stored secrets come back
  redacted and are echoed back redacted; typing replaces them; turning one
  into a plain variable needs a new value (mirrors the registry rule); a
  value that no longer decrypts is flagged "Re-enter".
- Each driver section says what isolates its instances: home-based drivers
  isolate sign-in and settings per home, while Cursor, OpenCode (and Pi
  later) share the CLI's sign-in and differ only by env and binary.
- Disable and remove surface the G9 in-use guard: the instance's
  references (threads, automations, user default) are fetched first, and a
  change to an instance in use asks to confirm, then runs with force.
  Default instances offer "Reset to defaults" instead of remove.
- Custom models editor per instance: id, display name, and options kept,
  dropped or copied from a model the engine reports; writes
  engine_instance.custom_models.
- The instance card shows its accent colour and takes an `actions` slot.

Logic lives in components/settings/engines/instance-management.ts with
unit tests; the env editor, card and driver section have render tests.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 270 test files, prettier, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: the composer's engine submenu (composer actions → Engine) now has
one entry per driver. A driver with a single instance stays a plain
entry; a driver with several becomes a section headed by the driver
label, with its instances indented underneath, each with its accent
colour dot (neutral when it has none). The "Engine" row shows the
selected instance's label, with its accent dot when its driver has
several instances, instead of the raw driver kind.

Grouping and the row summary are pure helpers in
chat-composer/engine-menu.helpers.ts with unit tests. Selection keys are
still instance ids, so nothing else in the composer changes; switching a
thread's instance stays out of scope (the P10c guard holds).

Why: P10c left a flat list of instances; with the add-instance UI a
driver can have several, which need grouping to tell apart.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 271 test files, prettier, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: the new-automation modal and the automation detail screen pick the
engine as a driver, plus an instance (with accent colours) when the driver
has several or the automation already uses a non-default one, and offer a
picker per model option (OpenCode's agent and variant, and any select
option a model or custom model describes, except the reasoning effort,
which keeps its own field). Picks are stored as automation.model_options.

- Each option picker starts at "Model default (…)"; values left there are
  not stored, so automations without picks behave as before (empty
  model_options, model defaults).
- Only values the selected model still offers are stored; changing the
  model drops picks the new model does not offer. Stored model_options are
  read back through parseEngineOptionSelections.
- Picking another driver keeps the current instance when it belongs to
  that driver, else takes the driver's default instance when it can run,
  else its first one that can.
- An automation whose instance is gone (removed or disabled) keeps it
  listed, disabled, and the form says to pick another engine instead of
  sending a bad request.
- "Use default model" clears the stored picks.

The shared pickers live in automations/automation-engine-fields.tsx, the
logic in automation-form-helpers.ts (getAutomationEngineOptions is
replaced by driver and instance options); both have tests.

Why: P10c deferred finding 5 (no UI wrote model options, so automations
always ran on model defaults) and the P11 automations instance picker.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 272 test files, prettier, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: saving an automation now writes model_options only when it knows
the selected model's options: null for "Use default model", the picks the
model offers when the model is in the catalog, and nothing (the stored
value is kept) when it is not, for example while the catalog loads or
while the instance is unavailable. resolveAutomationModelOptionsForSave in
automation-form-helpers.ts decides, with tests.

Why: the previous commit computed model_options from the descriptors of
the model found in the catalog. Saving the detail screen before the
catalog arrived, or with the automation's instance unavailable, found no
model and wrote null, wiping the stored picks. Before that commit the
forms never sent model_options, so this restores the keep-what-is-stored
behaviour for those cases.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 272 test files, prettier, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rrides

What:
- Environment editor: a row keeps its stored (redacted) secret only while
  it has the stored name and no typed value. Renaming the row leaves the
  secret behind and asks for a new value ("Enter a value for NEW: the
  stored secret of OLD does not move to a new name."); renaming it back,
  or clearing a typed replacement, keeps the stored secret again.
  EnvVarDraft.storedSecret becomes storedSecretName.
- Registry: encodeEnvironment refuses a redacted echo for a name with
  nothing stored ("There is no stored value of X to keep. Enter its
  value.") instead of storing an encrypted empty value.
- Home directory: the new-session warning also fires when the driver's
  home variable (CODEX_HOME, CLAUDE_CONFIG_DIR, COPILOT_HOME) is added,
  removed or changed among the environment variables, and the field shows
  that the variable set there is the one used.

Why: review of the P11 instance lane. Renaming a stored-secret row sent
{name: NEW, valueRedacted: true, value: ""}; the server found nothing
stored under NEW and stored "" for it, which then overrode the app's own
value of NEW for the instance while the dialog reported success. Clearing
a typed replacement did the same under the old name. The registry lets an
environment home variable win over config.homePath for the effective home
and continuation key, so changing it moved the home without the warning.

Tests: client drafts (rename, rename back, clear), registry (refused
redacted echo on create and update, stored value kept), a round trip from
the dialog's drafts through registry.update, and home-variable notices.
Mutation: without the registry guard the refusal test fails.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 272 test files, prettier, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- Claude and Codex now replay the conversation whenever they start a
  fresh native session, not only when no session is stored: the replay
  condition follows the resume decision (Claude's
  shouldResumeExistingSession, Codex's resumableCodexState), so a thread
  mode change (plan to chat, "Implement plan") also carries the prior
  transcript into the new session, as Copilot already does.
- The plan-mode preamble now comes before the replayed history, so the
  history's closing "New message:" introduces the user's own words with
  nothing in between (Claude's fresh plan session and Codex's fallback for
  app-servers without collaboration mode).

Why: review of the P11 instance lane. The replay keyed on whether a
session id was stored, but a mode change starts a fresh session even with
one stored, so "Implement the latest approved plan from this thread" went
to a session that had never seen the plan. And with the plan contract
after "New message:", the model could read it as part of the user's
message.

Tests: Claude and Codex mode-switch replay, and the preamble order
(Claude fresh plan session, Codex without collaboration mode; the old
Codex order test is updated). Mutation: with the previous condition and
order the three new tests fail.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 272 test files, prettier, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- Saving the automation detail screen while it still selects the stored
  instance and model keeps the stored model_options entries of options
  the model does not describe right now (not a picker option and not the
  reasoning option). Picker options still follow the form, and choosing
  another model or instance still starts over.
  resolveAutomationModelOptionsForSave takes the stored selection
  ({instanceId, modelId, modelOptions}) as an optional fourth argument.
- pruneAutomationOptionValues only drops values an option the model
  describes no longer offers. Values of options it does not describe stay
  in the form (they are not saved for that model), so the pickers show
  them again when the catalog entry recovers.

Why: review of the P11 instance lane. 9c25b78 kept the stored picks only
while the model was missing from the catalog. A model listed without its
option descriptors (OpenCode fallback models carry no agents or variants,
and a probe can lose the agent list) made the prune effect clear the form
values, and the next save of any field (the title, say) wrote
modelOptions: null.

Tests: the degraded catalog (no descriptors, only some), and another
model or another instance with the same model id; the prune test now
covers unoffered choices and undescribed options.

Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 272 test files, prettier, 59 tripwires).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After merging the P11 lanes the instance card also hosts the sign-in
panel and the install/update controls, which need a tRPC provider; the
instance-card render test stubs them (they have their own tests).

Gate G1: typecheck ok, all test files pass, format ok, tripwires ok,
next build ok.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant