Repository navigation
Conversation
- run-tests.mjs now runs every test file (it used to exit on the first
failure, hiding ~500 tests), prints a failure summary, supports
--jobs/--filter/--bail, includes scripts/ tests, and gives each file
its own throwaway SENTINEL_STATE_PATH/DB/MEDIA so runs never touch a
developer's real ~/.sentinel data
- make run_task and shell tests independent of the host `node` binary
- extend the bun:test type stub (test, spyOn, beforeAll, skipIf, ...)
- exclude .claude and worktrees from tsc
- CI: trigger on scripts/, desktop/, package.json, bun.lock and config
changes, add a build job, run tests in parallel
- add scripts/verify/{gate,tripwires}.mjs for revival phase gates
Gate G1: typecheck ok, 183/183 test files (1150 tests) pass, format ok,
next build ok.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Node 21.7.3 is end-of-life and below the minimum of the upgrade targets (ai 7, @ai-sdk/*, Electron 44, better-sqlite3 13, Copilot SDK 1.0 and Pi all require Node >= 22). Electron 44 embeds Node 24, which is also what the packaged Next server runs on, so pin the 24 LTS line everywhere: .nvmrc, .node-version, engines, @types/node, and the CI setup action (now reads .nvmrc). Gate G0: typecheck ok, 183/183 test files pass, format ok, tripwires ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bun 1.4 is the current release line. The full suite (183 files, 1150 tests, 260 mock.module call sites) passes unchanged on 1.4.2, so pin it for packageManager, engines and the CI setup action. Gate G0: typecheck ok, 183/183 test files pass on bun 1.4.2, format ok, tripwires ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tRPC 11.19, TanStack Query 5.104, drizzle-orm 0.45.3 / drizzle-kit 0.31.11 (staying on 0.x; 1.0 is still a release candidate), sqlite-vec 0.1.9 (first non-alpha), Tailwind 4.3.3, tiptap 3.31, lucide-react 1.52, hugeicons, @pierre/diffs 1.5, notion 5.27, mongodb 7.7, mysql2, pg, react-hook-form 7.89, resumable-stream, turndown, electron-updater 6.8.9, prettier 3.9.9, lefthook 2.1.17, postcss, wait-on. Gate G0: typecheck ok, 183/183 test files pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Prettier 3.9 collapses short union type annotations onto one line. Formatting only, no behaviour change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- googleapis 183: build OAuth clients with google.auth.OAuth2 instead of importing the undeclared google-auth-library, which now resolves to a different major than the one googleapis uses - @linear/sdk 97, @slack/web-api 8 (fetch transport; Sentinel uses no axios-specific options or ErrorCode checks), motion 14 (no breaking changes for React), shiki 4 (only removed misspelled aliases, unused) - commitlint 21, lint-staged 17, concurrently 10 (verified the commit-msg hook and `concurrently -k` still behave) Gate G0: typecheck ok, 183/183 test files pass, format ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- remove tailwind-variants, highlight.js and rehype-highlight (no imports anywhere) and the two Prettier plugins, which never loaded because the repo has no Prettier config listing them - officeparser 6.1.1 (latest 6.x). Holding below 8: officeparser 8 pulls pdfjs-dist 6, which hangs under bun (the test runtime) on PDF parsing even though it works on Node. Revisit when bun or pdfjs fixes that. - add PDF and ODT regression tests for the officeparser extraction path (the document parsers are loaded through runtime specifiers, so they were easy to mistake for unused dependencies) Gate G0: typecheck ok, 183/183 test files pass, format ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zod 4.6.5 with @t3-oss/env-nextjs 0.13.11 and @hookform/resolvers 5.9.1.
This also satisfies claude-agent-sdk's zod ^4 peer, and the lockfile now
resolves a single zod copy (copilot-sdk's nested zod 4.3.6 is deduped
onto the root), so instanceof ZodError checks see one class.
Code changes for the v4 breaks:
- z.record() takes an explicit key schema at all 19 one-argument sites
(integration tools, providers router, thread message JSON schema)
- appearance font size: invalid_type_error/required_error -> error fn,
same messages as before
- tRPC errorFormatter uses z.flattenError; tRPC 11 still calls zod's
parseAsync, so input errors keep a ZodError cause and the zodError
payload keeps its { formErrors, fieldErrors } shape
- threadCreate threadId: z.guid() keeps zod 3's 8-4-4-4-12 check (v4
uuid() is RFC-strict); the only producer is crypto.randomUUID()
- ZodTypeAny -> z.ZodType
ai 6.0.116 -> 6.0.301 (and @ai-sdk/react 3.0.304 so ai stays single):
provider-utils before 4.0.47 forced additionalProperties:false onto
every zod 4 object schema, so record inputs (Mongo queries, Airtable
fields, Notion properties: 23 tool fields) reached models as objects
that accept no keys. ai 6.0.265 is the first release with the fix; the
rest of the AI SDK stays where it is for P7.
Checked and unchanged: no .default().optional() or .partial() over
defaults, no superRefine ctx.path, .default() values already match the
output types, env.js still validates and SKIP_ENV_VALIDATION still
skips. Remaining zod 3 idioms that v4 still accepts (.url(), .email(),
ZodIssueCode.custom, superRefine) are left as they are.
Tool JSON schemas vs zod 3 (now snapshotted): .int() adds safe-integer
min/max, records add propertyNames, the email field adds a pattern,
computer_action's discriminated union is oneOf instead of anyOf, and
.optional().nullable() drops the old anyOf/not:{} wrapper. Sentinel
sends no strict tool schemas, and the Google converter keeps oneOf and
drops the other new keywords.
Tests: tool input JSON-schema snapshot over all 48 built-in and 159
integration tools (fails on objects that accept no properties), tRPC
zodError shape, env.js validation and skip, appearance messages,
thread/workspace schemas, engine-state round trip, settings defaults.
Tripwires: invalid_type_error/required_error, one-argument z.record,
ZodTypeAny.
Gate G0: install, typecheck ok, 188/188 test files pass, format ok,
tripwires clean (4 active).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
typescript ~6.0.3, not 7: TS 7 has no compiler API yet, which the Next TS plugin and Next's tsconfig handling still need. tsconfig.json for the TS 6 defaults and deprecations: - drop "baseUrl" (deprecated in 6, an error in 7); "paths" are already relative to the tsconfig, and Next, bun and esbuild all resolve them without it - "types": ["node"], since TS 6 no longer loads every @types package by default; everything else Sentinel uses comes in through imports - include .next/dev/types/**/*.ts, where Next 16 dev writes route types - keep the "next" language-service plugin noUncheckedSideEffectImports is on by default in TS 6 and flagged `import "@/styles/globals.css"`. next-env.d.ts declares *.css but is generated and gitignored, so src/types/css.d.ts declares it the same way Next 16 does (`declare module "*.css" {}`); side-effect imports of code modules stay checked. TS 6 also reported an always-nullish `?? null` in resolveOpenCodeTraitValueForThreadMode; the redundant fallback is gone, behaviour is unchanged. Tripwire: "baseUrl" in tsconfig*.json. Gate G0: install, typecheck ok, 188/188 test files pass, format ok, tripwires clean (5 active). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ai 6.0.301 bump brought @ai-sdk/provider-utils 4.0.57. On Node, its
default download fetch (generateVideo URL results, prompt asset and
transcription downloads) loads undici via
module.createRequire(<caller file>)("undici") to pin connections to
validated DNS results. Once webpack bundles and minifies that code into
.next/server/chunks, nft cannot see the require, so `output:
"standalone"` left undici out: a G1 build had no
.next/standalone/node_modules/undici, and resolving it from chunk 119 in
the packaged layout failed with MODULE_NOT_FOUND. That would reject every
default download in the desktop server (ELECTRON_RUN_AS_NODE makes
isNodeRuntime() true). bun skips this path, so the tests never saw it.
- scripts/desktop/untraced-server-packages.mjs lists such packages
(undici)
- next.config.js adds them to outputFileTracingIncludes for every route
- the bundle audit fails if the packaged server lacks one
- a test checks the trace include and that the top-level copy, which the
bundled chunks resolve, satisfies every installed dependent's range
(the three provider-utils 4.0.57 copies want ^6.28.0; root is 6.29.0)
A narrower ai bump would not avoid this: undici arrived in provider-utils
4.0.45 and the record additionalProperties fix in 4.0.47, and AI SDK 7's
provider-utils 5 depends on undici ^7.28.0.
Verified with a production build (scripts/next-build.mjs):
.next/standalone/node_modules/undici now exists and survives
prepare-production + prune-server. From a copy outside the repo, the
chunk-relative require resolves server/node_modules/undici and builds an
Agent.
Behaviour note: since provider-utils 4.0.47 these downloads also reject
hostnames that resolve to private, loopback or link-local addresses.
Literal localhost and private-IP URLs were already rejected in 4.0.19.
Sentinel sets no global dispatcher or proxy, so the dedicated Agent
bypasses nothing.
Gate G0: install, typecheck ok, 189/189 test files pass, format ok,
tripwires clean (5 active).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two zod 4 behaviour changes that the P3 commit let through:
- computer_action: zod 4 emits JSON Schema oneOf for discriminated
unions (they are always exclusive), where zod 3 sent anyOf. The Google
provider forwards oneOf unchanged, but Gemini's function-declaration
Schema (an OpenAPI 3.0 subset; @google/genai's Schema type) has anyOf
and no oneOf. Gemini requests that activate the computer tools would
likely be rejected. The tool input now uses z.union over the same
action options. Parsed data is identical for valid input, and the
discriminated union stays in automation-types for everything else. The
tool schema test now flags any oneOf/allOf, and the snapshot changes
only that keyword. Trade-off: a malformed action gets the union's
per-branch issues instead of the single discriminated-union issue.
- codexWriteConfig: zod 4 rejects an absent key for z.unknown() at
runtime ("expected nonoptional"); zod 3 accepted it as undefined.
value is now z.unknown().optional(), which restores both the runtime
check and zod 3's inferred `value?: unknown`. An audit of z.unknown,
z.any, z.undefined, z.void, z.custom and undefined unions found no
other production object key. The test mock now keeps input schemas so
the router test can check this.
The one-argument z.record tripwire comment now says it matches single
lines with at most two nested parenthesis levels and that typecheck is
the main guard: a multi-line call fails with TS2554, checkJs included.
Gate G0: install, typecheck ok, 189/189 test files pass, format ok,
tripwires clean (5 active).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Moves ai 6.0.301 -> 7.0.130 and every @ai-sdk package to its v7 major:
provider 4.0.24, react 4.0.133 (still not imported), mcp 2.0.69, openai
4.0.86, anthropic 4.0.74, google 4.0.90, google-vertex 5.0.104,
amazon-bedrock 5.0.108, azure 4.0.94, xai 5.0.17, groq 4.0.57, cohere
4.0.57, mistral 4.0.59, deepseek 3.0.61, moonshotai 3.0.65,
black-forest-labs 2.0.57, bytedance 2.0.59, fal 3.0.57, klingai 4.0.58,
replicate 3.0.57, plus @openrouter/ai-sdk-provider 3.1.0. Adds
@ai-sdk/provider-utils 5.0.56 (the version ai pins) for the type-only
ProviderOptions import: ai does not export it, and SharedV3ProviderOptions
is tied to provider spec V3. undici moves 6.29 -> 7.30 for
provider-utils 5; it is still loaded through createRequire, so the
standalone-server include from the zod/TS lane still applies.
Codemods (@ai-sdk/codemod 4.0.3 v7, each run through jscodeshift so the
result was visible): kept isStepCount, system -> instructions (title,
router, tool repair), onStepEnd/onEnd, transcribe and createGoogle.
Reverted the false positives: theme `system` keys renamed to
`instructions` (appearance page, workspace sidebar), Sentinel's own
metadata.usage.reasoningTokens and the Codex protocol usage rewritten to
outputTokenDetails, a doubled outputTokenDetails in runtime/reasoning.ts,
and `context: experimental_context` in prepareStep.
Silent v7 changes handled:
- Thread agent: runtimeContext replaces experimental_context in
prepareCall/prepareStep, and repairToolCall replaces
experimental_repairToolCall. prepareStep now returns the full
instructions from every branch: v7 carries returned instructions into
later steps, so a validation or task directive would otherwise stick
to every later step. Without a reroute the default branch reproduces
prepareCall's instructions exactly; after a reroute it describes the
rerouted tool set instead of the stale initial prompt.
- allowSystemInMessages: true on the agent. Context compaction feeds its
summary back as a synthetic role "system" message built server-side,
which v7 rejects by default.
- createAgentUIStream is replaced by createThreadAgentUIStream
(agent/ui-stream.ts): same flow, but the transcript is only validated
structurally, as in ai 6.0.116. The agent builds its tools in
prepareCall, so agent.tools is empty there. Since ai 6.0.301 the stock
helper throws "No tool schema found for tool part edit" when the
transcript holds an approval-responded tool part (approving a tool call
failed), and v7 also replaces every earlier static tool output with
"Tool output omitted because the tool is no longer available."
- createUIMessageStream gets an explicit onError; v7 redacts errors to
"An error occurred." by default. The agent stream already passed one.
- Reasoning tokens come from totalUsage.outputTokenDetails.reasoningTokens
(the top-level field is gone).
- generateObject -> generateText with Output.object (memory autosave,
commit messages).
- validateUIMessages drops Sentinel's approval.decision and
approval.response (6.x did too); messages/ui.ts copies them back after
validation. Dynamic tool titles now survive v7's schema (6.0.301
dropped them).
- Ollama uses the OpenAI Chat Completions model (.chat) through
createProviderLanguageModel; languageModel() on an OpenAI provider is
the Responses API.
- The OpenRouter providerOptions key is now "openrouter" (3.x ignores
"openai"). No OpenRouter catalog model has a reasoning config yet, so
nothing was sent under either key; P8 must use OpenRouter's
`reasoning: { effort }` payload when it adds one.
- MCP 2: the HTTP transport sets redirect: "follow" (2.x default
"error"), and both transports set protocolVersionDiscovery: false
(2.x probes server/discover before initialize), keeping the 1.x
handshake for user-configured servers. The OAuth provider reports its
client info as dynamically registered, so invalid_client still leads
to re-registration as in 1.x.
- createVertex -> createGoogleVertex.
Decisions and holds:
- OpenAI Responses function tools keep @ai-sdk/openai 4's explicit
strict: false. 3.x omitted strict, which let the Responses API apply
its own strict normalisation (optional tool parameters arrived as
empty strings, vercel/ai#11869).
- OpenAI reasoningSummary default ("detailed" once an effort is set):
nothing to change, Sentinel already sends reasoningSummary "detailed"
with every effort.
- Tool-level needsApproval stays (still honoured, now under test); the
move to toolApproval is P14.
- xAI 5 is Responses-only and createXai().languageModel() maps to it.
The catalog's grok-3, grok-4 and grok-4-fast ids were retired by xAI
on 2026-05-15 and redirect to grok-4.3; the catalog refresh is P8.
- streamText now runs tools after the model call finishes instead of
mid-stream; no test depended on the old timing.
- Not marked breaking: release-please has bump-minor-pre-major false,
so `!` would cut 1.0.0.
Tests: agent/ui-stream.test.ts drives a real ToolLoopAgent with
MockLanguageModelV4: an approved tool call runs and earlier tool outputs
reach the model, needsApproval still yields an approval request,
provider error text is not redacted, the compaction system message is
accepted, and usage/reasoning tokens map into thread metadata. The agent
test checks the per-step instruction reset and allowSystemInMessages,
messages/ui.test.ts the approval fields and titles, the factory test
Ollama chat and OpenAI/xAI Responses, the MCP test the redirect and
discovery options, and the orchestrator test both onError hooks. Mocks
use the v7 names (isStepCount, onStepEnd, onEnd, embeddingModel). The
tool JSON-schema snapshot is unchanged.
Tripwires (P7): stepCountIs, experimental_context,
experimental_repairToolCall, SharedV3ProviderOptions, generateObject,
experimental_transcribe, onStepFinish, createGoogleGenerativeAI.
Gate G0: frozen install ok, typecheck ok, 191/191 test files pass,
format ok, tripwires clean (13 active).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
extractTaskState read step.toolResults[].result, the AI SDK 4 shape.
Since AI SDK 5 a tool result carries the tool's return value in .output,
so the task map was always empty: the allTasksResolved stop condition
never fired and buildStepProgressAddon always returned "". It now reads
.output, so the agent stops once every tracked task is completed or
blocked, and steps with open tasks get the "Step Progress" directive.
Behaviour to watch: the loop now stops right after the step whose
manage_task call resolves the last task, as the stop condition intended,
so a run can end without a separate closing message after that call.
The tool router's evidence builder had the same dead branch ("result" in
toolResult ? toolResult.result : toolResult); it now reads .output
directly, which is what the old fallback always resolved to.
A test drives prepareStep and the stop condition with AI SDK 7 tool
result objects and checks that the AI SDK 4 shape no longer counts; it
fails on the previous commit. Tripwire P7-ai-v4-tool-result-shape
flags `toolResult(s)` followed by `.result` and `result.result`.
Gate G0: frozen install ok, typecheck ok, 191/191 test files pass,
format ok, tripwires clean (14 active).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The previous commit made the allTasksResolved stop condition work, but it only saw manage_task results from the current agent.stream call. Each approval resume or follow-up is a new call, so a continuation run that updated one task stopped as soon as that task was terminal, even with other plan tasks still pending. - The orchestrator passes the plan's task statuses from run start as the new planTasks call option; prepareCall keeps them and this run's manage_task results are applied on top. - Task state only counts once this run has changed a task, so leftover or already finished plan tasks never stop or steer an unrelated run. - The stop fires one step after the last open task resolves. That step lets the model report back (a text answer ends the loop by itself), and a tool call there ends the run. Opening a new task in it keeps the run going. - The Step Progress addon drops the step number, so the instructions only change with the task counts and no longer break prompt-prefix caching on every step. Tests: loop.test.ts drives the real ToolLoopAgent from createThreadAgent with MockLanguageModelV4 (routing, tools and instructions stubbed). It checks that a step directive reaches exactly one model call (AI SDK 7 carry-forward), that the run continues while earlier plan tasks are open, that resolved plan tasks do not stop a run, and the reporting step. Two of these fail on the previous commit. Unit tests cover the stop condition and progress text, and a thread-chat test checks planTasks is passed. Gate G0: frozen install ok, typecheck ok, 192/192 test files pass, format ok, tripwires clean (14 active). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ictness xAI: @ai-sdk/xai 5 only has the Responses API, whose `store` option defaults to true. xAI then stores prompts and responses for later retrieval, and Zero Data Retention teams get an API error. Sentinel's earlier Chat Completions path stored nothing and Sentinel never reads stored responses back. createProviderLanguageModel now wraps xAI models in defaultSettingsMiddleware with providerOptions.xai.store = false, so every caller gets it (the agent, routing, titles, memory, commit messages). With store off the provider also requests reasoning.encrypted_content, so reasoning still round-trips. A per-call xai option still wins. OpenAI/Azure strict tools (decision, no runtime change): AI SDK 7 sends Responses function tools with `strict: false` unless a tool opts in. AI SDK 6 omitted the field. The Responses API then tried strict mode when a schema looked compatible and normalised it, which made optional parameters arrive as "" (vercel/ai#11869, fixed by vercel/ai#15889). That omitted-field behaviour cannot be reproduced from the client, and sending strict: true has no non-strict fallback (an unsupported keyword fails the whole request) and would also switch on strict tools for Anthropic and others. So Sentinel keeps strict: false. A factory test pins the value sent for OpenAI and Azure so a provider bump cannot change it silently. Tests stub fetch and check the request bodies: xAI sends store false and the encrypted reasoning include (fails without the wrapper), and OpenAI and Azure send strict: false. The agent and thread-chat tests that mock 'ai' by name now also export defaultSettingsMiddleware and wrapLanguageModel. Gate G0: frozen install ok, typecheck ok, 192/192 test files pass, format ok, tripwires clean (14 active). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
AI SDK 7 runs the start/status polling for xAI, Kling AI and ByteDance
video models itself (experimental_generateVideo's `poll`). The providers
now ignore pollTimeoutMs in providerOptions, and ByteDance also returns a
"deprecated setting" warning that generate_video copied into the tool
output shown to the user and the model.
buildVideoProviderOptions becomes buildVideoPollOptions and passes
poll: { timeoutMs: 600_000 } for those three providers. That is the SDK
default today, so timing does not change, but the setting is live again
and the warning is gone. These video models have no doGenerate, so passing
poll does not change which flow the SDK picks. Other providers get no poll
option, as before.
The video test now checks the poll option and that no providerOptions are
sent.
Gate G0: frozen install ok, typecheck ok, 192/192 test files pass, format
ok, tripwires clean (14 active).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- next 16.4.0, react/react-dom 19.3.0, @types/react(-dom) 19.3.0 - next.config.js: drop the eslint key (removed in Next 16) and set compress: false; this is a loopback desktop app and Next's compression buffers streamed responses (the chat resume SSE route already forced Content-Encoding: none; the planned tRPC SSE subscription needs the same) - next.config.js: agentRules: false; Next 16 `next dev` otherwise writes a managed AGENTS.md into the repo root whenever it detects a coding agent - dev/dev:desktop: drop --turbo, Turbopack is the Next 16 dev default - scripts/next-build.mjs: production builds pass --webpack (decision D6: Turbopack standalone output with serverExternalPackages can emit aliased externals/symlinks that break desktop packaging); SENTINEL_NEXT_BUNDLER=turbopack opts into Turbopack - tripwires: --turbo in package.json, eslint key in next.config Audit: request APIs are already async, the only parallel slot (@settings) has default.tsx, no middleware, images, revalidateTag, runtime config, AMP or process.argv checks; no global smooth scroll-behavior, so no data-scroll-behavior. React 19.3 StrictMode double-invokes effects during hydration in dev; the thread session store refcounts subscribers and aborts the superseded SSE fetch, so only one stream stays open. Hold: next dev/build rewrite tsconfig.json (mandatory jsx: react-jsx, plus .next/dev/types in include); tsconfig is owned by the TS 6 lane, so those edits are left for it. Gate G0: frozen install ok, typecheck ok, 183/183 test files pass, format ok, tripwires clean (3 active). next build --webpack and an isolated next dev (port 3300) verified at the lane tip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- @heroui/react and @heroui/styles 3.2.6; 3.2.1+ no longer bundles react-aria, so its peers are now direct dependencies: react-aria 3.52.1, react-aria-components 1.21.1, @react-aria/ssr 3.10.1, @react-aria/utils 3.34.1, @internationalized/date 3.12.4 (single copy of each in bun.lock) - 3.2.0 makes Switch/Checkbox/Radio `*.Content` the clickable <label> and requires the control inside it. All 33 toggles are migrated: control-only switches wrap the control in `*.Content`; labelled ones nest the control in `*.Content` with font-normal (Content now sets font-medium, which descriptions would otherwise inherit); option cards in the user-input and plan renderers move their card styling onto `*.Content` so the whole card stays clickable, keeping label and description inside it. The plan renderer keeps the control offsets HeroUI used to apply. - globals.css: drop the `--default-hover` overrides. 3.0.1 ignored them (hover was derived from `--default`); 3.0.5+ reads them, which would have made light-mode hovers lighter than the default background - globals.css: zero ScrollShadow's new 10px scrollbar gutter in its fade mask; Sentinel hides scrollbars, so the gutter showed as an unfaded strip - workspace sidebar: overlay triggers are inline-block since 3.0.2; keep the linked-folder tooltip trigger block-level so the row stays full width - test: controlled switch/checkbox fields render control, input and label inside the clickable content (fails on the old composition) Audited, no change needed: Text->Typography (unused), tooltip delays (every tooltip sets `delay`; close delay is React Aria's 500ms as before), Tabs.ListContainer, Select/ComboBox/ListBox, Modal/Drawer (overlay z-index now 100000), Toast (Sentinel uses sileo). Accepted upstream restyles: tooltip padding p-2, radius tokens, accessible soft-foreground palette. Gate G0: frozen install ok, typecheck ok, 184/184 test files pass, format ok, tripwires clean (3 active). G1: next build --webpack ok; standalone keeps better-sqlite3 (with .node) and sqlite-vec. Isolated next dev on :3300: settings switches render and toggle, label text toggles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Next 16 treats `jsx: react-jsx` as mandatory and rewrites tsconfig.json on every build or dev run when it is set to `preserve`, which left the tree dirty. Set it explicitly (`.next/dev/types` is already included). Gate G0 (rebased on zod 4 + TS 6 + AI SDK 7): typecheck ok, all test files pass, format ok, tripwires ok; next build (webpack) ok and leaves tsconfig.json untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…qlite3 13 electron 44.6.0, electron-builder 26.17.0 (v26 tag), better-sqlite3 13.0.3 and @types/better-sqlite3 9.6.0 move together because the server's native SQLite module has to load under Electron's embedded Node (24.21). Electron binary (42+ no longer downloads it in postinstall): - scripts/desktop/electron-binary.mjs runs the package's install-electron script when node_modules/electron/dist is missing or from another version; exposed as `bun run electron:install` - dev-launch.mjs and the packaging preflight call it, so build.electronDist and the dev launcher always find a binary - .github/actions/setup-desktop-build runs it, covering desktop-verify and publish-release Main process (desktop/main/index.mjs): - await clipboard.writeText (returns a Promise in 44) - console-message listener reads the details object (level is now a string; the positional form is deprecated and logs a warning) - Electron 43 opens dialogs in Downloads when no defaultPath is passed; remember the folder of the last pick so the workspace and file pickers reopen there, as the OS did before - checked webview + did-attach-webview, guest debugger attach, setWindowOpenHandler, webRequest, permission handlers, media access, nativeImage, dock, dialogs and updater net.request against the 41-44 breaking changes; no other changes needed. The app uses no login items, Notification or renderer clipboard. Packaging: - drop the ia32 and armv7l targets (no Electron 44 binaries) - better-sqlite3 13 is N-API with bundled prebuilds: the packaged copy keeps only prebuilds/<platform>-<arch>.node and drops binding.gyp, build, deps and src instead of rebuilding against Electron headers (cross-arch builds now use the bundled prebuild too). The host smoke test (select 1 under ELECTRON_RUN_AS_NODE) stays - rebuild-node-native.mjs only smoke-tests better-sqlite3 and keeps the node-pty repair path; node-gyp is now a declared devDependency for it (Linux has no node-pty prebuilds) - audit-bundle asserts the packaged better-sqlite3 has the target prebuild and no other prebuilds or build inputs Docs: README and docs/product install list macOS 13+, 64-bit Windows 10+ and Linux x64/arm64, plus `bun run electron:install`. Tripwires (P6): ia32/armv7l in scripts/desktop, .github and package.json; prebuild-install in scripts; ELECTRON_SKIP_BINARY_DOWNLOAD; un-awaited clipboard.writeText and positional console-message in desktop/. Holds: node-gyp stays on 12.x (13 needs Node >=24.15; the repo allows 24.14). electron-updater stays 6.8.9 (6.8.10 has the same builder-util-runtime 9.7.0). Gate G0: frozen install ok, typecheck ok, 185/185 test files pass, format ok, tripwires clean (6 active). Unsigned build:desktop:mac and its bundle audit pass; the packaged binary (ELECTRON_RUN_AS_NODE) loads better-sqlite3 13 and runs select 1. sqlite-vec does not load from the packaged server; fixed in the next commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sqlite-vec 0.1.9 locates vec0.<ext> with a dynamic require.resolve of sqlite-vec-<os>-<arch>, which Next's output tracing does not follow (0.1.7-alpha.2 joined a path from __dirname, which it did). Since the sqlite-vec bump the packaged server shipped sqlite-vec without its platform package, so initVectorDb failed quietly and knowledge vector search was off in packaged builds (the v0.0.66 release still has sqlite-vec-darwin-arm64). - rebuild-standalone-native copies the target's sqlite-vec platform package next to sqlite-vec (an error on host builds when it is not installed, a warning on cross-target builds) and its Electron-as-Node smoke test now also loads sqlite-vec and reads vec_version() - audit-bundle fails when the packaged server has sqlite-vec but not node_modules/sqlite-vec-<os>-<arch>/vec0.<ext> Gate G0: frozen install ok, typecheck ok, 185/185 test files pass, format ok, tripwires clean (6 active). G2 packaging: unsigned `bun run build:desktop:mac` passes (preflight, host smoke "better-sqlite3 select 1 = 1" and "sqlite-vec v0.1.9", electron-builder 26.17.0, bundle audit). Packaged import smoke with ELECTRON_RUN_AS_NODE from Resources/server, using both Sentinel and "Sentinel Helper (Plugin)": Electron 44.6.0 / Node 24.21.0, better-sqlite3 13.0.3 select 1, sqlite-vec v0.1.9 loaded, node-pty spawned a pty. The packaged server.js on port 3399 with an isolated HOME and state paths served /api/health and DB-backed tRPC queries (migrations ran, vectors.db created). Sizes vs the v0.0.66 release: dmg 162.7 MB vs 146.8 MB (+10.9%), zip 158.9 MB vs 142.4 MB (+11.6%), Resources/server/node_modules 94.8 MiB vs 86.6 MiB (+9.4%). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Electron 44 needs macOS 13 and the packaged Info.plist now says so, but electron-builder 26.17 writes no minimumSystemVersion into latest-mac.yml. Installed builds (electron-updater 6.8.3 in v0.0.66) auto-download any newer feed entry, so a macOS 12 user would be updated to an app that no longer opens. - scripts/desktop/update-feed.mjs reads LSMinimumSystemVersion from the packaged .app, maps it to the Darwin kernel version electron-updater compares with os.release() (macOS 13 = 22.0.0) and writes it into every dist/*-mac.yml. Unmapped or minor-version floors fail loudly instead of guessing. - package.mjs stamps the feed right after electron-builder (builds run with --publish never; publish-release uploads dist/latest-mac.yml afterwards). - audit-bundle fails a mac build whose feed is missing the field or disagrees with the bundle's floor. - Tests cover the mapping, idempotent stamping, the parsed feed under electron-updater's own js-yaml/semver (Darwin 21.6 blocked, 22.1 allowed) and the package/audit wiring. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install, typecheck, 186/186 test files, format check, tripwires clean). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
better-sqlite3 13 dropped its prebuild-install script and sets
`gypfile: false`, but bun still runs an implicit `node-gyp rebuild` for
any trusted package with a binding.gyp. The rebuild compiles nothing (the
binding skips when a prebuild exists), yet node-gyp configure needs
Python, and Visual Studio on Windows. With no Python on PATH, a frozen
`bun install` of the previous commit fails with "install script from
better-sqlite3 exited with 1". 12.x only needed prebuild-install there.
- package.json gets an explicit trustedDependencies list: electron-winstaller,
esbuild, lefthook, node-pty and sharp, which are the packages whose
scripts run today under bun's default allowlist. better-sqlite3 is left
out, and tesseract.js stays blocked as before. An explicit list replaces
bun's default one, so package-config.test.ts now scans node_modules and
fails when any installed package with install scripts (or an implicit
gyp build) is neither trusted nor deliberately untrusted.
- Docs: README, install.md and environment-and-build.md no longer promise
a better-sqlite3 source build. They say it loads its bundled N-API
prebuild, document the Linux glibc 2.34 / libstdc++ (GCC 11) floor of
those prebuilds (GLIBC_2.34 and GLIBCXX_3.4.29 in prebuilds/linux-*.node),
say build tools are only needed for node-pty (always on Linux), note
that macOS 12 installs stop getting updates, and add the
`bun run electron:install` note to environment-and-build.md.
- rebuild-node-native's better-sqlite3 error names the glibc floor on
Linux. Its comment no longer mentions an install-time source build.
Verified: a full `bun install --frozen-lockfile` with PATH limited to
node, bun, git and /bin (no python3) exits 0, and better-sqlite3 select 1
plus require('node-pty') work. The same install of the parent commit
fails in node-gyp configure.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install,
typecheck, 186/186 test files, format check, tripwires clean).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Electron 43 opens dialogs in Downloads when defaultPath is omitted. The
earlier workaround kept one in-memory folder shared by the workspace and
file pickers. So the first dialog after every launch still opened in
Downloads, and attaching files moved the workspace picker to the
attachment folder. Before Electron 43 the OS reopened the last folder,
even across restarts.
- dialog-paths.mjs now keeps the last folder per dialog kind ("directory"
and "files") in userData/dialog-paths.json, using the existing
scripts/desktop/state.mjs helpers, which are already packaged. Until the
first pick, or once the remembered folder no longer exists, it falls back
to the home folder rather than Downloads. Unreadable state is ignored,
and a failed save only logs a warning.
- The store is created in registerIpc (after app ready) so userData is
final. PICK_DIRECTORY and PICK_FILES each use their own kind.
- Tests cover the fallback, separate kinds, restoring after a restart,
cancelled dialogs, deleted folders and corrupt state.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install,
typecheck, 186/186 test files, format check, tripwires clean).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pwires The P6 tripwires only matched the exact old lines: an un-awaited clipboard.writeText( and a `(_event, level, message` parameter list. desktop/main is plain .mjs that typecheck does not cover, so other Electron 44 breaks would only show up at runtime. - P6-unawaited-clipboard (desktop/main): has, read, readText, write and writeText now return Promises and must be awaited or returned. This replaces P6-unawaited-clipboard-write. - P6-removed-clipboard-api (desktop): availableFormats and the HTML, image, RTF, bookmark, buffer and find-text read/write methods. - P6-renderer-clipboard (desktop/preload): Electron's clipboard imported or required in the preload, which Electron 44 removed from renderers. - P6-positional-console-message now matches any console-message listener with a second positional parameter, including across line breaks and with other parameter names, plus named (e, level, message) style handlers. - tripwires.mjs gains an opt-in `multiline` flag: the pattern runs on the whole file and the hit reports the line the match starts on. The per-file scan moves into an exported findTripwireHitsInContent so the patterns can be unit-tested. Line-based behaviour is unchanged. - tripwires.test.ts checks hits and non-hits for each P6 pattern. Run against the pre-P6 desktop/main/index.mjs (b88ef14), the widened set reports both un-awaited writeText calls and the positional listener; the current tree has 0 hits. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install, typecheck, 187/187 test files, format check, tripwires clean, 8 active). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The packaged server resolves undici from the top-level node_modules (provider-utils loads it through an untraced createRequire). After the Electron 44 lane added node-gyp 12 (undici ^6) bun hoisted undici 6, which no longer satisfies @ai-sdk/provider-utils 5 (^7.29), so default downloads in packaged builds would have loaded the wrong major. - declare undici ^7.30.0 directly so the hoisted copy is the one the server needs; build tooling keeps its own nested copy - scope the trace test to the packages that load undici at runtime (UNTRACED_SERVER_PACKAGE_LOADERS) instead of every installed dependent Gate G0: typecheck ok, all test files pass, format ok, tripwires ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sentinel engine's MODEL_CATALOG stopped at claude-opus-4-6 and still listed models that providers have shut down. Every provider list now starts with its current flagship, verified against the installed @ai-sdk/* model-id unions, the providers' model and deprecation pages, the AI Gateway model list and OpenRouter's /api/v1/models: - OpenAI: GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol/Luna, GPT-5.6 Sol/Terra/Luna ahead of the GPT-5.x line; 922K/272K prompt windows. - Anthropic: Opus 5.5, Sonnet 5.5, Fable 5.1, Haiku 4.5, then Opus 5, Sonnet 5, Fable 5, Opus 4.8/4.7/4.6, Sonnet 4.6 and legacy 4.x; 1M. - Google and Vertex: Gemini 3.8 Flash, 3.1 Pro preview, 3.7/3.6/3.5 Flash, 3.5 Flash-Lite, 3 Flash preview, then the 2.5 models. - xAI Grok 4.7/4.6/4.5/4.3/4.20, DeepSeek V4 Pro and Flash, Kimi K3, K2.7 Code and K2.6, Mistral Large/Medium/Small -latest, Codestral, Ministral 14B, Cohere Command A Plus, Groq GPT-OSS, Bedrock Claude 5.5 inference profiles, Gemma 4 / Qwen 3.5 / GPT-OSS on Ollama, and current Gateway and OpenRouter ids (dotted versions, spacexai/ and x-ai/). Retired, renamed or never-valid ids move to RETIRED_MODEL_REPLACEMENTS with the provider's recommended successor (gpt-5-codex, o1, o3-mini, o4-mini, gpt-4.1-nano, claude-opus-4-1, Claude 3.x, Gemini 1.5/2.0 and 3 Pro preview, grok-3/4, deepseek-chat/-reasoner, moonshot-v1, Kimi K2/K2.5, Magistral 2507, Pixtral Large, ...). Threads, automations, stored defaults (normalizeSelectedModelId) and the composer resolve a retired id to its successor, unless the user still has that id enabled (for example as a custom model). Unknown ids keep their previous path. Reasoning options follow each provider's current values without switching to AI SDK 7's top-level `reasoning`: Claude 4.6+ send effort with adaptive, summarized thinking (Opus 5.5 defaults to medium; xhigh maps to max on Opus/Sonnet 4.6), Gemini 3.x thinking levels with thought summaries and per-model defaults, GPT-6/5.6 efforts, Codex models low to xhigh, o3 low to high, GPT-5 sends `minimal` instead of the unsupported `none`, Grok 4.5+ up to xhigh, DeepSeek V4 thinking plus reasoningEffort, Kimi K3 low/high/max (max by default), Mistral high/none. Helper models (titles, tool routing) are catalog entries again: deepseek gets deepseek-flash instead of an OpenAI fallback, Gateway/OpenRouter use the provider-prefixed google/gemini-3.5-flash-lite, Bedrock uses the us.anthropic.claude-haiku-4-5-20251001-v1:0 profile, Ollama llama3.2, Moonshot kimi-k2.6, OpenAI gpt-6-luna. resolveHelperModel asks them for the least reasoning they accept (the router previously hard-coded `minimal`). The Codex fallback list follows the catalog to GPT-6 Astra, 6.1 Sol and Luna (same ids as the P9 Codex lane, which replaces it). Native attachments now derive from the vision capability of built-in anthropic/google/vertex/openai models instead of a duplicate allow-list. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 199/199 test files, format, 25 tripwires) on the lane tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The media catalogs pointed at models the installed providers no longer serve: @ai-sdk/google 4 and @ai-sdk/google-vertex 5 throw for any image model id that does not start with `gemini-` (so both Imagen entries were dead), OpenAI shut DALL-E 2/3 down on 2026-05-12 and shuts gpt-image-1 down on 2026-10-23, and Google shut Veo 2/3.0 down on 2026-06-30. Images: GPT Image 2.5 Sunburst/Flare and GPT Image 2 (with reference image edits, which @ai-sdk/openai supports for them) ahead of the expiring GPT Image 1.5/1 Mini; Nano Banana 2.1 and Gemini 3 Pro Image for Google, Nano Banana 2.1 and Gemini 2.5 Flash Image for Vertex (no mask editing: Gemini image models reject masks); Grok Imagine Image 2.0; FLUX 3 (reference images, no seed). Video: Veo 3.1 GA ids first on Vertex, the Veo 3.1 previews on AI Studio, Grok Imagine Video 1.5. A stored image or video model that a provider retired now resolves to its replacement (RETIRED_IMAGE_MODEL_REPLACEMENTS / RETIRED_VIDEO_MODEL_REPLACEMENTS) instead of leaving the provider with no valid model; custom ids keep their previous handling. Transcription: OpenAI deprecated whisper-1 and the GPT-4o transcribe models (shutdown 2027-02-26); gpt-transcribe, on the same /audio/transcriptions endpoint, is the new default for users who never picked a model. Explicit choices are kept and still listed. Embeddings: add gemini-embedding-2 (Gemini API and Vertex, 3072 dims) and Cohere embed-v4.0 (1536 dims) profiles. Existing profile ids and the text-embedding-3-small default are unchanged, because stored memory is tied to its profile's vector size. Tripwire P8-retired-image-catalog-ids keeps Imagen and DALL-E out of the image catalog. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 199/199 test files, format, 25 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What and why: - Mistral: mistral-large-latest is Mistral Large 3 (mistral-large-2512, GA, 256K, no adjustable reasoning) per docs.mistral.ai/models, and @ai-sdk/mistral 4.0.59 only sends reasoning_effort for ids on its own allow-list (mistral-large-4 but not mistral-large-latest), so the effort picker on the default Mistral model did nothing. mistral-large-latest is now listed as Large 3 without a reasoning config (262,144 context), and Mistral Large 4 (public preview since 2026-10-06) is added as mistral-large-4 with the none/high config (524,288 context, the cap the AI Gateway and OpenRouter serve; Mistral lists 1M). The GA model stays first, so the default for new Mistral threads is unchanged from before the refresh. - Claude preserved thinking: Opus 5.5, Sonnet 5.5 and Fable 5.1 bind thinking blocks to the conversation, and accounts created on or after 2026-08-31 get a 400 when a replayed block's earlier history changed. Sentinel's compaction keeps recent turns behind a new summary and the tool router changes the active tool set between steps, so these models now send thinking.blockBinding.prefixMismatchBehavior "drop_block" (@ai-sdk/anthropic adds the thinking-binding-controls-2026-08-01 beta), with or without a selected effort. Other Claude models are unchanged. - xAI: grok-4.5 offers low/medium/high (xAI's reasoning guide says xhigh exists on grok-4.6+ and is treated as high on 4.5); grok-4.3 drops xhigh (its model page lists none/low/medium/high). - Retired ids: resolveStoredCompositeModelId only switches to the successor when the user has it available, matching normalizeSelectedModelId; otherwise the stored id passes through to the caller's existing fallback, as before the refresh. - Azure: gpt-5 is first again. Azure ids are deployment names, and gpt-5 was the default before the refresh. - Bedrock: Claude Sonnet 4.5 has no in-region endpoint, so the entry is us.anthropic.claude-sonnet-4-5-20250929-v1:0 and the bare id maps to it. Tests: - New reasoning-requests.test.ts sends every effort the picker offers for every catalog model of anthropic, cohere, deepseek, google, groq, mistral, moonshotai, openai and xai through the real provider package and asserts the value reaches the request body. Before the Mistral fix it failed on mistral-large-latest; every other provider passed. It also checks the drop_block request body and beta header for a compacted transcript that replays a signed thinking block. - models.test.ts and context/model.test.ts updated for the above. - Tripwire P8-bedrock-in-region-claude-ids forbids bare anthropic.claude-* catalog ids (zero hits). Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install, typecheck, all 200 test files, format check, tripwires clean, 26 active). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bump @anthropic-ai/claude-agent-sdk ^0.2.84 -> ^0.3.292 (bundles Claude Code 2.1.292) and add its
new peers @anthropic-ai/sdk ^0.131.0 (peer >=0.93) and @modelcontextprotocol/sdk ^1.32.1 (peer
^1.29); zod 4.6.5 stays the only zod. The old SDK's @img/sharp 0.34 optional deps drop out, so
sharp 0.35.5's own platform packages dedupe to the top level.
Compatibility with the new SDK:
- options.env replaces process.env since 0.2.113: buildClaudeSdkEnv always layers the runtime env
over process.env (run, probe, and any caller of buildClaudeSdkBaseOptions).
- permissionMode is always explicit (base default "default"); omitted it follows the settings
defaultMode, possibly "auto", since 0.3.286. chat -> default (+sandbox), full ->
bypassPermissions + allowDangerouslySkipPermissions, plan -> plan, as before.
- Native binaries are spawned by the SDK itself (stderr tail in exit errors, graceful shutdown).
Only Node-script CLIs (.js/.mjs/.cjs or a node shebang, e.g. npm shims) get Sentinel's spawner,
which runs them under process.execPath with ELECTRON_RUN_AS_NODE=1 (GUI launches often have
no node on PATH). The same launcher verifies `claude --version`. Windows npm .cmd/.bat/.ps1
shims resolve to bin/claude.exe or cli.js (ported from t3code ClaudeExecutable.ts, MIT).
- The chat-mode sandbox sets failIfUnavailable:false (0.2.91 default flip) so runs keep going
unsandboxed when the sandbox cannot start; Bash auto-allow then turns off and Bash is prompted
(checked in the CLI: auto-allow requires isSandboxingEnabled()).
- Status probe: a prompt stream that never yields (no turn can reach the API) and no hooks, MCP
servers or IDE auto-connect (t3code ClaudeProvider probe options). initializationResult() and
close() still exist; models/account are read as before.
- AskUserQuestion is answered the documented way: canUseTool resolves
{behavior:"allow", updatedInput:{...input, answers}} with answers keyed by question text
(multi-select comma-joined), the card's extra context as annotations notes, and unmatched free
text as `response`. The synthetic tool_use_result user message is gone; an answer that arrives
before canUseTool is held and applied when Claude Code asks.
- Task tools replaced TodoWrite (0.3.142) and are off by default on new models (0.3.233/0.3.268);
native builds swap Grep/Glob for Bash find/grep (0.3.162). allowedTools names all six on top
of the claude_code preset (an explicit tools list would freeze the tool surface). The mirror
accumulates Task* results by task id (structured tool_use_result, seeded from the transcript
when the session resumes) and stores the list on output.tasks; new claude_task* renderers show
it. claude_todowrite stays for persisted messages.
- Effort is now sent: clamped to the model's supportedEffortLevels from the last probe snapshot,
passed through for unknown models, none/minimal -> low. xhigh is offered; max waits for the
P10 ReasoningEffort widening. Default efforts follow t3code's model manifest instead of the
lowest level (which would now be sent). Thinking is adaptive with summarized display.
- rate_limit_event is recorded per window (and on the run control) for P11's usage limits, and
logged at debug. A success result with is_error now fails the run with its error text instead
of showing it as the answer. local_command_output and session_state_changed are unchanged.
- Context windows and fallback models follow the current lineup (Opus 5.5, Sonnet 5.5, Fable 5.1
as default, Haiku 4.5; aliases via ModelInfo.resolvedModel; [1m] -> 1M; Opus 4.7/4.8 1M).
- Commit messages run the resolved Claude binary with its managed env, not `claude` from PATH.
- Image attachments in types the API rejects are sent as a text note.
Tripwires (zero hits; verified against the pre-change files): the synthetic AskUserQuestion
answer, TodoWrite-only Claude renderers, `env ?? process.env` in the Claude SDK options, and
spawning `claude` from PATH.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (frozen install, typecheck, 202 test files,
prettier, 27 tripwires clean). One run hit a load-sensitive Sentinel-engine test
(thread-chat/index.test.ts, unrelated) that passes alone 6/6 and on the rerun.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: an auth panel in every engine instance card (components/settings/engines/engine-auth-panel.tsx). It lists the methods the instance offers (api.engines.auth.methods), starts and cancels flows, and shows the step the server waits on: - a sign-in page to open (desktop opens it in the system browser) with a link to copy; - a device code to copy and the page to enter it on; - on desktop, an embedded terminal running the CLI's sign-in (engine-auth-terminal.tsx, over terminal.createCommand and the flow's launch ticket); its exit is reported back to verify the result; - in a browser (or when the terminal cannot start), the command to copy and a Done button; - a credentials form (password inputs) whose values go to the server once. Sign-out asks for confirmation first. The panel polls auth.status every second while a flow runs, and EngineEventsBridge refetches it on auth events. When a flow finishes, snapshots, the composer catalog, models and methods are refetched (the server already re-probed the instance). engine-instance-card.tsx gains one line (the panel) and use-engine-snapshots.tsx a small handler for auth events. Why: critique G1, driver-contract §2.5; the settings page only said "Login needed". Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (275 test files); bun run build passed (the new /api/engines/auth/terminal route is in the route table, no client module pulls server code). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: loggerLink skips engines.auth.* (credentials, device codes, launch tickets), engines.codex.login (API keys) and engines.instances.create / update (secret instance variables), through isSensitiveTrpcOperation in src/trpc/sensitive-operations.ts. Credential answers are also capped at 16 fields by the schema. Why: in development loggerLink logs every operation with its input and result, and the desktop app forwards the renderer console to its own output, so a sign-in would have printed the key the user entered. In production it logged failed operations, inputs included. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (276 test files). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…R helper What: platform/auth/browser-helper.ts writes a /bin/sh script (a .cmd on Windows) into an instance's state directory. With BROWSER pointing at it, an agent that opens its own sign-in page prints the URL behind a marker on stderr instead; parseAuthBrowserMarker reads it back (https, or http on loopback, only) for the driver to show as a browser interaction. The script never evaluates the URL. The composer's "needs authentication" message now points to Settings → Engines. Why: critique G13. Antigravity and ACP agents (P13) need this to sign in from Sentinel, and the helper must not depend on a Node interpreter in packaged builds (t3code uses a Node helper). Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t a runtime What: the flow store catches a failure of its background runner and ends the flow as failed instead of leaving it waiting until the TTL. The auth panel offers "Sign out" only when the instance's runtime was found, so an engine that is not installed shows its install hint rather than a sign-out that cannot run. Why: review of the P11 auth lane. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: claude-sdk exports forgetClaudeEngineStatus(instance), which drops the instance's in-memory status caches (bumping the generation so a background refresh cannot write them back) and deletes its on-disk last-known-good snapshot. The Claude auth controller calls it after `claude auth login`, after storing an API key, and after `claude auth logout` (whether or not the command or the key removal did the work). Why: the status probe answers an empty model list or a failure, which is how a signed-out Claude Code looks, from that snapshot for up to 7 days, forced refresh or not. The flow store's verifying refresh therefore still saw "authenticated" with the old account after a sign-out, and the panel reported "Signed out, but Claude Code still finds credentials", kept Sign out visible and the composer kept offering Claude. A sign-in that switched accounts could likewise show the previous account when the probe timed out. Tests: a probe without models after forgetting reports auth_unavailable with no account and the snapshot file is gone; the controller forgets after sign-in and sign-out. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files, install, typecheck, format, tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t sign-out What: - EngineAuthController gains logoutNotice(instance), returned by api.engines.auth.methods as `logoutNotice` and shown in the panel's sign-out confirmation. sharedLoginNotice() (drivers/auth/shared.ts) covers engines whose login lives in the CLI's usual configuration unless the instance has a home of its own: Claude (CLAUDE_CONFIG_DIR), Codex (CODEX_HOME), Copilot (COPILOT_HOME). Cursor always warns (one login per OS user, shared by every Cursor instance); OpenCode warns unless the instance sets XDG_DATA_HOME. Copilot also says that it restarts on the instance, which stops chats running on it. - Copilot sign-out removes only COPILOT_GITHUB_TOKEN, the variable this panel stores. GH_TOKEN / GITHUB_TOKEN set on the instance stay, and the outcome (or the failure) names them with how to sign out completely. Why: signing out a default instance ran `claude auth logout`, `agent logout` or account.logout against the user's global login with only "Sign out of X?" as warning, and Copilot sign-out silently deleted token variables a user may have set for MCP servers or gh in agent shells. The Copilot sign-in/out restart is kept: like any change to an instance's configuration (which the snapshot service already retires immediately), the runtime has to restart to read the new login; the confirmation now says so. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (277 test files, install, typecheck, format, tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- The embedded terminal step remembers the command's exit and, once it
exited, shows "Check sign-in" next to Cancel, which sends the exit code
again. FlowStep is exported as EngineAuthFlowStep.
- Electron main ends command terminals (kind "command", the sign-in PTYs)
when the main frame navigates to another document or its renderer
process goes away. Shell terminals are left as they were.
- Render tests (renderToStaticMarkup) of the flow steps: embedded terminal
with Cancel only, command to copy with Done, credentials in empty
password fields with autocomplete off, device code, verifying.
Why: when respond({type: "terminal"}) failed after the PTY exited
(network error, server restart), the panel sat on an exited terminal with
only Cancel. After a renderer reload nothing can reattach to a command
PTY (its ticket is spent and the session pool is gone), so the login
process kept running until the app quit.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (278 test files,
install, typecheck, format, tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: - The flow store's auth events now carry neither the interaction nor the message: flowId, instanceId, phase, methodId and purpose only. Clients already refetch auth.status (per user) when one arrives. - toAuthTerminalInvocation passes an absolute cmd.exe for Windows .cmd / .bat shims (new resolveWindowsComSpec: ComSpec when absolute, else SystemRoot/windir, else C:\Windows, + System32\cmd.exe). - EngineAuthClientCapabilities documents the rule for terminal methods: offered to every client, embedded on desktop and shown to copy in a browser (no launch ticket); a driver leaves one out only when it cannot work as a copied command. Why: engines.onEvents streams every event to every subscriber, so driver error text and messages naming the instance reached other users' streams while the details stayed per-user behind auth.status. On Windows a GUI-launched app can lack ComSpec; buildSpawnInvocation then falls back to a relative "cmd.exe", which Electron main's absolute-path check refuses, so the embedded sign-in terminal could never start for npm-installed CLIs. The P13 note in the lane report said browsers should not get terminal methods, which contradicts the behaviour the scope asks for (a copyable command in browser mode); the contract now states the actual rule. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (278 test files, install, typecheck, format, tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: - flow-store.test.ts: a terminal interaction from the flow store goes through the renderer's toDesktopTerminalCommand and Electron main's prepareTerminalCommand, with a fetch that redeems the ticket as POST /api/engines/auth/terminal does. Main runs exactly the vended command, args (with a space), cwd and non-secret env, refuses the spent ticket, and refuses a renderer that changes the args (that ticket is spent too). Main's module is imported through a computed URL so tsc does not type-check its plain JS. - copilot-sdk.test.ts: resolveCopilotLoginCli picks the user's own copilot on PATH when the instance runs on the bundled runtime, and the instance's runtime (as a Node script) when it is a CLI. Why: both sides of the ticket check had their own hand-written fixtures, so a drift between what the renderer echoes and what main compares would only show in the app. resolveCopilotLoginCli had no tests. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (278 test files, install, typecheck, format, tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The shared ACP engine (P12) replaced engines/cursor-acp; the Cursor auth controller now finds `agent` with resolveCursorBinary from acp/agents/cursor, the same lookup the driver uses. Gate G0: typecheck ok, all test files pass, format ok, tripwires ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- contract/usage-limits: helpers for building, sorting and merging usage
limits (sparse merge by window id that keeps known reset times,
"unsupported" stays authoritative, a failed read keeps the last good
windows), an update schema, and an optional `unavailable.action`
("read-keychain") the UI can offer.
- platform/usage/limits-store: the one place limits live per (user,
instance). Full reads through `driver.usageLimits.read` at most once per
TTL (5 min; 1 min after a failure; 30 min for accounts that cannot
report), one at a time, bounded by a timeout that aborts the read; live
reports merge in between. `reportEngineUsageLimits()` for runtimes.
- platform/usage/enricher: the "usage" snapshot enricher. It seeds limits a
probe got for free, starts a background read when one is due and the
instance is usable, and overlays what the store knows. The snapshot
service takes the store as a dep (wired in getEngineSnapshotService):
store changes are republished on cached snapshots, reportUsageLimits
goes through the store, instance changes forget the instance's usage.
Enrichment input gains `userId`.
- Readers (ported from t3code, MIT): Claude through the SDK's experimental
usage request on an idle query (looked up at run time, unsupported when
absent); Codex through the instance's app-server
account/rateLimits/read (API-key and Bedrock sign-ins skipped); Cursor
through the dashboard API with CURSOR_AUTH_TOKEN or the CLI auth file;
OpenCode Go only when an `opencode-go` credential is configured in
OpenCode, otherwise unsupported without any request.
- Cursor's macOS Keychain login is read with /usr/bin/security only from
api.engines.usage.readCursorKeychain (an explicit user action); the
token stays in process memory per instance and is dropped when Cursor
refuses it. Without it, Cursor on macOS reports "read-keychain".
- Live updates: Claude rate_limit_event and Codex
account/rateLimits/updated now feed the run instance's usage.
- api.engines.usage.{get, refresh, readCursorKeychain}.
- Catalog: reportsUsageLimits for codex, claude, cursor and opencode.
Why: P11 usage-limits row (t3features "Usage limits"); the contract had
the snapshot field and a passthrough "usage" enricher but nothing filled
them, and the existing Codex rateLimits endpoint had no consumer.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 277 test files, format, 59 tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: - Composer: a compact usage chip next to the context ring for the selected instance (engines whose driver reports usage limits). It shows the fullest plan window and, in its tooltip, every window with its reset time. Nothing renders before usage is known or for other engines. - Settings → Engines: a "Usage limits" section listing usable instances that report usage, with bars per window, the time of the last read, a Refresh button (api.engines.usage.refresh) and, when Cursor on macOS needs it, "Read login from Keychain" behind a confirmation dialog that explains the macOS prompt (api.engines.usage.readCursorKeychain). - useEngineUsageLimits(instanceId) reads api.engines.usage.get; the EngineEventsBridge folds snapshot events into those queries (syncUsageLimitsFromEvent), so the chip follows live rate-limit updates without polling while the event stream is up (5-minute poll otherwise). - components/engines/usage-limits.ts: pure presentation helpers (tone, percent, reset phrasing, which instances to list, notices). Why: P11 usage-limits UI (composer chip and settings section); the store and readers landed in the previous commit. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 278 test files, format, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- lib/ai/chat/engines/slash-commands.ts (client-safe): Codex's
Sentinel-run commands (compact, review, rollback), normalisation of
runtime-reported commands, the composer menu built from an instance's
commands ("native" ones are inserted as `/name ` text for the runtime,
"sentinel" ones run through a driver action), the long-standing
per-driver commands as a fallback until an instance reports its own,
and canRunSentinelSlashCommands for the thread gate.
- Claude: the status probe keeps the initialize response's `commands`
(skills included) on the status and in its last-known-good snapshot;
the driver reports them as the instance's slashCommands.
- Codex: the driver reports its three app-server commands once the CLI is
detected.
- ComposerEngineOption carries `slashCommands` when an instance reports
any (so the composer catalog refreshes when they change).
- Composer: the slash menu comes from the selected instance's commands,
drops commands a listed skill already offers, and shows argument
hints. Execution is a per-driver action map instead of a Codex-only
branch; thread-screen gates it through canRunSentinelSlashCommands.
Why: P11 "Slash commands per agent" (t3features): commands were a static
Codex-only list plus a hard-coded Claude list that included commands the
SDK session does not run. Copilot, OpenCode and Cursor report none yet
(ACP available_commands_update arrives with P12).
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 279 test files, format, 59 tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Claude skills natively What: - lib/skills: discovery, watching and lookups take `globalDirectories`, the global folder a source kind uses instead of `<base>/.claude/skills` and the like; snapshots are cached per workspace and set of folders. Installs and uninstalls take `globalSkillsDirectory` for global scope. - lib/skills/instance-skills: Codex, Claude and Copilot keep global skills in `<home>/skills` (CODEX_HOME, CLAUDE_CONFIG_DIR, COPILOT_HOME, via platform/instance-homes). resolveSkillInstanceContext picks the selected instance for its driver and every other driver's default instance. - api.skills: list, get, install and uninstall accept an optional `instanceId` and use those folders (installCustom and the registry use the default instances). Without an instance home the calls are exactly as before, so ~/.claude/skills, ~/.copilot/skills and skillsBasePath keep working. - Composer: lists the selected instance's skills when it is not its driver's default instance. - Claude: a `$skill` chip of a Claude skill becomes Claude Code's own `/skill args` invocation, sent as the message's last text block, with earlier mentions rewritten inline (planClaudeSkillDispatch, ported from t3code ClaudeSkillDispatch.ts, MIT). `.agents` skills stay prose. Why: critique G11 and P10c deferred finding 4: Claude and Copilot skills ignored an instance's CLAUDE_CONFIG_DIR or COPILOT_HOME (only Codex followed its default instance's home), and Claude treated `$skill` as prose. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 281 test files, format, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: the usage enricher clears the store's limits for (user, instance) and publishes no usage when the probe reports the instance unauthenticated. The store gains clear(userId, instanceId), which also drops a read still running for it. Why: without it, the bars of the account that signed out stayed on the snapshot (the store overlay and the previous-snapshot fallback both kept them) until the next successful read after signing in again, possibly with another account. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 281 test files, format, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: - contract/usage-limits: a runtime's live windows now replace an `unsupported` read, and resolveEngineUsageLimitsAfterRead takes `liveSinceLastRead`: an unsupported read keeps windows a runtime reported since the previous read instead of wiping them. - limits-store: tracks live reports per entry and passes that to the merge; the TTL now counts from when a read starts, and a read is due 10% before its TTL (USAGE_LIMITS_DUE_SLACK_RATIO). - Claude reader: an SDK without the usage request is remembered per query factory, so no Claude Code process is started for it again; `rate_limits_available: true` with no `rate_limits` (usage endpoint failed) is now probeFailed, not unsupported; a subscription login that cannot be read on demand says its usage shows during runs. Why (review findings): - "unsupported" was authoritative, so a missing experimental SDK method or a token without profile scope silenced every live rate_limit_event for that instance, and the 30-minute re-read kept it that way. An account without plan limits never streams windows, so letting first-hand windows win costs nothing for API-key logins. - Background reads start when the 5-minute refresh tick's probe ends and the TTL counted from when the read ended, so the next tick landed a few seconds short and reads ran about every 10 minutes. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 281 test files, format, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hange
What:
- cursor-keychain: tokens are keyed by (user, instance);
readCursorKeychainToken/getCursorKeychainToken take the user id and
forgetCursorKeychainToken({instanceId?, userId?}) drops one user's,
an instance's or all logins.
- Driver contract: usageLimits.read receives `userId` (the usage store
passes it) and an optional usageLimits.forget(instanceId, userId?).
- The usage enricher calls forget for the account that signed out; the
Cursor driver's invalidate (run by handleInstanceChange when an
instance is reconfigured or removed) drops that instance's logins.
- api.engines.usage.readCursorKeychain passes the signed-in user.
- New routers/engines/usage.test.ts: the procedures hand the session
user to the service, map EngineUsageError and Keychain errors to tRPC
codes, and the Keychain is read only from its own procedure.
Why (review findings): the token sat in a process-global map keyed by
instance id only, so another session user with the same default
"cursor" instance reused a login they never approved, and a signed-out
or reconfigured instance kept reading usage with the old login.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 282 test files, format, 59 tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- skill-dispatch: getClaudeSkillRoots({cwd, env}) lists the folders the
CLI loads skills from (<cwd>/.claude/skills and <CLAUDE_CONFIG_DIR or
~/.claude>/skills); getClaudeDispatchSkillNames(context, roots) keeps
a Claude skill chip only when its folder sits directly in one of them.
Chips without a folder stay prose.
- run.ts computes the roots from the run's cwd and the env the CLI is
started with (options.env, instance home included).
Why (review finding): dispatch trusted the composer's target/sourceKind.
With skillsBasePath set, Sentinel lists <base>/.claude/skills as Claude
skills that Claude Code never loads, so `/name` reached a CLI that does
not know it and the prompt could become an unknown-command reply. Such
chips now keep the pre-dispatch behaviour (prose plus the
<referenced-skills> prefix).
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 282 test files, format, 59 tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: - usage-limits: describeUsageCheckedAt (`Read at 14:32` today, `Read Oct 5, 14:32` before) and isUsageLimitsStale (instance not usable, or read more than 15 minutes ago). - Settings → Engines → Usage limits: the label carries the day when the read is not from today (full time on hover), stale windows are dimmed with a warning-toned label, and a not-ready instance says the bars are its last known usage. - The composer chip's tooltip shows when the usage was read. - New engine-usage-limits.test.tsx renders the section (static markup): the Keychain offer reads nothing on render, carried-over windows are marked, and accounts that never report stay hidden. Why (review finding): an instance that is installed but not usable keeps the usage of its last snapshot (status.json, up to 7 days old), and the row showed it as plain bars with a time of day only, which read as current. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 283 test files, format, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- slash-commands: getRunnableComposerSlashCommands(commands, {actions,
canExecute}) keeps inserted ("native") commands always and
Sentinel-run ones only when the composer has an action for them and
can run it on this thread.
- chat-composer/index.tsx uses it for the menu instead of an inline
filter (same result: the editor already hid execute commands without a
handler).
- Test: Claude-reported commands are inserted with their hints on any
thread; Codex compact/review/rollback are hidden without a Codex
thread or an action.
Why (review finding): the generalized slash menu was only covered at the
helper level; the filtering that decides what a thread offers lived in
the component.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 283 test files, format, 59 tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hers Snapshot enrichment inputs now carry the instance owner's userId (usage limits are per user); the manifest and maintenance enricher tests build their inputs with it. Gate G0: typecheck ok, all test files pass, format ok, tripwires ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: when a Claude or Codex thread has no native session to resume (its instance's continuation key changed, for example the home directory was edited, or the stored session state is gone), the first prompt of the new session now carries the thread's prior transcript, wrapped in a <conversation_history> block ahead of the new message. Copilot already did this; its bootstrap prompt moved unchanged into the shared runtime/history-replay.ts, which Claude and Codex now use too. - The history comes from the stored active transcript (truncated at the checkpoint anchor), never from client-sent messages: everything before the edited message for an edit, else the transcript without the message being sent. - Claude and Codex keep the newest messages within 200k characters and say how many earlier ones were left out. Copilot's prompt is unchanged (golden test). - A resumable session is still resumed without any replay, and a plan/chat mode switch on a resumable session still starts fresh without replay, as before. - On Codex app-servers without collaboration mode, the replay precedes the plan-mode preamble. Why: P10c deferred finding 7 (driver-contract §3 "null on mismatch ⇒ fresh session + replay"). It must land before Settings exposes editing an instance's home, which changes its continuation key. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 267 test files, prettier, 59 tripwires). The new Claude and Codex tests fail with the replay disabled. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: Settings → Engines lists each driver's instances under a driver heading and can now add, edit, enable/disable, remove or reset them, and edit their custom models (api.engines.instances, P10). - Add instance (implemented external drivers): name, accent colour (presets or any #rrggbb), binary path, home directory for drivers that have one (CODEX_HOME, CLAUDE_CONFIG_DIR, COPILOT_HOME) and environment variables. The server derives the id from the name. - Edit sends only what changed; config keeps the keys the dialog does not edit (launchArgs, …). Changing the home warns that threads on the instance start new native sessions with their conversation sent along (the replay landed in the previous commit). - Environment editor: per-variable secret toggle. Stored secrets come back redacted and are echoed back redacted; typing replaces them; turning one into a plain variable needs a new value (mirrors the registry rule); a value that no longer decrypts is flagged "Re-enter". - Each driver section says what isolates its instances: home-based drivers isolate sign-in and settings per home, while Cursor, OpenCode (and Pi later) share the CLI's sign-in and differ only by env and binary. - Disable and remove surface the G9 in-use guard: the instance's references (threads, automations, user default) are fetched first, and a change to an instance in use asks to confirm, then runs with force. Default instances offer "Reset to defaults" instead of remove. - Custom models editor per instance: id, display name, and options kept, dropped or copied from a model the engine reports; writes engine_instance.custom_models. - The instance card shows its accent colour and takes an `actions` slot. Logic lives in components/settings/engines/instance-management.ts with unit tests; the env editor, card and driver section have render tests. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 270 test files, prettier, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: the composer's engine submenu (composer actions → Engine) now has one entry per driver. A driver with a single instance stays a plain entry; a driver with several becomes a section headed by the driver label, with its instances indented underneath, each with its accent colour dot (neutral when it has none). The "Engine" row shows the selected instance's label, with its accent dot when its driver has several instances, instead of the raw driver kind. Grouping and the row summary are pure helpers in chat-composer/engine-menu.helpers.ts with unit tests. Selection keys are still instance ids, so nothing else in the composer changes; switching a thread's instance stays out of scope (the P10c guard holds). Why: P10c left a flat list of instances; with the add-instance UI a driver can have several, which need grouping to tell apart. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 271 test files, prettier, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: the new-automation modal and the automation detail screen pick the engine as a driver, plus an instance (with accent colours) when the driver has several or the automation already uses a non-default one, and offer a picker per model option (OpenCode's agent and variant, and any select option a model or custom model describes, except the reasoning effort, which keeps its own field). Picks are stored as automation.model_options. - Each option picker starts at "Model default (…)"; values left there are not stored, so automations without picks behave as before (empty model_options, model defaults). - Only values the selected model still offers are stored; changing the model drops picks the new model does not offer. Stored model_options are read back through parseEngineOptionSelections. - Picking another driver keeps the current instance when it belongs to that driver, else takes the driver's default instance when it can run, else its first one that can. - An automation whose instance is gone (removed or disabled) keeps it listed, disabled, and the form says to pick another engine instead of sending a bad request. - "Use default model" clears the stored picks. The shared pickers live in automations/automation-engine-fields.tsx, the logic in automation-form-helpers.ts (getAutomationEngineOptions is replaced by driver and instance options); both have tests. Why: P10c deferred finding 5 (no UI wrote model options, so automations always ran on model defaults) and the P11 automations instance picker. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 272 test files, prettier, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: saving an automation now writes model_options only when it knows the selected model's options: null for "Use default model", the picks the model offers when the model is in the catalog, and nothing (the stored value is kept) when it is not, for example while the catalog loads or while the instance is unavailable. resolveAutomationModelOptionsForSave in automation-form-helpers.ts decides, with tests. Why: the previous commit computed model_options from the descriptors of the model found in the catalog. Saving the detail screen before the catalog arrived, or with the automation's instance unavailable, found no model and wrote null, wiping the stored picks. Before that commit the forms never sent model_options, so this restores the keep-what-is-stored behaviour for those cases. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 272 test files, prettier, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rrides
What:
- Environment editor: a row keeps its stored (redacted) secret only while
it has the stored name and no typed value. Renaming the row leaves the
secret behind and asks for a new value ("Enter a value for NEW: the
stored secret of OLD does not move to a new name."); renaming it back,
or clearing a typed replacement, keeps the stored secret again.
EnvVarDraft.storedSecret becomes storedSecretName.
- Registry: encodeEnvironment refuses a redacted echo for a name with
nothing stored ("There is no stored value of X to keep. Enter its
value.") instead of storing an encrypted empty value.
- Home directory: the new-session warning also fires when the driver's
home variable (CODEX_HOME, CLAUDE_CONFIG_DIR, COPILOT_HOME) is added,
removed or changed among the environment variables, and the field shows
that the variable set there is the one used.
Why: review of the P11 instance lane. Renaming a stored-secret row sent
{name: NEW, valueRedacted: true, value: ""}; the server found nothing
stored under NEW and stored "" for it, which then overrode the app's own
value of NEW for the instance while the dialog reported success. Clearing
a typed replacement did the same under the old name. The registry lets an
environment home variable win over config.homePath for the effective home
and continuation key, so changing it moved the home without the warning.
Tests: client drafts (rename, rename back, clear), registry (refused
redacted echo on create and update, stored value kept), a round trip from
the dialog's drafts through registry.update, and home-variable notices.
Mutation: without the registry guard the refusal test fails.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 272 test files, prettier, 59 tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What: - Claude and Codex now replay the conversation whenever they start a fresh native session, not only when no session is stored: the replay condition follows the resume decision (Claude's shouldResumeExistingSession, Codex's resumableCodexState), so a thread mode change (plan to chat, "Implement plan") also carries the prior transcript into the new session, as Copilot already does. - The plan-mode preamble now comes before the replayed history, so the history's closing "New message:" introduces the user's own words with nothing in between (Claude's fresh plan session and Codex's fallback for app-servers without collaboration mode). Why: review of the P11 instance lane. The replay keyed on whether a session id was stored, but a mode change starts a fresh session even with one stored, so "Implement the latest approved plan from this thread" went to a session that had never seen the plan. And with the plan contract after "New message:", the model could read it as part of the user's message. Tests: Claude and Codex mode-switch replay, and the preamble order (Claude fresh plan session, Codex without collaboration mode; the old Codex order test is updated). Mutation: with the previous condition and order the three new tests fail. Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install, typecheck, 272 test files, prettier, 59 tripwires). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
What:
- Saving the automation detail screen while it still selects the stored
instance and model keeps the stored model_options entries of options
the model does not describe right now (not a picker option and not the
reasoning option). Picker options still follow the form, and choosing
another model or instance still starts over.
resolveAutomationModelOptionsForSave takes the stored selection
({instanceId, modelId, modelOptions}) as an optional fourth argument.
- pruneAutomationOptionValues only drops values an option the model
describes no longer offers. Values of options it does not describe stay
in the form (they are not saved for that model), so the pickers show
them again when the catalog entry recovers.
Why: review of the P11 instance lane. 9c25b78 kept the stored picks only
while the model was missing from the catalog. A model listed without its
option descriptors (OpenCode fallback models carry no agents or variants,
and a probe can lose the agent list) made the prune effect clear the form
values, and the next save of any field (the title, say) wrote
modelOptions: null.
Tests: the degraded catalog (no descriptors, only some), and another
model or another instance with the same model id; the prune test now
covers unoffered choices and undescribed options.
Gate: node scripts/verify/gate.mjs G0 --jobs 3 passed (install,
typecheck, 272 test files, prettier, 59 tripwires).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After merging the P11 lanes the instance card also hosts the sign-in panel and the install/update controls, which need a tRPC provider; the instance-card render test stubs them (they have their own tests). Gate G1: typecheck ok, all test files pass, format ok, tripwires ok, next build ok. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Draft PR that tracks the Sentinel revival. Each phase lands here as one or more commits, and every commit passes the gate before the next one starts. The gate checks frozen install, typecheck, the full test suite, format, tripwires and
next build.Goals:
Phases
Notes for reviewers
@ai-sdk/mcp2: OAuth MCP servers authorized under MCP 1.x have to be re-authorized once. The stored tokens carry no authorization-server pin.strict: false, pinned by a test. AI SDK 6 left the field out, andstrict: truerejects schema keywords Sentinel uses.@ai-sdk/xai5 is Responses-only. Sentinel now sendsstore: falseto keep the old no-retention behaviour.minimumSystemVersion, so macOS 12 installs are not offered an update they can't run.COPILOT_CLI_PATHstill overrides the bundled runtime.claudebinary; its native CLIs are not shipped.updatedInput.item/tool/requestUserInput, permissions and elicitation requests.thread/rollbackwas replaced bythread/revert.total/lastshape.codex.cmdsupport.session/cancelis sent as a notification, so Stop can no longer hang.Test plan
🤖 Generated with Claude Code