Skip to content

Release v2.6.0 - #59

Merged
Alex375 merged 34 commits into
mainfrom
dev
Sep 24, 2026
Merged

Alex375 merged 34 commits into
mainfrom
dev

Conversation

@Alex375

@Alex375 Alex375 commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Release v2.6.0. Voir les commits de dev depuis la dernière release.

🤖 Generated with Claude Code

Alex375 and others added 30 commits September 21, 2026 14:01
…s live turns

A conversation started from the TOSSE tasks view showed its first prompt
TWICE — once optimistically, then again at the tail of the thread once the
first burst of work was over.

`loadConversationHistory` is additive and assumes a fresh entry (its own doc
says so, and the app-control send_message/read_conversation paths pre-hydrate
for exactly that reason). "Start" creates the conversation, sends, and STAYS on
the tasks page (`startStaysOnTasks`), so `ConductorConversation` never mounts
and the loader never runs. By the time the thread is opened — typically after
the end-of-turn notification — the conversation has a session id, so the whole
transcript is replayed ON TOP of the turns already streamed: the disk user line
carries the transcript uuid (≠ the optimistic `user_N` id) and lands as a second
bubble at the end, while every same-id assistant message gets its blocks
appended a second time.

`addConversation` now marks a row hydrated on insert when it arrives WITHOUT a
session id: a conversation born in this run has no cold transcript, everything
it will ever show arrives live. Rows inserted WITH a session id (undo of a
delete, `reactivateDiskConversation`, a Codex fork) still load their history.

Covers the whole "conversation that runs before it is ever opened" class, not
just the TOSSE path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The repository was renamed; pushes to the old name only work through
GitHub's redirect. The updater endpoint in tauri.conf.json and the TOSSE
repository row already point at the new name, and release.yml goes through
$GITHUB_REPOSITORY — so the repo name was the one piece of the technical
identity that did follow the rebrand, and the context said the opposite.

Synced from the TOSSE repository context (single source of truth).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…MCP steps

A call to the TOSSE MCP server rendered like any other: a generic plug glyph, a raw
`claude ai TOSSE : update_task_status` label, and a run header counting "3 tools". Which
task was filed, which status moved, which level of context was touched — none of it was
readable without expanding rows.

Two shapes, decided in one pure module (`tosseTool.ts`) and honoured by every renderer:

- a WRITE gets its own card (`TosseToolCard`): the task's title, the exact tool that ran,
  the status it moved FROM → TO, the assignee, and a click that opens the task in the
  conversation's side panel (browser fallback when the CRM is not signed in);
- a READ stays a step inside its run, with the CRM's mark and a human label — "Read tasks"
  plus a "12 tasks" count badge — instead of its wire name.

Notable decisions:

- The CRM never sends a previous status back, so "À faire → En cours" is DERIVED from the
  thread: every earlier TOSSE call that saw the task is a sighting of its status, and the
  last one strictly BEFORE this call is where it moved from. No sighting → one pill, never
  an invented origin. A substring pre-filter guards the parse: `list_projects` comes back
  at ~200 kB and this runs in a store selector, on every streamed token.
- Writes FOLD with the work under clean output (like `message`, unlike `artifact`): a
  /pickup or /done fires a burst of them, and holding them all in clear would empty the
  fold of its purpose.
- The server is recognised by "segment contains tosse" + a placeable tool, never by one
  hard-coded name — the same CRM arrives as `claude_ai_TOSSE`, a plugin variant, or plain
  `tosse` on the Codex side.
- Zero silent error: a refused write is red with the CRM's reason and shows NO status pills
  (it moved nothing); a call in flight reads "Saving…" and holds the fold open.
- Parsing is tolerant throughout — a drifted field costs only itself, never the card.

Reversible via Settings → TOSSE → "TOSSE actions in the conversation" (on by default). The
effective value ANDs the preference with the CRM session: the tab does not exist while
signed out, so leaving the rendering on there would hand someone cards with no switch.

Also: `Ico` can render the CRM rose by name; `?demo=tosse` fixture for visual verification;
`cleanFoldIdentity` now mounts under a QueryClientProvider (the gate reads the CRM session).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…only design

"Tinted conversation rows" and "Status inside the composer" were regression
escape hatches, not real choices — nobody wants the old look back. Drop both
prefs and keep the new design unconditionally.

- display.ts: remove `sidebarStatePills` and `composerStatusBand`.
- Settings → Display → Appearance: drop the two toggles (and their search-index
  entries); the "Time on sidebar rows" hint no longer points at a gone toggle.
- ConductorSidebar: the pill IS the row, always — `StatusDot`, the `data-attn`
  tint and the `data-rows="pills"` opt-in attribute are gone.
- ReviewBar (the classic full-width bar) is deleted; `useMarkSeenShortcut` moves
  into ComposerStatusBand, its only remaining caller.
- CSS: unscope the pill block from `[data-rows="pills"]` (source order now
  carries the override), drop the dead `[data-attn]` rules and the
  `.cv-reviewbar` tones nothing renders anymore. The bar's base + error tone
  stay — AuthWarningBar reuses them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…o it

Build artifacts had reached 21 GB across tosse-code and flightdeck-server
(19.5 GB of pure artifact) on a Mac with 31 GB free. Nothing in the workflow
ever reclaimed them: /release has no cleanup step at all, so landing — not
releasing — is the only hook that runs on every feature.

- New /cleanup skill: the single place that knows every location (the four
  target dirs, orphan com.tosse.desktop.<slug> identities, never-gc'd repos)
  and every guard rail (protected prod/dev identities, live features, builds
  in flight, no history rewrite). Full purge by default, --caches to keep
  incremental state.
- /land step 7c delegates there instead of reimplementing the purge.
- Fix step 7b: it removed HTTPStorages/$ID but not $ID.binarycookies, which
  macOS writes NEXT TO that directory — every landed feature leaked one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A repository on a paired server looked exactly like a folder on this Mac:
`Repo.machineId` drove real behaviour (Claude account, title push, IDE
refusal) but nothing ever rendered it. Two folders can even carry look-alike
paths once truncated, so you could message an agent believing it works on
your Mac while it runs on a server.

Adds one mark — globe + the server's name, with user@host:port in the tooltip
— on the three surfaces that answer "where does this run?": the sidebar repo
header, the Flight Deck swimlane header and the stream card. A local
repository stays unmarked (the default, and the overwhelming majority).

No setting: this is missing information, not a deliberate change of
experience, so there is nothing to opt back out of.

A `machineId` naming no paired machine reads as an amber "unknown server",
never as local — degrading it would resurrect the exact confusion this mark
exists to end.

Colour is left unspent on purpose: the machine HEALTH state (unreachable /
degraded) lands on this same chip next and needs the attention colours.

Sidebar layout: the repo name now carries its real flex basis so the chip,
which shrinks eight times faster, gives way before the name does; on header
hover the chip drops to the globe alone and hands the space back.

`?demo=remote` seeds a local folder, one on a paired server and one orphaned,
so all three read differently in dev.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The TOSSE badge of a repository hosted on a paired server showed the
"attention" flag and "This folder's git remote is unreadable", while its
origin matches the CRM's url perfectly.

Nothing was wrong with that folder. `repo_tosse_links` selected only
`id, path, tosse_repository_id`, so the association check ran THIS Mac's
`git -C <path>` on a path that only exists on the server. It exits 128 with
"cannot change to …" — byte for byte what a deleted folder returns — which
landed in `remote_error` and became a warning on the badge. A fault of ours,
reported as the user's.

It is decided now before the spawn rather than guessed from its stderr
afterwards: the query carries `machine_id` (LEFT JOIN for the server's label,
so an un-paired one stays remote instead of dropping out of the view), and
`remote_probe` returns `Skip(machine)` for it. No git runs, no `remote_error`,
and the link carries the machine instead — an ordinary limit like "not a git
repository", not a failure.

The card said "This folder has no git remote", which is untrue: it has one, we
just never read it. `whyUnmatched` now names the server and points at the
manual pin — which works on a remote folder, being pure local SQLite — and the
facts row stops labelling a server path "Local folder". The unchecked state
says it too, so "Refresh" is not read as a promise no refresh can keep.

Resolving those remotes over SSH would make the automatic match work there as
well; that is a separate task, and it costs a round trip per remote folder on
a path called at load.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The remote mark shipped as a globe + the machine's name, permanently. At a
glance the only question worth answering is "is this on this Mac?", which one
small glyph answers — the name competed with the repo name beside it for room
it did not need to take. Which server is a question you ask by pointing.

So the mark is now just the globe, in the quietest text tone, and the machine
name plus its ssh target moved to hover.

`title="…"` could not carry that: it waits about a second, which reads as
nothing happening for information the UI deliberately hides at rest, and it
cannot be styled. Adds `ui/Tooltip.tsx` — appears immediately, portaled to
<body> so the Flight Deck swimlane's overflow cannot clip it, placed by the
pure `tooltipPlacement` (centred, above by preference, flipped only when it
genuinely does not fit, both axes clamped on screen). It closes on scroll
rather than chasing the trigger, and carries `aria-label`/`aria-describedby`
so a pointer-only tooltip is not invisible to everyone else.

Reverts the sidebar flex changes the old chip needed: with no text to hold,
the mark takes ~13px and the repo name is no longer squeezed, so `.cv-repo-n`
and `.cv-repo-title` go back to their original `flex:1`.

Colour stays reserved for the machine HEALTH state landing on this same glyph
next; amber "unknown server" remains the one exception, a fault not a state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two things the dev build surfaced.

A globe was the wrong mark for a paired server. It says "the internet",
where a VPS of your own is a MACHINE — and at the 13px the remote mark
renders at, its four crossing strokes turn to mush. Adds a `server` glyph (a
rack unit with its status light, the conventional pictogram) and moves the
machine-sense icons onto it: the remote repo mark, the IDE view's repo
picker, the sidebar "+" menu's per-machine entry, and Settings → Remote
servers (SSH). `globe` stays for the NETWORK sense it was right for all
along — the claude.ai bridge, the phone relay, an http MCP server.

A bare stack of bars was the other readable option at that size, but it reads
as a list; the status light is what makes it a server.

SIDEBAR_MIN drops 190 → 150. Measured at every width down to 120: nothing in
the sidebar overflows its box, names simply ellipsis earlier and the header
buttons are `flex:0 0 auto`. 190 was a guess that stopped people reclaiming
horizontal space they were entitled to. 150 is where a repo name still shows
enough characters to tell apart — a reason to stop, unlike "it might break".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A folder on a paired server could only ever be associated BY HAND. Its
`origin` matches the CRM's url perfectly — the probe just never ran where the
folder is, so the automatic match had nothing to work with.

It now runs there. `tosse_probe_remote_origins` asks each server, in one
connection per machine, for the `origin` of every folder that lives on it,
and caches each answer in SQLite (v15). Matching is unchanged: the url simply
arrives from elsewhere and goes through the same `normalize_remote_url`.

The read is deliberately NOT on `tosse_repo_links`'s path. That one is what
the sidebar waits on at load, and an SSH round trip does not belong there; it
reads the cache instead. So the match is instant on every later load, and it
still works with the server switched off — which is why the answer is
persisted rather than memoised. The sweep runs beside the query and the front
refetches only when a url actually MOVED, or the app would loop: sweep →
invalidate → refetch → sweep.

Two columns, not one. `remote_origin_probed_at` is what says we ever LOOKED;
`remote_origin_url` is what we found. A single nullable url would make "never
asked" and "asked, this repo has no origin" the same row, and that difference
is the one this whole feature is built to keep. It reaches the UI as
`machine.originRead`, so the card can tell "we have not been able to read it
on <server> yet" from "that folder has no origin there" — the same sentence
for both would blame a server that is working fine.

An unreachable server is not an error: it is logged, the folders keep the
answer they already had, and nothing is reported as a fault — the false alarm
the machine-aware probe removed does not come back through this door. A
server with no `git` is asked about once, up front, so its absence cannot read
as "none of your folders are repositories"; `gone` and `norepo` cache nothing,
since writing "no origin" for an unmounted folder would un-match a repository
that is merely asleep.

The two pre-migration fixtures now seed `repos`: v15 is the first migration to
touch it, and a real database left by an older app has always had that table
(v1 creates it), so leaving it out described a state that cannot exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Alexandre's pick from the candidate sheet, over the rack unit that shipped
yesterday. What matters at a glance is that the thing is REACHING you from
somewhere, and a mast on a base says that where a stack of bars says "list".

Every machine-sense icon follows it, since they all read the one `server`
entry: the remote repo mark, the IDE view's repo picker, the sidebar "+"
menu's per-machine entry, and Settings → Remote servers (SSH).

⚠️ The two arcs are spaced for the 13px the mark renders at. Arcs any closer
land under a pixel apart once scaled down and merge into one smudge, which is
also why there are two and not three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…5.5)

CLI 2.1.280 moved the `opus` alias to Opus 5.5, so the catalogue's "Opus 5" row had
been running Opus 5.5 under the wrong name — in the picker, in the live label (every
Opus id contains "opus") and in the Rust model-change notices — while Opus 5 itself
could no longer be picked. A fresh install also hid that row and defaulted to Opus 4.8.

Catalogue (transcribed from the 2.1.280 registry): the alias row reads "Opus 5.5" and
carries the id it resolves to (`modelId`), and Opus 5, Fable 5 and Mythos 5.1 get
pinned rows. Rows now name their `family`.

Factory picker = the first row of each family, except Mythos, derived rather than
listed: today Fable 5.1, Opus 5.5, Sonnet 5, Haiku 4.5. The factory default model is
the newest Opus (`opus`, front and Rust spawn fallback), so it follows each release.

Existing prefs: a new `seen` map (value -> model it ran) gives any model the user has
never seen the factory treatment — new older rows and Mythos arrive hidden, and a
moved alias has its old hide lifted (that hide was about Opus 5, which keeps it on its
own row). Blobs predating `seen` are read against a frozen LEGACY_SEEN. Stored default
models are left alone.

Labels: Claude ids are matched on each row's identity at an id boundary, so
`claude-opus-5` no longer claims `claude-opus-5-5`; the Rust `model_label` parses the id
instead of a `contains("opus")` ladder. Spend: Opus 5.5 priced at 4/20, Opus 5 history
on its own row and chart slot, Mythos priced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t be runnable

Settings -> Models showed every Codex model as "shown", and named GPT-6 Astra as the
Codex default on a machine whose codex (0.144.4) did not offer it at all.

- codexFactoryHidden: the Codex half of "newest by default". Each line (Astra, Sol,
  Terra, Luna, Mini) keeps its newest; a line-less model (gpt-5.5) goes to Available
  once a newer generation exists. Applied to the static list and, authoritatively, to
  the live `model/list` (noteCodexOffered) for models the user has never seen — a
  model they showed or hid keeps their choice (setHidden now marks it seen).
- effectiveDefaultModel: a new Codex conversation starts on the stored default only
  while the installed binary offers it, else on the binary's own `isDefault`. The
  stored choice is kept and wins again once the binary can run it.
- Static catalogue + effort ladders refreshed to codex-cli 0.156.1 (GPT-6 Sol and
  Luna added; GPT-5.4 Mini, gone from the binary, moved to the retired rows).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Expanding a live Claude MCP server in the extensions manager (claude.ai cloud
connectors included) now lists its tools with a permission each — Default, Allow,
Ask, Block — so e.g. reading mail can be allowed while sending asks first.

The choices are Claude Code's own permission rules, written as exact tool names
(`mcp__claude_ai_Gmail__send_message`) in ~/.claude/settings.json: the CLI watches
the file, so a running conversation applies them from its next tool call. Block
removes the tool from Claude's context in every mode; Ask prompts even in Auto and
Bypass; Default leaves it to the permission mode.

- Rust `extensions::permissions`: reads every rule that can reach an MCP tool from
  the managed, local, project and user settings files (a broken file is a warning,
  a broken USER file refuses writes); writes only exact `mcp__<server>__<tool>`
  names through the shared `write_settings` lock, then reads the file back.
- `mcp_status` now keeps each tool's description and the server's readOnly /
  destructive annotations (verified live: bare tool names, `readOnly` spelling).
- Front `mcpToolPermissions.ts` (pure, tested) mirrors the CLI: server-name
  normalization, deny > ask > allow whatever the file, allow globs only after a
  literal server prefix. A choice another rule would override is disabled, with the
  winning rule named; "Default" shows what applies without the user's rule.
- One-click "Read-only" per server (reads allowed, the rest asks — from the
  annotations, else a name guess drawn dashed that errs toward asking) and "Reset".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… Extensions

Tool permissions now have two scopes, as Armand asked: the global ones apply to
every conversation, the ones set in a conversation apply to it alone.

- Conversation (its ⌘E panel): the rules live in the session's flag settings layer
  (`apply_flag_settings{settings:{permissions}}` — verified live on 2.1.280: a
  second apply REPLACES the key, null clears it, `list_permission_rules` reports
  them as flagSettings). The layer dies with the process, so the app keeps them per
  conversation (localStorage `tosse:convToolPerms`), hands them to every spawn
  (`SpawnFlags.sessionPermissions` → re-applied right after `initialize`, a failure
  surfacing as a control error) and pushes changes to the running session — only
  committed once the CLI accepted them. deny > ask > allow still holds across
  sources, so a conversation can tighten the global rules, never loosen them; the
  rows say which global rule wins.
- Global (new Settings → Extensions tab, Claude-only): connectors & user MCP servers
  listed by a short-lived conversation-less `claude` (`fetch_global_mcp_status`, no
  model turn), their per-tool rules written to ~/.claude/settings.json, plus
  plugins (global enable + hot-reload bar), the user's skills and sub-agents, and a
  link to add connectors on claude.ai.
- Allow all / Ask for all / Block all next to Read-only (only what would hold,
  only what changes); the bar's buttons no longer shrink below their label.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ays it is remote

The app already knew how to ask a server whether it was there — `machine_diagnose`
is what draws each card's headline in Settings → Control → Remote servers. It just
never ran anywhere else, so the rest of the app found out a server had died the hard
way: a spawn that fails, a title push that quietly `console.warn`s, after the user
had already written to an agent that was never going to answer.

The same command now runs ambiently, and its verdict lands on the remote mark the
sibling task introduced — which had been holding the error colour in reserve for
exactly this.

- `ServerDiagnosis` gains `reachable`. It could not be derived from what was already
  there: `DiagnosisState::Failed` is also what a perfectly reachable server with a
  stopped `flightdeckd` produces, so the only other discriminant would have been
  matching `Failed`'s reason STRING — a reworded message silently becoming a wrong
  verdict. A Rust test pins the two meanings apart.
- `store/machineHealth.ts` holds the one answer, binary (reachable / not) per
  Alexandre. Unknown is its own state and renders exactly as before: a server is
  never accused on a guess, and a probe that could not even run is filed as such
  rather than swallowed.
- `MachineHealthHost` polls only machines that are paired AND host a repository, once
  per machine rather than per repo, paused while the window is hidden or Settings is
  open (which does its own), staggered. A failed remote spawn or title push triggers
  an immediate check — a trigger, never a verdict, since both fail for reasons that
  have nothing to do with reachability.
- The mark turns red and swaps to a crossed-out mast (shape, not colour alone), its
  tooltip states the reason and how long the server has been gone, and it becomes a
  button onto the server panel that already knows how to repair it. Conversations on
  that machine get a bar above the composer — non-blocking on purpose: the verdict is
  a snapshot, and locking someone out of their own conversation on a stale one would
  be the worse failure.

No setting: not seeing that a server is down is a defect, not an experience anyone
would choose.

`?demo=remote` grows a second, dead server so all four states can be seen at once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…band, not a floating card

It was built on `.cv-reviewbar` — the pre-`cv-sband` design, kept alive only by
`AuthWarningBar`. The conversation's status moved INTO the composer card two
redesigns ago (3a3fba3, then 89225f2 "make the tinted rows and composer status band
the only design"), and the stylesheet says so itself right above `.cv-reviewbar`. A
warning that looked unlike every other composer band read as a different kind of
thing.

It is now a `.cv-sband`: the card's own header, tinting the card's border and halo
through `:has()`, same icon/label/detail/action grammar as the settled-status band.

`ComposerBand` routes between the two, because `.cv-sband` is a HEADER — it pulls up
into the card's padding and carries its top corners, so two stacked would give the
composer two rounded tops, a doubled border and two competing tints. An unreachable
server wins: whatever the conversation's status is, it is about a turn that already
finished, while the server being gone is about every turn from here on — "Continue"
would fail on the spot and "Mark as seen" would tidy away the colour with the real
problem left invisible.

Still its own leaf, so the composer re-renders on neither status nor machine health.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, and retire cv-reviewbar

`AuthWarningBar` was the last thing floating above the composer in the pre-redesign
`.cv-reviewbar` card. It says the same kind of thing as the unreachable-server band
next to it — "your next message will fail" — so it is now the same band, through the
same `ComposerBand` router, and `.cv-reviewbar` is gone from the stylesheet.

Order, most-blocking first, each one making the ones below it moot: CLI missing →
account signed out → server unreachable → the conversation's own settled status.

⚠️ Behaviour change, and the point of doing this here: the two auth warnings are now
raised only on a LOCAL conversation. `binaryAvailable` and `useAccountsLoggedOut` both
describe THIS Mac, while a remote conversation runs its `claude` on the server against
the server's own credential store — `spawn_session` refuses a local account for one
outright ("Remote (SSH) conversations run on the server's own Claude account"). On a
remote conversation the old bar nagged about a binary and a sign-in that had nothing
to do with the agent being talked to, and sent the user to a Settings page that could
not fix it. The server's own equivalents are facts on its card in Settings → Control →
Remote, which the unreachable band leads to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… or global

The conversation's extensions panel (⌘E) now opens with a scope picker, "This
conversation" by default, and every setting below follows it: per-tool permissions,
a server on/off, and a plugin on/off.

- Conversation: the session's flag layer, it alone. Now also carries
  `enabledPlugins` (verified live on 2.1.280: `{id:false}` + reload_plugins drops
  the plugin's commands, clearing it brings them back) — renamed SessionOverrides,
  re-applied after `initialize` with a plugin reload when it holds any.
- Repository: `<root>/.claude/settings.local.json` of the working tree (a worktree's
  own), on this machine. Git is first told to ignore it in the clone's own
  `info/exclude` (never a tracked .gitignore) — git/mod.rs `toplevel` +
  `ignore_locally`. The shared `.claude/settings.json` is never written.
- Global: `~/.claude/settings.json` (also Settings → Extensions).
- A server is turned off by a `deny` on its whole rule name (`mcp__claude_ai_Gmail`)
  at the scope; off from a broader scope it can't be turned back on and says who
  did it. A server Claude Code itself disabled for the folder shows "Turn back on".
- Plugins show their state at the scope (its own say, else what it inherits), a
  note when a level above decides otherwise for this conversation, and write
  nothing of their own when the choice matches what they'd follow anyway.
- Auth buttons say they apply to every conversation; skills and sub-agents stay
  read-only with their source badge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e server" on screen

Refresh, a repair and a server sign-in all already re-diagnosed and filed the result in
the machine-health store, so the card, the sidebar mark and the composer band agreed.
"Retry" (phone access) did not: it lives in the PARENT's state, it makes a real ssh
round trip to the daemon, and nothing re-examined the machine afterwards. A retry that
visibly succeeded left the red headline and the red band exactly where they were.

The group now hands each card a `recheckToken` and bumps it when one of its own actions
has just talked to that server. The card re-runs `refresh` — not its mount effect, so
the facts the user is staring at are not blanked back to "Checking…" mid-round-trip.

Deliberately a trigger, not a conclusion: a retry can fail for reasons that say nothing
about reachability (a daemon too old, a refusal), so the diagnosis stays the single
verdict source for every surface.

Also: `record` now stamps the probe clock, whoever paid for the round trip. The Settings
card calls `machine_diagnose` directly (it needs the full diagnosis back, which
`probeMachine` does not hand over), so without this the ambient poll dialled the same
server again seconds after the user hit Refresh.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…els override both ways

Armand's review of the scope picker: each scope must only show what exists at its
level, show its OWN value (else the broader one it inherits — never a narrower one),
and a narrower level must be able to LOOSEN a broader one ("cautious by default,
permissive in this one repository"), not just tighten it.

Claude Code can't do that with its files — its rules resolve deny > ask > allow
whatever their source — so Flight Deck now keeps the cascade itself and hands each
conversation ONE consistent rule per tool through its session layer:

- store/mcpPolicy.ts: Global, per-repository (the app's repo record, worktrees
  included) and per-conversation levels of tool rules, server on/off and plugin
  on/off (localStorage). A change is pushed to every live Claude conversation it
  reaches; a spawn gets the resolved set (SpawnFlags.sessionOverrides). The first
  per-conversation store is migrated.
- mcpToolPermissions.ts: the cascade (narrowest level wins; within a level a tool
  beats its server), and `sessionOverridesFrom` — a server off is one deny, or a
  deny per OTHER known tool when a narrower level re-allows some (per-server tool
  cache kept for that).
- Plugins: Global stays Claude Code's own enabledPlugins; repository / conversation
  override it through the session layer (plugin reload).
- Each scope lists only what exists at its level (Global: the user's skills /
  sub-agents, user servers, connectors, user-installed plugins; Repository: the
  repo's own skills, everything usable in it).
- Claude Code's own MCP rules (its settings files) still apply underneath: listed
  as such, blocking only the choices they'd override, removable from the user file.
- Rust: the Repository file target, set_plugin_override and git ignore_locally
  are gone (Flight Deck no longer writes repository files); set_mcp_tool_permissions
  only serves to remove a rule from ~/.claude/settings.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ver the CLI allows

Armand: Claude Code's permissions are only the reference for what's default — what
is set in Flight Deck must win whenever it can be made to, and what can't be done is
simply not offered, without any "Claude Code prevents this" notice.

- Claude Code's own rules (its settings files) are now the BASELINE under Global:
  "Default" falls back to them ("→ Ask · from Claude Code").
- A file `ask` no longer limits anything: when Flight Deck says Allow, the session
  answers that prompt for the user (`auto_allow`, fed from the SessionOverrides allow
  list at spawn and on every change) — only for a prompt a settings-file rule raised
  (`decision_reason_type == "rule"`), never one from the mode / classifier / a safety
  check, nor a tool that itself requires a human.
- What really can't be overridden — a file `deny` (the tool never reaches the app), a
  server the files turn off entirely, the organization's `ask` — is just not offered.
- The external-rules box, its warnings and the "overridden" notes are gone; so is the
  Rust write path to ~/.claude/settings.json (permissions.rs is read-only again).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Armand: the Extensions page is Claude's (Codex's own stay in its ⌘E panel), so it
belongs with the rest of Claude Code's settings rather than as its own rail tab.

Settings → Claude Code gains a fourth sub-tab, Extensions, rendering the same global
page (connectors + their tool permissions, plugins, skills, sub-agents). The page
heading's subtitle follows the sub-tab so the ⌘E override note stays with it. The
"extensions" section is gone from the rail and from SettingsSection; the two search
entries now land on claudeCode/extensions (SETTINGS_SUBS updated, invariant tests green).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A multi-agent review of the 21 commits dev adds to main (8 dimensions × 4
file batches, every finding put to 3 independent refuters) confirmed 29
distinct defects. All of them are fixed here.

/cleanup — the only one that destroyed data
  Its repository root was hard-coded to a path that exists on no machine,
  so every rm -rf, the git gc AND the "is this feature still alive?" guard
  all tested somewhere else. The guard therefore always passed: running it
  today would have deleted the live side-panel-artifacts worktree's
  isolated database. Roots are derived now, and a root we cannot find
  purges nothing and says so. Three more: the identity inventory matched
  production's OWN sidecar files (com.tosse.desktop.plist, .binarycookies)
  past a guard that could not catch them; `reflog expire --all` destroyed
  every stash while the prescribed fsck check could not see it; and the
  "no build running" guard matched its own PATH, so the purge step was
  skipped every single time and the ~19 GB never came back.

Remote-origin sweep — a server that answered, reported as unreachable
  The sweep threw away every firm answer that was not a url (gone, not a
  repository, no git), which left remote_origin_probed_at NULL — the field
  that means "we could not ask". The card then blamed the server and
  offered "try again once the server is reachable", advice that could
  never work. Those answers are carried through to the UI now (migration
  v16), and the card names the real situation. Also: the refetch was
  edge-triggered on a boolean VALUE, so a second sweep returning true
  again never fired and Refresh visibly did nothing; a first probe finding
  no origin reported "nothing changed" although it had just flipped
  origin_read; the concurrency guard answered "nothing moved" for a sweep
  that never looked; a failed cache write was logged and swallowed; one
  sweep cost N tosse_repo_links invocations (one per mounted badge); a
  remote folder's path was used as a scan root for clones on this Mac;
  and the 30s timeout dropped the future without killing ssh.

  A remote's credentials are redacted at the point of capture: a clone can
  carry a PAT in its url, and this one was persisted to SQLite and printed
  on the card. Matching is unaffected — normalize_remote_url already drops
  the userinfo.

Machine health — the regression that undid 69da5cc
  Three producers wrote a diagnosis with no ordering guard, so a slow
  mount probe landed on top of the fresher Retry re-check and re-raised
  "unreachable" on three surfaces — exactly what recheckToken had just
  been added to prevent. Answers are ordered by when they were FIRED, in
  the panel and in the store. A verdict that arrives after the panel
  closed is now filed instead of dropped, and MachineHealthHost stops
  accumulating timer handles for the life of the effect.

Model preferences
  setDefaultEffort clamped against the STORED model while the row offered
  the EFFECTIVE one's ladder: picking "Ultra" silently stored `max`.
  A Codex default picked from the binary's live model/list was validated
  against the static catalogue and reset to the factory model on every
  launch. Hiding the model effectiveDefaultModel resolves to went
  unrepaired, so new conversations kept using it.

TOSSE card, thread and composer
  A task id from the agent's tool input reached a CRM URL with no check —
  ../admin/… normalises out of /tasks/ entirely; the canonical-UUID gate
  is now a shared module used by both surfaces. A resultless write in a
  SETTLED turn read "Saving…" for ever. The status arrow credited a call
  with a move made elsewhere. A blank title left the card headless.
  project_id was filtered out of an update_task's changed fields, hiding
  the only change it made. The refusal fallback read "refused to created
  this task". And with a warning band on screen ⌘↵ no longer acknowledged
  — it sent the draft, which then failed for the very reason of the band.

Verified: 2390 front tests, 1129 Rust tests, tsc --noEmit, pnpm build.
Bindings regenerated and committed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Making the "is a build running?" test precise (pgrep -x on the process
NAME) removed a guard the old pgrep -f held by accident: it matched every
process whose PATH contains .cargo/bin, which includes the Flight Deck dev
build app itself. So it refused to purge anything, always, for the wrong
reason — and the real hazard underneath was never named.

/build-dev and /build-app leave an app running from
target/release/bundle/macos/…, i.e. INSIDE the directory the purge step
erases. macOS keeps the open inode alive, so nothing crashes on the spot,
but a relaunch fails and anything loaded lazily is gone.

Hit for real while landing the review fixes: no build was running, the
guard was correctly silent, and target/release held the app in use.
target/debug (3.1 GB) was reclaimed, target/release deliberately was not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
clousty8 and others added 4 commits September 24, 2026 09:06
… c9bf1482)

Real incident (prod 2.5.0): the app's SSH key had been removed from a server's
authorized_keys. A remote conversation spun on "Reasoning… Meditating…" forever:
run_actor treated ssh's "Permission denied" as a network loss (one "reconnecting…"
notice, infinite retries), Settings said "could not reach the server", suggested
"Install Claude Code", and printed the raw ssh error for phone access.

- ssh_link.rs: one classifier for a transport that closed before fd_attach
  (exit 255 + stderr) → KeyRefused / HostKeyChanged (terminal) or Unreachable
  (transient). KeyRefused/HostKeyChanged end the session with a
  "Can't reach this server" notice naming the fix; a message in flight is flagged
  with the existing send_failed notice and busy is cleared — no more endless spinner.
  Server-controlled stderr is ANSI-stripped and capped before reaching the UI.
- SessionStatePayload.link (Connecting / Reconnecting{attempt}) drives the working
  indicator ("Connecting to the server…" / "Reconnecting…") instead of thinking verbs.
- tailscale.rs: "Tailscale looks off on this Mac" — fixed app-bundled CLI path,
  bounded 2 s, only for tailnet hosts, once per outage, never claimed on a guess.
- machine_diagnose keeps ssh's verdict: link_issue (key refused / host identity
  changed / unreachable) + tailscale_off_locally; headlines say which; no Claude-side
  repair is suggested when the server could not be reached.
- RepairAction::ReconnectMac: re-installs only this Mac's key with the server's
  login password (inline prompt, password never kept); honest error when the server
  refuses it. Host identity changes are never auto-trusted.
- Phone-provisioning and run_ssh_on_machine errors use the same sentences.
- Front: notice heading, working-indicator link text, ambient machine-health probe on
  a link edge (reuses Alexandre's RemoteRepoMark/ComposerBand/machineHealth), Settings
  card states + password prompt wording, mock scenarios for every new state.

Tests: classifier table, run_actor (key refused / host key changed terminal and never
retried, message flagged undelivered, unreachable keeps retrying, reconnect attempt
count), diagnose/repair, sanitizer, front surfaces. Live (Docker fixtures, local sshd,
127.0.0.1/.invalid, never a real server): the incident reproduced end to end, host key
change through the real transport, ReconnectMac restoring access, the real Tailscale
state — plus the whole bootstrap live suite (28/28).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three entries added for the review fixes just landed:
  - the schema is at v16, not v8 — v15 (remote_origin_url + probed_at) and
    v16 (remote_origin_note) were missing, and the list read as a total;
    plus scan_local_git_repos only scans LOCAL folders;
  - redact_remote_url alongside normalize_remote_url in git/mod.rs;
  - a new entry for the remote-origin sweep's contract: it returns
    RemoteOriginSweep, never a bool, and a firm non-url answer persists.

Two more came along that predate this work: the Codex default moving to
gpt-6-astra, and the derived FACTORY_HIDDEN_MODELS paragraph. Both were
already in the CRM context and had simply never been synced down.

Verified before applying: the regenerated file expands existing lines
only — no section lost or truncated — and the trailing newline the
generator drops was restored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Alex375
Alex375 merged commit eb7ffcd into main Sep 24, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants