Skip to content

Release v2.2.0 - #53

Merged
Alex375 merged 39 commits into
mainfrom
dev
Sep 14, 2026
Merged

Alex375 merged 39 commits into
mainfrom
dev

Conversation

@Alex375

@Alex375 Alex375 commented Sep 14, 2026

Copy link
Copy Markdown
Owner

Release v2.2.0. Voir les commits de dev depuis la dernière release.

🤖 Generated with Claude Code

Alex375 and others added 30 commits September 8, 2026 19:36
The status ladder added a fourth control to a footer that was already full, and
the buttons in it could be squeezed: a flex item shrinks by default, so at a
narrow panel width « Mark as Done » and « Open in TOSSE » wrapped onto two and
three lines and spilled out of their fixed 22px height.

Three changes, in order of what they fix:

- `.act` / `.open` no longer shrink and never wrap their label. This is the bug:
  a button now keeps its own width, whatever the row does.
- The way out to TOSSE becomes the icon alone (`compact`, as on a project card).
  Spelled out it was the widest thing in the footer (117px) and the least
  important — it was crowding out the very controls the panel exists for.
- The footer itself wraps. The panel is resizable down to 380px and the
  conversation's side region is narrower still, so there is a width at which the
  row cannot hold; there it breaks into a second line of whole buttons rather
  than a line of broken labels. The way out keeps to the right end of whichever
  line it lands on via its own auto margin, the `.spacer` before it having only
  ever pushed on one line.

Verified in the browser against the mock at the width that reproduced it: at a
417px panel the full « Approve & Done / Open / Discuss / Start ▾ / TOSSE » row
sits on one line; at the 380px minimum it wraps cleanly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… block

The fold body is mounted inside `.cv-collapse`, a grid used to animate the
disclosure height (0fr ↔ 1fr). Its child is therefore a grid item, whose
automatic minimum size is its MIN-CONTENT size on BOTH axes — and only
`min-height:0` was released. Any card in the fold with a wide min-content
(a diff, a code block, a table) grew the item past the reading column: the
card visibly stuck out to the right and the whole thread scrolled sideways.

Clean output is where this bit, because it is the mode that puts full cards
inside the fold; the same card outside is a flex item of `.cv-aibody`
(column), whose box stays capped.

Same hole on `.cv-motion-in`, the grid item of the live travel wrapper.

Measured in the mock (Playwright, 442px column): fold body 2007px → 442px,
thread scrollWidth 2075px → 552px = clientWidth. Same for a step detail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The footer still needed 401px to hold one line once a task had a conversation to
open (413px from « Review », whose ladder button is the longest) — and the panel
is draggable down to 380px. Below that it wrapped: correct, but two rows.

« Open » loses its word and keeps its bubble, so the footer now needs 364px —
374px with a count, 380px with a two-digit one. It fits at the panel's narrowest,
and the wrap stays underneath as the guard it was meant to be.

The count survives `compact`; only the word goes. It is the whole reason that
button differs from the single-conversation one, and reading it out of a tooltip
would mean hovering every task to find where the agents are.

Panel footer only. The task ROW keeps the word: it shows these buttons on hover
against a title that gives up its width first, so it has the room to spell things
out and no reason to be read as an icon puzzle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…itch

Flight Deck could only ever hold ONE Claude identity. It can now hold several,
each conversation says which one it runs on, and an opt-in policy moves work off
an account approaching its usage limit.

The isolation mechanism, verified live against claude 2.1.263:
CLAUDE_SECURESTORAGE_CONFIG_DIR scopes the credential store and NOTHING else.
CLAUDE_CONFIG_DIR — the obvious candidate — would have isolated transcripts,
settings.json, plugins, skills and MCP too, fragmenting the user's conversations.
With the secure-storage variable, `projectsDirectory` stays ~/.claude/projects,
so resume / fork / rewind are untouched by a switch. The Keychain service name
is derived exactly as the CLI derives it (sha256 of the NFC path, first 8 hex),
which is what lets us read a NON-ACTIVE account's usage — the only way to show
every account's rate limits from one panel. The default account deliberately
sets no variable at all: an empty string is not "unset", and would move the
credentials of anyone already scoping their CLI with CLAUDE_CONFIG_DIR.

- accounts/slot.rs: the isolation primitive, with a live probe test.
- usage/: fetch_plan_usage_for(slot) — per-account token + Keychain item.
- Sessions: handles are filtered BY ACCOUNT before get_usage. Asking another
  account's session returns a wrong-but-plausible figure the auto-switch acts on.
- SQLite v11: claude_accounts + conversations.claude_account_id (no FK; removing
  an account detaches its conversations rather than corrupting them).
- Settings → Accounts: one card per account, each with the SAME 5h/7d bars the
  context ring draws (PlanUsageBars, extracted so the two cannot drift).
- Composer: an account chip, separate from the model picker (what answers is not
  who pays). It costs one slot of the bar budget; it is hideable.
- Auto-switch: OFF by default, 90% trigger / 75% target ceiling, 10-min cooldown.
  Applied ONLY at a safe boundary — never mid-turn, never while background work
  runs — by ClaudeAccountApplyHost; until then the chip shows the account really
  in use. Every switch, and every armed-but-impossible switch, is stated in the
  thread.

A single-account setup spawns exactly as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… instructions block

A new backend-specific settings tab, shown only while a Claude account is connected:
which model each sub-agent runs on, what the helpers cost, and the CLAUDE.md policy
that lets Claude choose well.

## Routing (extensions/routing.rs, extensions/agent_edit.rs)

Three layers rather than one mechanism, because no single one does the job:
  1. baseline `CLAUDE_CODE_SUBAGENT_MODEL` — one value, no file, no frozen prompt;
  2. per-agent `model:`/`effort:` frontmatter, which beats the baseline (level 2 > 3);
  3. a definition FILE for `Explore` and `Plan` only — verified: the baseline env var
     does not reach those two, and that is the sole reason to duplicate anything.

`_FORCE` is offered as an explicit "lock the budget" checkbox, never on by default, and
visibly greys out the rows it overrides — which is exactly what it does to them.

Frontmatter writes rewrite ONLY the keys we own and preserve every other byte, including
the body (an agent's body is its system prompt). Round-tripped against every real agent
definition on disk, plus CRLF, BOM, block scalars, and a `---` inside the prose.

## Spend (agentspend/)

Read from the `agent-*.jsonl` the CLI already writes — no instrumentation. Measured on a
real 931-file / 194 MiB corpus: 24 066 assistant turns collapse to 50 buckets, so one
scan returns everything and the front pivots it for every table, filter and chart. A
substring prefilter before `serde_json` and a lean borrowed struct took the cold scan
from 2.3 s to 1.2 s (debug); a (size, mtime) cache makes a re-open ~10 ms. Verified
against a known-good tally: five of six models match to the token.

Costs use an EDITABLE rate card (localStorage), never buried constants, and the UI says
in place that these are API list prices, not what a subscription is billed.

## Drift canary

The resolution order changed once already (2.1.251) and a built-in's name is an
undocumented contract — either can break an override with no error anywhere. So the
dashboard checks intent against reality: an agent configured for one model but seen
running another raises a banner. Compares normalized ids, so `haiku` vs
`claude-haiku-4-5-20251001` never cries wolf.

## Instructions (memoryfile/)

Writes only between its own markers in `~/.claude/CLAUDE.md`; damaged markers are
refused rather than repaired. Ships the routing-policy block — the one lever that makes
the dosing intelligent rather than merely uniform, since Claude's own guidance otherwise
forbids downgrading a worker for looking easy.

Scope warnings are COMPUTED (`git check-ignore`, worktree detection), not hardcoded, and
"could not check" is never rendered as "fine".

Charts are inline SVG (no new dependency); the categorical palette was validated with
the dataviz validator against this app's own surface — all five checks pass.

636 Rust tests, 1609 front tests, tsc clean. Bindings regenerated.
…ings sub-tabs + search

The voice agent wrapped every answer ("C'est bon, j'ai lancé la conversation
dans le repo, je reviens vers toi tout de suite"). A sentence budget was never
going to fix that — it just fits filler inside the budget — so the brief now
bans the filler shapes by name and gives the skeleton of the two lines it says
most. The announcement lines repeat the rule: they are the last thing in context
before the agent speaks, which is exactly where the padding crept back in.

- src/voice/instructions.ts: the default brief + `resolveInstructions`
  (empty override = the default, so clearing the box IS the reset).
- The brief is EDITABLE from Settings and applies to a live session
  (instructions, unlike the voice, can change mid-session).
- Voice picker: the catalogue and the default live Rust-side (`voice/mod.rs`),
  the choice travels to `voice_agent_client_secret` and is sanitized there — a
  stale key degrades to the default instead of 400ing the whole session. The
  voice is fixed at mint time, so an idle armed session is re-armed at once and
  anything busier waits for the next one; Settings says which happened,
  including when the re-arm itself failed.
- Settings: sub-tabs for the two tabs that had become a scroll of unrelated
  cards (General, MCP Control), the voice card split into Voice agent /
  Microphone / Wake word, and a search box over a flat index — picking a result
  lands on the right tab AND sub-tab, then flashes the row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…record false triggers

The wake word fires on background noise, onomatopoeia and silence. This is the
instrumentation + latent-bug pass; the detection gates (patience, VAD veto, RMS
floor) come next, on top of a VAD that now works.

**The Silero VAD never once reported speech.** It needs a 576-sample tensor —
64 samples of context carried from the previous hop, prepended to the 512-sample
hop, which silero-vad 5.x does inside its own wrapper. We fed a bare 512. The
exported graph's sequence dimension is DYNAMIC, so onnxruntime ACCEPTS that and
silently returns ~0: measured on 7 s of loud, clear speech (RMS 0.12), peak
probability 0.003 at 512 versus 1.000 at 576. Nothing errored, nothing logged —
the VAD simply reported silence over speech, forever, which is indistinguishable
from a VAD that is switched off. It gates nothing today (see `ccf654d`), so this
changed no behaviour; it is what the planned fire-time veto has to stand on.
Guarded by a committed 0.6 s speech fixture, so CI checks the contract without a
microphone.

**Duplicate embeddings.** `Engine::feed` sliced the audio buffer's TAIL for every
step, so a callback carrying several 80 ms chunks handed each of them the SAME
audio — stuffing the rolling window with repeated embeddings and corrupting the
temporal pattern the classifier is trained on. Rare with CoreAudio's 512-frame
buffers, systematic with a larger one. `feed` now walks a per-step window end
through the buffer; a test asserts the score trajectory is identical whether the
audio arrives in 10 ms or 320 ms callbacks.

**Non-finite features.** A muted or dead input device delivers exact zeros and a
log-mel of zeros can come back -inf/NaN. That does not merely make the classifier
wrong — its output becomes meaningless and can read as a confident hit, which is
one way "it fires on silence" happens. Those steps are now dropped, with one
latched log line rather than one every 80 ms.

**Recording false triggers** (opt-in, off by default, Settings → Control). Every
fire writes a WAV of the ~4 s around it plus a JSON score trajectory carrying the
per-step VAD probability and window RMS. The trajectory says WHICH failure a
trigger is — a lone spike a "N consecutive steps" rule would swallow, a slow
climb wanting a higher threshold, or a fire at rms≈0 that no threshold can fix —
and the clips are both the hard negatives a retrain needs and replayable through
the real engine via `detect_wav`. It writes microphone audio to disk, so: opt-in
only, capped at 40 pairs, never sent anywhere. The capture directory is created
when the switch goes on, not at the first fire, so a folder it cannot write turns
the switch back off and says why instead of silently keeping no evidence.

Also: the score log line now carries vad and rms, and
`no_bundled_phrase_fires_on_silence_or_noise` covers every bundled phrase rather
than `alexa` alone (the custom `ground_control` classifier never saw room tone or
broadband noise in training).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
## The canary was accusing the user of its own success

Setting Code search to Haiku made the drift banner fire immediately: it compared
today's setting against a week of transcripts that had all run under the PREVIOUS
setting. A warning at the exact moment you change something is how a canary loses its
credibility.

`AgentRouting` now carries `configured_at_ms` — the mtime of whatever file holds the
setting (the agent definition, or settings.json for a baseline-driven row). No new
bookkeeping: the file that holds the setting already records when it was written. The
canary only weighs turns that ran strictly AFTER that day, and says "since you changed
this on <date>" so the window is visible.

## Tokens lead, money trails

A dollar figure in the first tile made a subscription read as a four-figure bill, and
the disclaimer underneath could not undo the impression the big number had made. The
tiles are now Output tokens · Turns · In workflow runs · At API rates ("not your bill"),
and the daily chart plots tokens.

Dropped the "same volume on Sonnet" counterfactual: comparing token volumes across
models is hard to read and easy to misread as a forecast.

## Bars, not mountains

The daily chart was an area chart, which reads as a silhouette — the eye follows the
outline instead of comparing days, and placing yourself meant counting along the axis.
Now discrete stacked bars, with a weekday initial under each and the date on Mondays,
and `Mon 7 Sep` as the tooltip heading.

The legend became a filter: clicking a model hides it and the chart rescales to what is
left, which is the only way to see a model with a hundredth of the spend. The last
series can't be switched off — an empty chart just looks broken.

## Tooltips no longer leave the screen

Hovering the right-hand side of a chart pushed the tooltip off the window. Placement is
now edge-aware and flips to the left of the pointer, shared by both charts.

## Smaller things

- Model suggestions name the TIER, not the release: "Suggested: Haiku", not "Haiku 4.5".
  Same in the routing policy written into CLAUDE.md, which outlives releases.
- The checkbox is drawn by hand instead of the near-white browser default, which shouted
  louder than the setting it controls.
- Instruction blocks are framed as a list ("Instructions we suggest adding") rather than
  one button, since the list is meant to grow.
- Unsaved instruction text is tinted like a diff's added lines, with a notice — the box
  used to look identical whether the text was live in the file or merely proposed.

1612 front tests, 636 Rust tests, tsc clean. Bindings regenerated. Verified in the
browser against the mock: canary window, legend toggle, tooltip flip, unsaved tint.
Throwaway integration branch so both feature branches stay clean: neither
`/land` should carry the other's work.

Two conflicts, both in Settings, and both caused by `dev` having moved on since
voice-brevity branched (it predates the new Behavior tab).

`SettingsPanel.tsx` — voice-brevity adds sub-tabs and a settings search; dev had
since added a Behavior tab and moved `PermissionPrefs` into it, out of General.
Kept both: the sub-tab structure stands, but `general/system` renders only
`CaffeinatePrefs`. Taking voice-brevity's block verbatim would have rendered the
bypass switch in TWO tabs at once. The Behavior and Models sections are gated on
`!searching` to match the search feature's own convention.

`settingsSearch.ts` — the index still pointed "bypass" at `general/system`, which
after that move lands on a page the switch is no longer on (the index test only
checks that a sub-tab EXISTS, not that the setting is there, so this was green
and wrong). Repointed to `behavior`, and indexed Output style, which had no entry
at all.

`VoiceAgentSection.tsx` — structural only: voice-brevity rewraps the component,
so the "Record false triggers" row just needed reseating inside the new
`SettingsGroup`. Both features intact.

cargo test --lib 596 · tsc --noEmit clean · vitest 1600 · bindings regenerate to
no diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four things from the second pass.

## The per-folder legend filters too

The daily chart's legend toggles models; this one did not, for no reason other than that
it was built first. Same behaviour now, and the bars re-sort and rescale to what is left.

## Long folder names no longer slide under the bars

`web_dentiste_middleware_api` ran straight into the chart. Widening the column is not a
fix — names have no common length, so sizing to the worst case starves every other row.

The chart is now laid out in HTML rather than SVG, which makes the actual fixes one CSS
property each: a hairline divider marks where the label column ends, the name fades out
as it reaches it (rather than an ellipsis, because the point is that it continues), and
hovering shows the full name plus the folder's absolute path, wrapped so a deep path
cannot push the tooltip off the window. Nothing in this chart needed SVG's coordinate
space — the bars are proportional widths.

## Table headers sat over the wrong columns

`.cc-table th` (specificity 0,1,1) was beating `.cc-num` (0,1,0), so the headers stayed
left-aligned above right-aligned figures and every number read as belonging to the column
beside it. Scoped the rule to `.cc-table .cc-num`.

## The instructions box shows a real diff

Tinting the whole box green said "something changed" but not WHAT — and the thing worth
seeing before writing to a file you also edit by hand is the lines that would DISAPPEAR.

New `lineDiff.ts` (plain LCS over lines; the managed block is instructions, not a
repository, so the O(n·m) table is free and anything cleverer would only cost
readability). Additions in green, removals kept on screen in red until you save them
away, unchanged lines in grey, long unchanged runs elided to two lines of context, and a
`+n −n` summary. The textarea keeps only a quiet left border now — tinting the editing
surface made every character look like an addition.

The demo fixture gained a long folder name and an existing managed block, since an empty
block only ever produces green and hides half of what the diff is for.

1624 front tests (12 new on the diff), 636 Rust tests, tsc clean. Verified in the browser:
legend filter, label fade and path tooltip, header alignment, red removals.
… a reply

Brevity wasn't the whole ask: a SHORT pleasantry is still a pleasantry. In text
you skim past an intro; spoken, every syllable is time you sit through, so
anything that isn't information is noise. The agent is a radio operator now —
it speaks for four reasons (report an event, answer what was asked, report an
action's outcome, ask the one thing it needs) and says nothing otherwise.
Greetings and sign-offs get silence, not a shorter pleasantry.

The prompt could only do half of it: `runToolCall` answered end_call with a
`response.create` and closed the mic only once that goodbye had finished
playing — the app was COMMISSIONING the « à la prochaine » it complained about,
and holding the microphone open for it. end_call now returns the tool result
without requesting a turn and closes the mic at once, which also retires
`endPending` and its 15 s fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pure silence has one ambiguity: "order taken, working on it" and "I did not
hear you" sound exactly alike. Radio solved this long ago, so the brief now
allows a fifth reason to speak — acknowledging an order there is nothing to
report on yet — as a FIXED call sign (« Bien reçu. » / "Roger."), not a licence
to improvise a polite sentence. It never prefixes a real report: the outcome
already proves the agent heard, and « Bien reçu, message envoyé » is exactly the
padding this brief exists to remove.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…n 25 recorded triggers

The instrumentation paid for itself immediately. 25 false triggers recorded across
silence, background noise, music and speech, all on `ground_control`:

- 24 of 25 were a SINGLE 80 ms spike out of nowhere — the score jumped from below
  0.5 straight past the threshold (median jump +0.53) and back.
- 23 of 25 fired with essentially no speech in the window (Silero peak over the
  trailing 1.3 s below 0.2, median 0.007).
- 20 of 25 fired at an RMS under 0.01, i.e. a quiet room.

Against that, a real utterance holds the threshold for 12-18 consecutive steps
(the ~2 s classifier window slides through the phrase slowly) with a VAD peak of
0.62-1.00. The separation is 18 versus 1 — not a threshold that needs tuning,
a missing confirmation step. Raising the threshold instead would have been the
wrong lever: even 0.9 leaves 4 of these 25 and would gut real recall.

So, two gates before a fire:
- **patience**: 3 consecutive steps above threshold (openWakeWord's own parameter,
  which this engine never had). ~160 ms of added latency.
- **confirmed speech**: Silero's peak over the trailing 16 steps must reach 0.5.
  ⚠️ Over the WINDOW, never the firing step alone — the classifier fires as the
  phrase ENDS, on audio already gone quiet (measured: vad≈0.01 at the firing step
  of a real utterance whose phrase steps read ~1.0). A per-step veto would reject
  every genuine detection. This gate only became possible now that Silero actually
  returns a probability.

Replayed through the new engine: 24 of the 25 recorded triggers are suppressed,
and 4 synthesized "Ground Control" utterances across 4 voices all still fire at
0.999-1.000. The one survivor is the single capture that contained real speech
(VAD 0.999, RMS 0.05) — it needs a human ear to say whether it was a genuine
wake or a false positive on speech, which no amount of analysis can settle.

Debug capture now also records SUPPRESSED candidates, named `-blocked-<gate>` and
tagged in the JSON. Without them a quiet capture folder cannot be told from gates
that are swallowing real detections — and that is the failure mode this change
introduces. Suppressed candidates are only produced while capture is on, so the
normal path allocates nothing extra, and they never reach the app.

Also adds `replay_corpus` (ignored): replays a whole directory of WAVs and reports
fired/suppressed per file. This is how the numbers above were produced, and how
the next gate change gets checked against the same corpus.

No new setting: firing on room tone is a defect, not a behaviour anyone chose.
The added latency and the recall trade-off are worth a look before this lands.

cargo test --lib 597 · tsc --noEmit clean · vitest 1600.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…merged

The panel had grown to 12 top-level tabs with a wall of switches in
General → Display, and alert settings split across two places.

- General → Display: one card of 11 toggles becomes three — Appearance
  (interface zoom, live workflow on the card), Thread (clean output,
  task notifications, last-message preview, minimap, message controls,
  clickable filenames) and Motion (the three animations, each of which
  the system's "reduce motion" already overrides).
- Conversation absorbs Models and Composer behind sub-tabs
  (Markdown / Models / Composer): all three shape what a conversation
  is, and the drag surfaces keep their own sub-page.
- Notifications absorbs the old General → Alerts behind sub-tabs
  (Channels / Fleet / Background): the fleet readout and the
  background-task alert sat in a different tab from the OS channels
  they belong with. General keeps Display / Durations / System.

The four merged sections take an `embedded` prop that drops their own
PageHead, since the tab now carries it (same shape as the Control
sub-groups). The search index follows: SETTINGS_SUBS gains the two new
sub-tab sets and every moved entry gets its section/sub/group, so a hit
still lands on a tab that renders it. Rail goes from 12 tabs to 10.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Settings panel: reconciled the sub-tab/search restructure with the
Behavior + Claude Code tabs dev landed meanwhile.

- TABS keeps dev's "Behavior" and "Claude Code" (with its needsClaude
  gating and the claudeSeen latch); "Models" and "Composer" are gone as
  top-level tabs because they are now Conversation sub-tabs.
- General → System no longer renders PermissionPrefs: dev moved bypass
  permissions to the Behavior tab, which is the better home. System
  keeps Caffeinate.
- dev's new "Per-agent detail in the workflow view" toggle moved into
  the Appearance card, right next to the live-workflow toggle it
  belongs with.
- Behavior and Claude Code render behind the same `!searching` guard as
  every other branch, so search results don't render under a tab pane.
- Search index follows dev: bypass permissions re-pointed at "behavior",
  plus new entries for Output style, the Claude Code groups and the
  per-agent workflow toggle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`Alex375/tosse-code` was renamed to `Alex375/FlightDeck`. It is a pure rename —
same repository (`id 1271050927`), and the old name still redirects — so nothing
is broken today. The redirect is the problem: it dies the moment anyone creates a
new repository under the old name.

## The one with teeth

`tauri.conf.json`'s updater endpoint is polled by every installed app. It was
resolving through a 301 rename redirect; if that redirect ever lapses, installs
stop receiving updates silently — no error, no banner, just an app that never
updates again. Verified before: 301 → 302 → 302 → 200. After: only GitHub's own
asset 302s, no rename hop.

## The shop window

The repository is public, so the README's release link and clone URL are how
people arrive. The clone snippet also said `cd tosse-code`, which is now the
wrong directory — a clone of `FlightDeck` creates `FlightDeck/`.

## Deliberately untouched

- The ~12 `Alex375/tosse-code` strings in `git/mod.rs` and `tosse/mod.rs` are all
  inside `#[cfg(test)]`. They are arbitrary URLs exercising `normalize_remote_url`;
  the repository name is incidental to what they test. Changing them buys nothing
  and risks the assertions.
- `CLAUDE.md:103` still names the old repo, and is left alone ON PURPOSE: that line
  lives inside the `[GENERATED] Repository Context` block, which `sync_claude_md()`
  regenerates from the TOSSE CRM. A hand-edit there is reverted at the next release
  — which is exactly what happened to the governance section (`b3dcb6e` corrected it
  on 4 Sept, `e8c77e8` overwrote it on 7 Sept). The CRM context is the source of
  truth; that is where this gets fixed.

Out of band: the CRM repository record's `url` was corrected to the new name. That
one is functional rather than cosmetic — folder↔repo association runs through
`normalize_remote_url` on the git remote, so changing the local `origin` without
changing the CRM would have broken the link.

1624 front tests, 636 Rust tests, tsc clean.
…ganisation

`feat/voice-brevity` is on dev now, so the throwaway integration branch became the
feature branch (renamed `feat/wake-detection-gates`) and follows dev instead.

Both conflicts resolved to dev's side: it had since done the same reorganisation
properly and further. Alerts moved out of General into Notifications, Models and
Composer became Conversation sub-tabs, Claude Code gained its own tab — our older
shape is superseded, not complementary. Same for the search index: dev had already
repointed the Behavior entries, so our version of that fix is redundant.

cargo test --lib 653 · tsc --noEmit clean · vitest 1641.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hreshold

The threshold slider felt inert, and the meter explained why: normal speech
(RMS ≈ 0.05) drew 16% of a bar whose threshold handle sat at 60%. The level could
not reach the handle no matter how loud you spoke, so the instruction printed
under it — "drop the handle just above your noise floor" — drove the threshold to
its minimum if followed.

Rescaling the bar would only have made the fiction plausible. The two numbers do
not share a scale at all: `server_vad.threshold` is a CONFIDENCE from OpenAI's
own speech detector, not a loudness, so no local level bar can be calibrated
against it. A bar that lines up by arithmetic while meaning something else is
worse than no bar, because it looks authoritative.

So each control now claims only what it can prove:
- The threshold is its own slider, described as what it is — how sure their
  detector has to be — with the honest instruction to find it by ear.
- The meter is a microphone CHECK: is the right input selected, is it hearing me.
  On a dBFS scale (-60..0), which puts room tone near 15% and speech near 55%
  instead of everything squashed into the bottom sixth.

Turn detection stays OpenAI's job. Gating the audio ourselves before it leaves
the machine was considered and dropped: it buys a threshold we could draw on a
bar, at the cost of a second detector to keep honest, a delay line so it does not
clip the first syllable, and hysteresis so it does not chop a sentence in two.

Both mic call sites now share `mic.ts`, so what the meter shows is what the agent
hears — they had drifted apart by accident, which is exactly how a diagnostic
starts lying. The meter also reads back what the track actually negotiated
(`getSettings`), because a constraint the platform ignored should be visible, not
assumed.

⚠️ Left deliberately unchanged, and written down in `mic.ts`: auto gain control is
ON. Its job is to make loud and quiet input come out at the same level — in a
quiet room it winds gain up until the noise floor reads like speech — so every
threshold downstream, OpenAI's included, is working on a deliberately flattened
signal. That is the leading suspect for the remaining over-sensitivity, and it is
a one-word change in one place now. It is also a real trade-off (quiet speech
gets quieter), so it wants a measurement rather than a guess.

tsc --noEmit clean · vitest 1648.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…dates was evicting the fires

First session with the gates on wrote 38 blocked candidates to 2 real fires. With
one shared cap of 40, that ratio prunes away exactly the evidence worth keeping:
the fires — rare, and the only captures that show something actually went wrong —
were gone within minutes, crowded out by the plentiful ones.

Fires and blocked candidates now have independent budgets (40 and 60), so a busy
stretch of blocked candidates can never cost a fire. Test asserts it directly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ces and a confirmed false positive

Armand recorded ten deliberate "Ground Control" utterances and confirmed which
earlier capture was a false positive, which is the first time this had real
evidence on both sides rather than false positives alone.

What the recordings show: the FIRST step of a real detection is a ramp-in — the
classifier window is only part-way into the phrase, so it scores 0.66-0.82 before
pinning at ~1.00 for the rest. So demanding three consecutive steps over a high
bar fails on the ramp, not on the phrase. At threshold 0.90, patience 3 lost 4 of
10 genuine utterances; patience 2 lost none. Two steps over a high bar beats three
over a low one, and it is 80 ms FASTER than what it replaces.

The threshold band was aimed at the wrong place entirely. Scores here are bimodal
— near 0 or near 1 — so nothing interesting happens below ~0.8, yet the slider
spent its whole upper half between 0.30 and 0.60. And it topped out at 0.90, so
the one region that separates real from false was not reachable at any setting.
The band is now 0.82-0.98, default 0.90.

Replayed through the engine:
- 10/10 real utterances fire.
- 25/25 confirmed false positives suppressed, including the one Armand identified
  (it peaks at 0.92 and touches it exactly once; real utterances hold 0.90 for at
  least two consecutive steps).
- The capture that fires out of the "false" folder is the one he confirmed as a
  genuine wake — correct, not a miss.

The one recording lost at this setting turned out to be a deliberately distorted
phrase, so rejecting it is the wanted behaviour rather than the cost it looked
like.

⚠️ This is tuned on ONE confirmed false positive, which places a boundary inside a
0.01 gap. It holds on the evidence we have and no further. The durable fix is
retraining a classifier that has never heard background noise or a real human
voice — TOSSE task 64a42f06, with every capture corpus listed.

cargo test --lib 655 · tsc --noEmit clean · vitest 1648.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…a sound

Armand set the threshold to its maximum, 0.90, and talking while emptying a
dishwasher still cut the agent off every ten seconds. Worse, the agent ANSWERED
the noise: asked "shall I grant permission X?", it heard a clatter and replied
that it was granting. That is a consequential action taken on crockery, and it
made the feature unusable.

No threshold fixes that, because loudness is not the problem. A plate on a
counter IS loud; `server_vad` is an amplitude gate with no way to know the sound
carries no words, and a threshold high enough to exclude dishes excludes the user
too. The slider was not failing to work — it was working, on the wrong question.

`semantic_vad` asks a different one: does what was said sound FINISHED. Audio with
no words in it is not a turn at all. It is now the default, with `eagerness: low`
so a pause mid-sentence is not treated as the end of one. The loudness gate stays
available as a mode, since it is the one that offers a number to turn, and
Settings now says plainly what each is and when the choice matters.

Two more layers, because turn detection alone should not be the only thing
standing between a noise and an irreversible action:

- The brief gains a rule that outranks everything else in it: a turn you could
  not make out is not a turn — do nothing, say nothing, and never infer an answer
  from a sound. "Yes", "go ahead" and "delete it" must come from a sentence the
  agent actually heard. It had no such rule, and a model asked a question will
  find an answer in whatever comes back.

- Server `error` events reached `console.error` and nowhere else. A `session.update`
  the server rejects is reported there and ONLY there, so turn-detection settings
  could be silently discarded while Settings showed them as applied — someone
  turning a slider the server threw away had no way to find out. They now raise
  the app error banner. (Checked against the current API docs: our `server_vad`
  shape is valid, so this was not the cause here. It could have been, invisibly.)

Stored mode/eagerness are coerced on read and on write: an unknown enum from an
older build would be rejected by the server as a whole-payload error, taking the
rest of the session config down with it.

tsc --noEmit clean · vitest 1649.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e MCP Control to Control

Semantic turn detection did not fix it and probably made it worse: the agent now
cut out mid-sentence roughly every thirty seconds, with no relation to how loud
the room was.

That "no relation to the noise" is the tell, and it points away from the room and
at the agent itself. Its voice comes out of the speakers and back into the open
microphone. Echo cancellation is on but imperfect on laptop speakers — and where
an amplitude gate at 0.90 might reject a quiet echo, a SEMANTIC detector is
looking for words, which is precisely what an echo of speech is made of. Asking
"did someone just finish a sentence" of the agent's own sentence gets a yes.
Switching to semantics sharpened the wrong instrument.

Armand's verdict settles the default regardless of cause: an agent that never
stops beats one that stops constantly. A long answer you must sit through costs
seconds; an interruption costs the whole exchange. So barge-in becomes a choice,
defaulting to off:

- **Let it finish** (new default) — nothing stops it mid-sentence. Speaking over
  it still WORKS: `create_response` stays on, so the turn is answered once the
  agent is done rather than over the top of it. Nothing is lost but the cutting.
- **With the wake word** — his own suggestion, and the best of the three. The
  wake detector is a far stricter judge than any turn detector: a specific
  phrase, two consecutive steps over 0.90, vetoed unless Silero agrees it was
  speech. It does not fire on a room. The app sends the cancel itself rather than
  handing the server barge-in, so the strict judge is the only judge.
- **By speaking** — the old behaviour, kept and labelled with what it costs.

Interrupting needs BOTH halves to look like it worked: `response.cancel` stops the
model generating, `output_audio_buffer.clear` drops the audio already sitting in
the WebRTC playout buffer. Cancel alone leaves it talking through what it had
queued.

`shouldFireWake` gains the one narrow exception to its self-trigger guard: in wake
mode the phrase MUST be heard while the agent is speaking, since that is exactly
when you need it. It still refuses a session that is off, unconfigured or broken.

Also renames the MCP Control tab to Control, as asked.

tsc --noEmit clean · vitest 1652.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
37 confirmed findings, ~15 root causes. The two critical ones broke the feature's
headline capabilities for every added account:

- Keychain item name was derived in the WRONG ORDER. The CLI calls `Sx("-credentials")`,
  so an isolated slot's service is `Claude Code-credentials-<sha8>`, not
  `Claude Code-<sha8>-credentials`. Invisible on a single-account machine (empty
  suffix), fatal for every added account: its usage could never be read, so per-account
  limits and auto-switch were dead. The old unit test only asserted the shape the code
  chose; it is now pinned to a literal digest, plus a live probe that shims `security`
  and asserts against the service the REAL CLI queries (verified: it asks for
  `Claude Code-credentials-60a207a7`, exactly what we derive).
- The pasted OAuth code was not bound to its account. One global in-flight login, one
  card per account: a code typed into a superseded card was redeemed into another
  account's store and stamped the wrong identity. The login now records its account,
  submission verifies it, and a superseded card closes its code box.

Data integrity:
- `conversations.claude_account_id` is now written on INSERT only; the dedicated
  `set_conversation_claude_account` is its sole updater. A stale in-memory copy
  re-upserted on any activity bump was resurrecting ids a removal had just detached.
  The front also mirrors the detach and clears a removed default.
- Removing an account ABORTS when `claude auth logout` fails: the Keychain item name is
  derived from the directory being deleted, so proceeding orphaned live, unrevokable
  tokens. "Remove anyway" names the exact item to revoke by hand. Refused while a
  session runs on it.
- Placeholder labels are a persisted fact (`label_is_generated`), not a prefix guess
  that clobbered names like "Account manager".

Silent failures now surfaced: unreadable CURRENT account (auto-switch was a silent
no-op), frozen usage after a terminal error, identity-capture failure, rename failure,
failed account-list read, unknown account typed as `UnknownAccount` instead of a
network error retried forever. The default-account preference no longer drops when
the list is unloaded (seeded at boot; unloaded ≠ empty). Remote (SSH) conversations
refuse an account the launcher cannot carry. The orphaned-account chip stays visible.
Threshold inputs commit on blur, reject out-of-range values, and the ceiling has its
own control; the hysteresis invariant now holds on every clamp path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Multiple Claude accounts with manual selection, per-account rate limits in
Settings → Accounts, and opt-in auto-switch near the usage limit.

Conflict: src/ipc/bindings.ts (generated) — both sides added exported types at
the same alphabetical position (dev: SpendBucket/SpendReport/SubagentBaseline/
SubagentRouting; feature: SpawnFlags). Resolved by regenerating from the merged
Rust source rather than hand-editing, so every type from both sides is present.

Integration: dev's new Settings search indexed the Accounts tab with a single
generic entry, so none of the feature's new rows (auto-switch, default account,
thresholds, per-account limits) were findable. Added them to SETTINGS_INDEX.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…g band

The page had grown into one ~340px ConnectionCard per account (gradient tile, coral
glow and border, provider line repeated, "Connected" shown twice, monospace name
field), so two accounts filled the screen before any setting. Chosen from three
mocked directions: tiles for the accounts, the switching band and the numbered
sign-in from the compact-list direction.

- Claude accounts: a two-column grid of tiles, each with a ring gauge per usage window
  (5h / 7d, amber past 80% like the context ring) and model-scoped caps on a quiet
  line. Rename / Make default / Sign out / Remove move into a portalled ⋯ menu; sign
  out and remove confirm first. A dashed tile adds an account.
- Sign-in: numbered steps inside the tile. Step one is only marked done once the
  browser really opened; if it failed, the manual link fallback shows instead.
- Switching: default account as a segmented control, the auto-switch toggle (disabled
  with its reason while there is nowhere to switch to), and a 0–100% band — the
  hysteresis gap hatched, amber past the trigger — with every account placed at its
  busiest window. Accounts closer than 12 points share one marker (their labels were
  printing over each other); the merged marker sits at its busiest member.
  Thresholds are steppers that commit on blur/Enter and never ratchet the ceiling.
- Codex: the same tile language.

No behaviour dropped: the account-bound login, superseded-login notice, refused
removal with "Remove anyway", identity-capture warning, typed usage errors, and the
shared-profile-cache guard on the default tile all carry over. ConnectionCard is
untouched for the TOSSE tab. The settings search index follows the new titles, and
the mock now starts added accounts signed out so the sign-in flow is exercisable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t after every tool call

Armand changed a setting mid-conversation and the banner read: "The voice agent
rejected a setting: Conversation already has an active response in progress:
resp_ENz8mbt061hUCNaoQzB5k. Wait until the response is finished before creating
a new one."

That line was wrong twice, and it hid a real bug.

Wrong label: the error came from a `response.create`. A `session.update` never
creates a response, so no setting was refused. The previous commit had pushed
EVERY server error into the banner under that one prefix, whatever caused it.

Wrong audience: a response id means nothing to a person. Zero silent errors does
not mean dumping the server's wording on screen.

The real bug underneath: `runToolCall` fired `response.create` the moment a tool
finished. A tool call is reported by `response.output_item.done` while the response
CONTAINING it is still open, and app-control tools run locally in milliseconds, so
the request routinely arrived first and was refused. That refusal was the only
trace: the agent carried out the action and never said so. It now waits for the
server to be quiet first (the same bounded wait announcements already use). The
tool output goes in BEFORE the wait, so any response that starts in the gap still
carries it.

Errors are now sorted by cause, which the API reports exactly: every client event
can carry an `event_id` and the error echoes it back (checked against the current
client-events reference).
- One of our tagged setting updates was refused → a plain note inside the
  Microphone card: the change applies the next time the agent starts. That is
  true, because a new session reads the stored preferences on connect.
- The active-response race → absorbed. After the wait above, it can only happen
  when the server starts its own response from the user speaking, and that
  response already carries the tool output. Logged, not shown.
- Anything else → a plain sentence in the banner. The server text stays in the log.

No "applies on the next start" notice appears on every change, as was briefly
considered: the same reference says `session.update` applies live for every field
except voice and model, and voice already has its own note. Showing that notice on
every change would be false.

The classifier is a pure module (`serverErrors.ts`) with tests, including one that
keeps ids and server wording out of anything user-facing.

tsc --noEmit clean · vitest 1657.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…atus

The attention tint (input/review/error) replaced the .on background, so the
open conversation was indistinguishable from other rows with the same status.
The selected row now gets a deeper wash of its status colour plus a thin ring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
clousty8 and others added 9 commits September 14, 2026 13:41
An account is its email address — two accounts cannot share one — so the tiles,
the default-account selector and the switching band all show the address. The
generated "Account 3" label only stands in until an address is known, and the
rename action is gone with the name it edited (`claude_account_rename` removed
from the IPC surface rather than left dead).

The default account needed real work to say this honestly: its address cannot be
read live once a second account exists, because `claude auth status` answers from
a profile cache every account SHARES and would name whichever signed in last —
which is why the old card hid it. Its identity is now captured at its OWN sign-in
and stored in `meta` (`claude_default_identity`), the same treatment the added
accounts already got in their row. Absent that capture (signed in outside the app)
the tile keeps the "Claude" fallback and says why, instead of showing an address
that may belong to another account.

Also: Codex is named by its address too (its status is its own, never shared), and
a marker near either end of the band anchors its label to that edge — a full
address centred on a marker at 8% hung outside the card.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each account's email is read from GET /api/oauth/profile with that account's
own token (claude_account_identity), instead of the CONFIG-dir profile cache
the CLI shares across accounts. The address now names accounts in the
Settings tiles, the default-account selector, the threshold band, the
composer account chip and its menu, and the auto-switch notices.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…"Other"

A pending AskUserQuestion is itself a can_use_tool, so the voice agent
announced it as "blocked on a permission prompt for the AskUserQuestion
tool" and relayed dictated answers with send_message, which only queued
behind the open question.

- Classify attention on the tool, not on "anything pending" (shared
  attentionFields for the OS ping, the journal and the announcement), and
  carry the question text.
- Expose get_pending_request + answer_request to the voice agent.
- answer_request: questions need no opt-in; the executor builds
  updated_input.answers from a loose `answers` payload (free text lands as
  the "Other" choice) and refuses answers that match no question. Real
  permission prompts stay behind Settings -> Control.
- Extract the framework-free questionnaire helpers into questionnaire.ts.
- Voice brief: questions are not permissions, never answer with send_message.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ts search

- Search no longer offers results in a hidden tab: picking a Claude Code result while no
  Claude account is signed in used to bounce the user back to General.
- The Claude Code tab now leaves the rail after a successful "signed out" read; the
  latch only survives a FAILED read, as it was meant to.
- A search highlight no row claims (a page heading, a tile, a label) disarms itself after
  2.5 s instead of flashing that row out of the blue whenever it appears later. Results that
  name a label inside a card ("Switch at", "Target below") flash the card via a new `flash`
  redirect, cross-checked by a unit test.
- git: `path_is_ignored` had been inserted between `normalize_remote_url`'s doc block and
  its `fn`, stealing its documentation. Moved above it; pure move.
- CHANGELOG v2.2.0: add the voice turn-detection change and the clean-output width fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t and the Claude Code tab

Multi-account
- Remote (SSH) conversations: always seeded on the default account (create, worktree,
  reopen, fork), refused a non-default one in the store, skipped by the auto-switch, and
  excluded from `claude_handles_for` so the default account's ring never shows the SERVER
  account's quota.
- Account restart only at a real boundary: not busy, no `activity`, no queued message, no
  pending permission, no background task, no spawn in flight — then a 3 s settle and a
  re-check against fresh state. `!busy` alone killed work the CLI takes back after a result.
- Auto-switch: transient read failures keep the last good figure; terminal ones are
  reported once per ACCOUNT into live conversations only; nothing is evaluated until every
  probe has answered once; only live conversations are moved; a user pick pins the
  conversation against the policy; told markers are voided by any account change.
- New non-terminal `UsageError::TokenExpired`: an idle added account's lapsed access token
  no longer reads as a revoked sign-in that stops polling forever.
- Account removal refuses while a sign-in for it is awaiting or redeeming its code, and an
  account-lifecycle RwLock closes the check-then-delete window against spawn / sign-in /
  identity capture.
- The default slot removes an inherited CLAUDE_SECURESTORAGE_CONFIG_DIR; a corrupt stored
  default identity is an error shown in Settings, not a silent null; the usage ring follows
  the account the live process actually runs on.
- Docs: `claude auth login/logout` DO rewrite the shared profile cache in ~/.claude.json
  (verified in the 2.1.270 bundle); the CLI refetches it at the next session start.

Claude Code tab
- Spend: fold the CLI's per-content-block lines by (message.id, requestId), keeping the
  final output count — totals were inflated ~4.5x.
- The budget lock follows the baseline (and clears with it); the forced model is shown.
- Effort edits send the file's own `model:` (new `defined_model`), never the baseline.
- Plugin agents are read-only with a visible reason; drift ignores locked and `inherit`
  rows and is worded as an observation; `settings.json` read failures surface as
  `SubagentBaseline.error`; unpriced models can be priced and never chart as $0; scan
  warnings and unparsed lines are shown; atomic saves write THROUGH symlinks; a failed
  CLAUDE.md save keeps the draft.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Alex375
Alex375 merged commit 906d674 into main Sep 14, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants