Repository navigation
Conversation
The status ladder added a fourth control to a footer that was already full, and the buttons in it could be squeezed: a flex item shrinks by default, so at a narrow panel width « Mark as Done » and « Open in TOSSE » wrapped onto two and three lines and spilled out of their fixed 22px height. Three changes, in order of what they fix: - `.act` / `.open` no longer shrink and never wrap their label. This is the bug: a button now keeps its own width, whatever the row does. - The way out to TOSSE becomes the icon alone (`compact`, as on a project card). Spelled out it was the widest thing in the footer (117px) and the least important — it was crowding out the very controls the panel exists for. - The footer itself wraps. The panel is resizable down to 380px and the conversation's side region is narrower still, so there is a width at which the row cannot hold; there it breaks into a second line of whole buttons rather than a line of broken labels. The way out keeps to the right end of whichever line it lands on via its own auto margin, the `.spacer` before it having only ever pushed on one line. Verified in the browser against the mock at the width that reproduced it: at a 417px panel the full « Approve & Done / Open / Discuss / Start ▾ / TOSSE » row sits on one line; at the 380px minimum it wraps cleanly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… block The fold body is mounted inside `.cv-collapse`, a grid used to animate the disclosure height (0fr ↔ 1fr). Its child is therefore a grid item, whose automatic minimum size is its MIN-CONTENT size on BOTH axes — and only `min-height:0` was released. Any card in the fold with a wide min-content (a diff, a code block, a table) grew the item past the reading column: the card visibly stuck out to the right and the whole thread scrolled sideways. Clean output is where this bit, because it is the mode that puts full cards inside the fold; the same card outside is a flex item of `.cv-aibody` (column), whose box stays capped. Same hole on `.cv-motion-in`, the grid item of the live travel wrapper. Measured in the mock (Playwright, 442px column): fold body 2007px → 442px, thread scrollWidth 2075px → 552px = clientWidth. Same for a step detail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The footer still needed 401px to hold one line once a task had a conversation to open (413px from « Review », whose ladder button is the longest) — and the panel is draggable down to 380px. Below that it wrapped: correct, but two rows. « Open » loses its word and keeps its bubble, so the footer now needs 364px — 374px with a count, 380px with a two-digit one. It fits at the panel's narrowest, and the wrap stays underneath as the guard it was meant to be. The count survives `compact`; only the word goes. It is the whole reason that button differs from the single-conversation one, and reading it out of a tooltip would mean hovering every task to find where the agents are. Panel footer only. The task ROW keeps the word: it shows these buttons on hover against a title that gives up its width first, so it has the room to spell things out and no reason to be read as an icon puzzle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…itch Flight Deck could only ever hold ONE Claude identity. It can now hold several, each conversation says which one it runs on, and an opt-in policy moves work off an account approaching its usage limit. The isolation mechanism, verified live against claude 2.1.263: CLAUDE_SECURESTORAGE_CONFIG_DIR scopes the credential store and NOTHING else. CLAUDE_CONFIG_DIR — the obvious candidate — would have isolated transcripts, settings.json, plugins, skills and MCP too, fragmenting the user's conversations. With the secure-storage variable, `projectsDirectory` stays ~/.claude/projects, so resume / fork / rewind are untouched by a switch. The Keychain service name is derived exactly as the CLI derives it (sha256 of the NFC path, first 8 hex), which is what lets us read a NON-ACTIVE account's usage — the only way to show every account's rate limits from one panel. The default account deliberately sets no variable at all: an empty string is not "unset", and would move the credentials of anyone already scoping their CLI with CLAUDE_CONFIG_DIR. - accounts/slot.rs: the isolation primitive, with a live probe test. - usage/: fetch_plan_usage_for(slot) — per-account token + Keychain item. - Sessions: handles are filtered BY ACCOUNT before get_usage. Asking another account's session returns a wrong-but-plausible figure the auto-switch acts on. - SQLite v11: claude_accounts + conversations.claude_account_id (no FK; removing an account detaches its conversations rather than corrupting them). - Settings → Accounts: one card per account, each with the SAME 5h/7d bars the context ring draws (PlanUsageBars, extracted so the two cannot drift). - Composer: an account chip, separate from the model picker (what answers is not who pays). It costs one slot of the bar budget; it is hideable. - Auto-switch: OFF by default, 90% trigger / 75% target ceiling, 10-min cooldown. Applied ONLY at a safe boundary — never mid-turn, never while background work runs — by ClaudeAccountApplyHost; until then the chip shows the account really in use. Every switch, and every armed-but-impossible switch, is stated in the thread. A single-account setup spawns exactly as before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… instructions block
A new backend-specific settings tab, shown only while a Claude account is connected:
which model each sub-agent runs on, what the helpers cost, and the CLAUDE.md policy
that lets Claude choose well.
## Routing (extensions/routing.rs, extensions/agent_edit.rs)
Three layers rather than one mechanism, because no single one does the job:
1. baseline `CLAUDE_CODE_SUBAGENT_MODEL` — one value, no file, no frozen prompt;
2. per-agent `model:`/`effort:` frontmatter, which beats the baseline (level 2 > 3);
3. a definition FILE for `Explore` and `Plan` only — verified: the baseline env var
does not reach those two, and that is the sole reason to duplicate anything.
`_FORCE` is offered as an explicit "lock the budget" checkbox, never on by default, and
visibly greys out the rows it overrides — which is exactly what it does to them.
Frontmatter writes rewrite ONLY the keys we own and preserve every other byte, including
the body (an agent's body is its system prompt). Round-tripped against every real agent
definition on disk, plus CRLF, BOM, block scalars, and a `---` inside the prose.
## Spend (agentspend/)
Read from the `agent-*.jsonl` the CLI already writes — no instrumentation. Measured on a
real 931-file / 194 MiB corpus: 24 066 assistant turns collapse to 50 buckets, so one
scan returns everything and the front pivots it for every table, filter and chart. A
substring prefilter before `serde_json` and a lean borrowed struct took the cold scan
from 2.3 s to 1.2 s (debug); a (size, mtime) cache makes a re-open ~10 ms. Verified
against a known-good tally: five of six models match to the token.
Costs use an EDITABLE rate card (localStorage), never buried constants, and the UI says
in place that these are API list prices, not what a subscription is billed.
## Drift canary
The resolution order changed once already (2.1.251) and a built-in's name is an
undocumented contract — either can break an override with no error anywhere. So the
dashboard checks intent against reality: an agent configured for one model but seen
running another raises a banner. Compares normalized ids, so `haiku` vs
`claude-haiku-4-5-20251001` never cries wolf.
## Instructions (memoryfile/)
Writes only between its own markers in `~/.claude/CLAUDE.md`; damaged markers are
refused rather than repaired. Ships the routing-policy block — the one lever that makes
the dosing intelligent rather than merely uniform, since Claude's own guidance otherwise
forbids downgrading a worker for looking easy.
Scope warnings are COMPUTED (`git check-ignore`, worktree detection), not hardcoded, and
"could not check" is never rendered as "fine".
Charts are inline SVG (no new dependency); the categorical palette was validated with
the dataviz validator against this app's own surface — all five checks pass.
636 Rust tests, 1609 front tests, tsc clean. Bindings regenerated.
…ings sub-tabs + search
The voice agent wrapped every answer ("C'est bon, j'ai lancé la conversation
dans le repo, je reviens vers toi tout de suite"). A sentence budget was never
going to fix that — it just fits filler inside the budget — so the brief now
bans the filler shapes by name and gives the skeleton of the two lines it says
most. The announcement lines repeat the rule: they are the last thing in context
before the agent speaks, which is exactly where the padding crept back in.
- src/voice/instructions.ts: the default brief + `resolveInstructions`
(empty override = the default, so clearing the box IS the reset).
- The brief is EDITABLE from Settings and applies to a live session
(instructions, unlike the voice, can change mid-session).
- Voice picker: the catalogue and the default live Rust-side (`voice/mod.rs`),
the choice travels to `voice_agent_client_secret` and is sanitized there — a
stale key degrades to the default instead of 400ing the whole session. The
voice is fixed at mint time, so an idle armed session is re-armed at once and
anything busier waits for the next one; Settings says which happened,
including when the re-arm itself failed.
- Settings: sub-tabs for the two tabs that had become a scroll of unrelated
cards (General, MCP Control), the voice card split into Voice agent /
Microphone / Wake word, and a search box over a flat index — picking a result
lands on the right tab AND sub-tab, then flashes the row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…record false triggers The wake word fires on background noise, onomatopoeia and silence. This is the instrumentation + latent-bug pass; the detection gates (patience, VAD veto, RMS floor) come next, on top of a VAD that now works. **The Silero VAD never once reported speech.** It needs a 576-sample tensor — 64 samples of context carried from the previous hop, prepended to the 512-sample hop, which silero-vad 5.x does inside its own wrapper. We fed a bare 512. The exported graph's sequence dimension is DYNAMIC, so onnxruntime ACCEPTS that and silently returns ~0: measured on 7 s of loud, clear speech (RMS 0.12), peak probability 0.003 at 512 versus 1.000 at 576. Nothing errored, nothing logged — the VAD simply reported silence over speech, forever, which is indistinguishable from a VAD that is switched off. It gates nothing today (see `ccf654d`), so this changed no behaviour; it is what the planned fire-time veto has to stand on. Guarded by a committed 0.6 s speech fixture, so CI checks the contract without a microphone. **Duplicate embeddings.** `Engine::feed` sliced the audio buffer's TAIL for every step, so a callback carrying several 80 ms chunks handed each of them the SAME audio — stuffing the rolling window with repeated embeddings and corrupting the temporal pattern the classifier is trained on. Rare with CoreAudio's 512-frame buffers, systematic with a larger one. `feed` now walks a per-step window end through the buffer; a test asserts the score trajectory is identical whether the audio arrives in 10 ms or 320 ms callbacks. **Non-finite features.** A muted or dead input device delivers exact zeros and a log-mel of zeros can come back -inf/NaN. That does not merely make the classifier wrong — its output becomes meaningless and can read as a confident hit, which is one way "it fires on silence" happens. Those steps are now dropped, with one latched log line rather than one every 80 ms. **Recording false triggers** (opt-in, off by default, Settings → Control). Every fire writes a WAV of the ~4 s around it plus a JSON score trajectory carrying the per-step VAD probability and window RMS. The trajectory says WHICH failure a trigger is — a lone spike a "N consecutive steps" rule would swallow, a slow climb wanting a higher threshold, or a fire at rms≈0 that no threshold can fix — and the clips are both the hard negatives a retrain needs and replayable through the real engine via `detect_wav`. It writes microphone audio to disk, so: opt-in only, capped at 40 pairs, never sent anywhere. The capture directory is created when the switch goes on, not at the first fire, so a folder it cannot write turns the switch back off and says why instead of silently keeping no evidence. Also: the score log line now carries vad and rms, and `no_bundled_phrase_fires_on_silence_or_noise` covers every bundled phrase rather than `alexa` alone (the custom `ground_control` classifier never saw room tone or broadband noise in training). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
## The canary was accusing the user of its own success
Setting Code search to Haiku made the drift banner fire immediately: it compared
today's setting against a week of transcripts that had all run under the PREVIOUS
setting. A warning at the exact moment you change something is how a canary loses its
credibility.
`AgentRouting` now carries `configured_at_ms` — the mtime of whatever file holds the
setting (the agent definition, or settings.json for a baseline-driven row). No new
bookkeeping: the file that holds the setting already records when it was written. The
canary only weighs turns that ran strictly AFTER that day, and says "since you changed
this on <date>" so the window is visible.
## Tokens lead, money trails
A dollar figure in the first tile made a subscription read as a four-figure bill, and
the disclaimer underneath could not undo the impression the big number had made. The
tiles are now Output tokens · Turns · In workflow runs · At API rates ("not your bill"),
and the daily chart plots tokens.
Dropped the "same volume on Sonnet" counterfactual: comparing token volumes across
models is hard to read and easy to misread as a forecast.
## Bars, not mountains
The daily chart was an area chart, which reads as a silhouette — the eye follows the
outline instead of comparing days, and placing yourself meant counting along the axis.
Now discrete stacked bars, with a weekday initial under each and the date on Mondays,
and `Mon 7 Sep` as the tooltip heading.
The legend became a filter: clicking a model hides it and the chart rescales to what is
left, which is the only way to see a model with a hundredth of the spend. The last
series can't be switched off — an empty chart just looks broken.
## Tooltips no longer leave the screen
Hovering the right-hand side of a chart pushed the tooltip off the window. Placement is
now edge-aware and flips to the left of the pointer, shared by both charts.
## Smaller things
- Model suggestions name the TIER, not the release: "Suggested: Haiku", not "Haiku 4.5".
Same in the routing policy written into CLAUDE.md, which outlives releases.
- The checkbox is drawn by hand instead of the near-white browser default, which shouted
louder than the setting it controls.
- Instruction blocks are framed as a list ("Instructions we suggest adding") rather than
one button, since the list is meant to grow.
- Unsaved instruction text is tinted like a diff's added lines, with a notice — the box
used to look identical whether the text was live in the file or merely proposed.
1612 front tests, 636 Rust tests, tsc clean. Bindings regenerated. Verified in the
browser against the mock: canary window, legend toggle, tooltip flip, unsaved tint.
Throwaway integration branch so both feature branches stay clean: neither `/land` should carry the other's work. Two conflicts, both in Settings, and both caused by `dev` having moved on since voice-brevity branched (it predates the new Behavior tab). `SettingsPanel.tsx` — voice-brevity adds sub-tabs and a settings search; dev had since added a Behavior tab and moved `PermissionPrefs` into it, out of General. Kept both: the sub-tab structure stands, but `general/system` renders only `CaffeinatePrefs`. Taking voice-brevity's block verbatim would have rendered the bypass switch in TWO tabs at once. The Behavior and Models sections are gated on `!searching` to match the search feature's own convention. `settingsSearch.ts` — the index still pointed "bypass" at `general/system`, which after that move lands on a page the switch is no longer on (the index test only checks that a sub-tab EXISTS, not that the setting is there, so this was green and wrong). Repointed to `behavior`, and indexed Output style, which had no entry at all. `VoiceAgentSection.tsx` — structural only: voice-brevity rewraps the component, so the "Record false triggers" row just needed reseating inside the new `SettingsGroup`. Both features intact. cargo test --lib 596 · tsc --noEmit clean · vitest 1600 · bindings regenerate to no diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four things from the second pass. ## The per-folder legend filters too The daily chart's legend toggles models; this one did not, for no reason other than that it was built first. Same behaviour now, and the bars re-sort and rescale to what is left. ## Long folder names no longer slide under the bars `web_dentiste_middleware_api` ran straight into the chart. Widening the column is not a fix — names have no common length, so sizing to the worst case starves every other row. The chart is now laid out in HTML rather than SVG, which makes the actual fixes one CSS property each: a hairline divider marks where the label column ends, the name fades out as it reaches it (rather than an ellipsis, because the point is that it continues), and hovering shows the full name plus the folder's absolute path, wrapped so a deep path cannot push the tooltip off the window. Nothing in this chart needed SVG's coordinate space — the bars are proportional widths. ## Table headers sat over the wrong columns `.cc-table th` (specificity 0,1,1) was beating `.cc-num` (0,1,0), so the headers stayed left-aligned above right-aligned figures and every number read as belonging to the column beside it. Scoped the rule to `.cc-table .cc-num`. ## The instructions box shows a real diff Tinting the whole box green said "something changed" but not WHAT — and the thing worth seeing before writing to a file you also edit by hand is the lines that would DISAPPEAR. New `lineDiff.ts` (plain LCS over lines; the managed block is instructions, not a repository, so the O(n·m) table is free and anything cleverer would only cost readability). Additions in green, removals kept on screen in red until you save them away, unchanged lines in grey, long unchanged runs elided to two lines of context, and a `+n −n` summary. The textarea keeps only a quiet left border now — tinting the editing surface made every character look like an addition. The demo fixture gained a long folder name and an existing managed block, since an empty block only ever produces green and hides half of what the diff is for. 1624 front tests (12 new on the diff), 636 Rust tests, tsc clean. Verified in the browser: legend filter, label fade and path tooltip, header alignment, red removals.
… a reply Brevity wasn't the whole ask: a SHORT pleasantry is still a pleasantry. In text you skim past an intro; spoken, every syllable is time you sit through, so anything that isn't information is noise. The agent is a radio operator now — it speaks for four reasons (report an event, answer what was asked, report an action's outcome, ask the one thing it needs) and says nothing otherwise. Greetings and sign-offs get silence, not a shorter pleasantry. The prompt could only do half of it: `runToolCall` answered end_call with a `response.create` and closed the mic only once that goodbye had finished playing — the app was COMMISSIONING the « à la prochaine » it complained about, and holding the microphone open for it. end_call now returns the tool result without requesting a turn and closes the mic at once, which also retires `endPending` and its 15 s fallback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pure silence has one ambiguity: "order taken, working on it" and "I did not hear you" sound exactly alike. Radio solved this long ago, so the brief now allows a fifth reason to speak — acknowledging an order there is nothing to report on yet — as a FIXED call sign (« Bien reçu. » / "Roger."), not a licence to improvise a polite sentence. It never prefixes a real report: the outcome already proves the agent heard, and « Bien reçu, message envoyé » is exactly the padding this brief exists to remove. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…n 25 recorded triggers The instrumentation paid for itself immediately. 25 false triggers recorded across silence, background noise, music and speech, all on `ground_control`: - 24 of 25 were a SINGLE 80 ms spike out of nowhere — the score jumped from below 0.5 straight past the threshold (median jump +0.53) and back. - 23 of 25 fired with essentially no speech in the window (Silero peak over the trailing 1.3 s below 0.2, median 0.007). - 20 of 25 fired at an RMS under 0.01, i.e. a quiet room. Against that, a real utterance holds the threshold for 12-18 consecutive steps (the ~2 s classifier window slides through the phrase slowly) with a VAD peak of 0.62-1.00. The separation is 18 versus 1 — not a threshold that needs tuning, a missing confirmation step. Raising the threshold instead would have been the wrong lever: even 0.9 leaves 4 of these 25 and would gut real recall. So, two gates before a fire: - **patience**: 3 consecutive steps above threshold (openWakeWord's own parameter, which this engine never had). ~160 ms of added latency. - **confirmed speech**: Silero's peak over the trailing 16 steps must reach 0.5.⚠️ Over the WINDOW, never the firing step alone — the classifier fires as the phrase ENDS, on audio already gone quiet (measured: vad≈0.01 at the firing step of a real utterance whose phrase steps read ~1.0). A per-step veto would reject every genuine detection. This gate only became possible now that Silero actually returns a probability. Replayed through the new engine: 24 of the 25 recorded triggers are suppressed, and 4 synthesized "Ground Control" utterances across 4 voices all still fire at 0.999-1.000. The one survivor is the single capture that contained real speech (VAD 0.999, RMS 0.05) — it needs a human ear to say whether it was a genuine wake or a false positive on speech, which no amount of analysis can settle. Debug capture now also records SUPPRESSED candidates, named `-blocked-<gate>` and tagged in the JSON. Without them a quiet capture folder cannot be told from gates that are swallowing real detections — and that is the failure mode this change introduces. Suppressed candidates are only produced while capture is on, so the normal path allocates nothing extra, and they never reach the app. Also adds `replay_corpus` (ignored): replays a whole directory of WAVs and reports fired/suppressed per file. This is how the numbers above were produced, and how the next gate change gets checked against the same corpus. No new setting: firing on room tone is a defect, not a behaviour anyone chose. The added latency and the recall trade-off are worth a look before this lands. cargo test --lib 597 · tsc --noEmit clean · vitest 1600. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…merged The panel had grown to 12 top-level tabs with a wall of switches in General → Display, and alert settings split across two places. - General → Display: one card of 11 toggles becomes three — Appearance (interface zoom, live workflow on the card), Thread (clean output, task notifications, last-message preview, minimap, message controls, clickable filenames) and Motion (the three animations, each of which the system's "reduce motion" already overrides). - Conversation absorbs Models and Composer behind sub-tabs (Markdown / Models / Composer): all three shape what a conversation is, and the drag surfaces keep their own sub-page. - Notifications absorbs the old General → Alerts behind sub-tabs (Channels / Fleet / Background): the fleet readout and the background-task alert sat in a different tab from the OS channels they belong with. General keeps Display / Durations / System. The four merged sections take an `embedded` prop that drops their own PageHead, since the tab now carries it (same shape as the Control sub-groups). The search index follows: SETTINGS_SUBS gains the two new sub-tab sets and every moved entry gets its section/sub/group, so a hit still lands on a tab that renders it. Rail goes from 12 tabs to 10. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Settings panel: reconciled the sub-tab/search restructure with the Behavior + Claude Code tabs dev landed meanwhile. - TABS keeps dev's "Behavior" and "Claude Code" (with its needsClaude gating and the claudeSeen latch); "Models" and "Composer" are gone as top-level tabs because they are now Conversation sub-tabs. - General → System no longer renders PermissionPrefs: dev moved bypass permissions to the Behavior tab, which is the better home. System keeps Caffeinate. - dev's new "Per-agent detail in the workflow view" toggle moved into the Appearance card, right next to the live-workflow toggle it belongs with. - Behavior and Claude Code render behind the same `!searching` guard as every other branch, so search results don't render under a tab pane. - Search index follows dev: bypass permissions re-pointed at "behavior", plus new entries for Output style, the Claude Code groups and the per-agent workflow toggle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`Alex375/tosse-code` was renamed to `Alex375/FlightDeck`. It is a pure rename — same repository (`id 1271050927`), and the old name still redirects — so nothing is broken today. The redirect is the problem: it dies the moment anyone creates a new repository under the old name. ## The one with teeth `tauri.conf.json`'s updater endpoint is polled by every installed app. It was resolving through a 301 rename redirect; if that redirect ever lapses, installs stop receiving updates silently — no error, no banner, just an app that never updates again. Verified before: 301 → 302 → 302 → 200. After: only GitHub's own asset 302s, no rename hop. ## The shop window The repository is public, so the README's release link and clone URL are how people arrive. The clone snippet also said `cd tosse-code`, which is now the wrong directory — a clone of `FlightDeck` creates `FlightDeck/`. ## Deliberately untouched - The ~12 `Alex375/tosse-code` strings in `git/mod.rs` and `tosse/mod.rs` are all inside `#[cfg(test)]`. They are arbitrary URLs exercising `normalize_remote_url`; the repository name is incidental to what they test. Changing them buys nothing and risks the assertions. - `CLAUDE.md:103` still names the old repo, and is left alone ON PURPOSE: that line lives inside the `[GENERATED] Repository Context` block, which `sync_claude_md()` regenerates from the TOSSE CRM. A hand-edit there is reverted at the next release — which is exactly what happened to the governance section (`b3dcb6e` corrected it on 4 Sept, `e8c77e8` overwrote it on 7 Sept). The CRM context is the source of truth; that is where this gets fixed. Out of band: the CRM repository record's `url` was corrected to the new name. That one is functional rather than cosmetic — folder↔repo association runs through `normalize_remote_url` on the git remote, so changing the local `origin` without changing the CRM would have broken the link. 1624 front tests, 636 Rust tests, tsc clean.
…ganisation `feat/voice-brevity` is on dev now, so the throwaway integration branch became the feature branch (renamed `feat/wake-detection-gates`) and follows dev instead. Both conflicts resolved to dev's side: it had since done the same reorganisation properly and further. Alerts moved out of General into Notifications, Models and Composer became Conversation sub-tabs, Claude Code gained its own tab — our older shape is superseded, not complementary. Same for the search index: dev had already repointed the Behavior entries, so our version of that fix is redundant. cargo test --lib 653 · tsc --noEmit clean · vitest 1641. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hreshold The threshold slider felt inert, and the meter explained why: normal speech (RMS ≈ 0.05) drew 16% of a bar whose threshold handle sat at 60%. The level could not reach the handle no matter how loud you spoke, so the instruction printed under it — "drop the handle just above your noise floor" — drove the threshold to its minimum if followed. Rescaling the bar would only have made the fiction plausible. The two numbers do not share a scale at all: `server_vad.threshold` is a CONFIDENCE from OpenAI's own speech detector, not a loudness, so no local level bar can be calibrated against it. A bar that lines up by arithmetic while meaning something else is worse than no bar, because it looks authoritative. So each control now claims only what it can prove: - The threshold is its own slider, described as what it is — how sure their detector has to be — with the honest instruction to find it by ear. - The meter is a microphone CHECK: is the right input selected, is it hearing me. On a dBFS scale (-60..0), which puts room tone near 15% and speech near 55% instead of everything squashed into the bottom sixth. Turn detection stays OpenAI's job. Gating the audio ourselves before it leaves the machine was considered and dropped: it buys a threshold we could draw on a bar, at the cost of a second detector to keep honest, a delay line so it does not clip the first syllable, and hysteresis so it does not chop a sentence in two. Both mic call sites now share `mic.ts`, so what the meter shows is what the agent hears — they had drifted apart by accident, which is exactly how a diagnostic starts lying. The meter also reads back what the track actually negotiated (`getSettings`), because a constraint the platform ignored should be visible, not assumed.⚠️ Left deliberately unchanged, and written down in `mic.ts`: auto gain control is ON. Its job is to make loud and quiet input come out at the same level — in a quiet room it winds gain up until the noise floor reads like speech — so every threshold downstream, OpenAI's included, is working on a deliberately flattened signal. That is the leading suspect for the remaining over-sensitivity, and it is a one-word change in one place now. It is also a real trade-off (quiet speech gets quieter), so it wants a measurement rather than a guess. tsc --noEmit clean · vitest 1648. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…dates was evicting the fires First session with the gates on wrote 38 blocked candidates to 2 real fires. With one shared cap of 40, that ratio prunes away exactly the evidence worth keeping: the fires — rare, and the only captures that show something actually went wrong — were gone within minutes, crowded out by the plentiful ones. Fires and blocked candidates now have independent budgets (40 and 60), so a busy stretch of blocked candidates can never cost a fire. Test asserts it directly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ces and a confirmed false positive Armand recorded ten deliberate "Ground Control" utterances and confirmed which earlier capture was a false positive, which is the first time this had real evidence on both sides rather than false positives alone. What the recordings show: the FIRST step of a real detection is a ramp-in — the classifier window is only part-way into the phrase, so it scores 0.66-0.82 before pinning at ~1.00 for the rest. So demanding three consecutive steps over a high bar fails on the ramp, not on the phrase. At threshold 0.90, patience 3 lost 4 of 10 genuine utterances; patience 2 lost none. Two steps over a high bar beats three over a low one, and it is 80 ms FASTER than what it replaces. The threshold band was aimed at the wrong place entirely. Scores here are bimodal — near 0 or near 1 — so nothing interesting happens below ~0.8, yet the slider spent its whole upper half between 0.30 and 0.60. And it topped out at 0.90, so the one region that separates real from false was not reachable at any setting. The band is now 0.82-0.98, default 0.90. Replayed through the engine: - 10/10 real utterances fire. - 25/25 confirmed false positives suppressed, including the one Armand identified (it peaks at 0.92 and touches it exactly once; real utterances hold 0.90 for at least two consecutive steps). - The capture that fires out of the "false" folder is the one he confirmed as a genuine wake — correct, not a miss. The one recording lost at this setting turned out to be a deliberately distorted phrase, so rejecting it is the wanted behaviour rather than the cost it looked like.⚠️ This is tuned on ONE confirmed false positive, which places a boundary inside a 0.01 gap. It holds on the evidence we have and no further. The durable fix is retraining a classifier that has never heard background noise or a real human voice — TOSSE task 64a42f06, with every capture corpus listed. cargo test --lib 655 · tsc --noEmit clean · vitest 1648. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…a sound Armand set the threshold to its maximum, 0.90, and talking while emptying a dishwasher still cut the agent off every ten seconds. Worse, the agent ANSWERED the noise: asked "shall I grant permission X?", it heard a clatter and replied that it was granting. That is a consequential action taken on crockery, and it made the feature unusable. No threshold fixes that, because loudness is not the problem. A plate on a counter IS loud; `server_vad` is an amplitude gate with no way to know the sound carries no words, and a threshold high enough to exclude dishes excludes the user too. The slider was not failing to work — it was working, on the wrong question. `semantic_vad` asks a different one: does what was said sound FINISHED. Audio with no words in it is not a turn at all. It is now the default, with `eagerness: low` so a pause mid-sentence is not treated as the end of one. The loudness gate stays available as a mode, since it is the one that offers a number to turn, and Settings now says plainly what each is and when the choice matters. Two more layers, because turn detection alone should not be the only thing standing between a noise and an irreversible action: - The brief gains a rule that outranks everything else in it: a turn you could not make out is not a turn — do nothing, say nothing, and never infer an answer from a sound. "Yes", "go ahead" and "delete it" must come from a sentence the agent actually heard. It had no such rule, and a model asked a question will find an answer in whatever comes back. - Server `error` events reached `console.error` and nowhere else. A `session.update` the server rejects is reported there and ONLY there, so turn-detection settings could be silently discarded while Settings showed them as applied — someone turning a slider the server threw away had no way to find out. They now raise the app error banner. (Checked against the current API docs: our `server_vad` shape is valid, so this was not the cause here. It could have been, invisibly.) Stored mode/eagerness are coerced on read and on write: an unknown enum from an older build would be rejected by the server as a whole-payload error, taking the rest of the session config down with it. tsc --noEmit clean · vitest 1649. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e MCP Control to Control Semantic turn detection did not fix it and probably made it worse: the agent now cut out mid-sentence roughly every thirty seconds, with no relation to how loud the room was. That "no relation to the noise" is the tell, and it points away from the room and at the agent itself. Its voice comes out of the speakers and back into the open microphone. Echo cancellation is on but imperfect on laptop speakers — and where an amplitude gate at 0.90 might reject a quiet echo, a SEMANTIC detector is looking for words, which is precisely what an echo of speech is made of. Asking "did someone just finish a sentence" of the agent's own sentence gets a yes. Switching to semantics sharpened the wrong instrument. Armand's verdict settles the default regardless of cause: an agent that never stops beats one that stops constantly. A long answer you must sit through costs seconds; an interruption costs the whole exchange. So barge-in becomes a choice, defaulting to off: - **Let it finish** (new default) — nothing stops it mid-sentence. Speaking over it still WORKS: `create_response` stays on, so the turn is answered once the agent is done rather than over the top of it. Nothing is lost but the cutting. - **With the wake word** — his own suggestion, and the best of the three. The wake detector is a far stricter judge than any turn detector: a specific phrase, two consecutive steps over 0.90, vetoed unless Silero agrees it was speech. It does not fire on a room. The app sends the cancel itself rather than handing the server barge-in, so the strict judge is the only judge. - **By speaking** — the old behaviour, kept and labelled with what it costs. Interrupting needs BOTH halves to look like it worked: `response.cancel` stops the model generating, `output_audio_buffer.clear` drops the audio already sitting in the WebRTC playout buffer. Cancel alone leaves it talking through what it had queued. `shouldFireWake` gains the one narrow exception to its self-trigger guard: in wake mode the phrase MUST be heard while the agent is speaking, since that is exactly when you need it. It still refuses a session that is off, unconfigured or broken. Also renames the MCP Control tab to Control, as asked. tsc --noEmit clean · vitest 1652. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
37 confirmed findings, ~15 root causes. The two critical ones broke the feature's
headline capabilities for every added account:
- Keychain item name was derived in the WRONG ORDER. The CLI calls `Sx("-credentials")`,
so an isolated slot's service is `Claude Code-credentials-<sha8>`, not
`Claude Code-<sha8>-credentials`. Invisible on a single-account machine (empty
suffix), fatal for every added account: its usage could never be read, so per-account
limits and auto-switch were dead. The old unit test only asserted the shape the code
chose; it is now pinned to a literal digest, plus a live probe that shims `security`
and asserts against the service the REAL CLI queries (verified: it asks for
`Claude Code-credentials-60a207a7`, exactly what we derive).
- The pasted OAuth code was not bound to its account. One global in-flight login, one
card per account: a code typed into a superseded card was redeemed into another
account's store and stamped the wrong identity. The login now records its account,
submission verifies it, and a superseded card closes its code box.
Data integrity:
- `conversations.claude_account_id` is now written on INSERT only; the dedicated
`set_conversation_claude_account` is its sole updater. A stale in-memory copy
re-upserted on any activity bump was resurrecting ids a removal had just detached.
The front also mirrors the detach and clears a removed default.
- Removing an account ABORTS when `claude auth logout` fails: the Keychain item name is
derived from the directory being deleted, so proceeding orphaned live, unrevokable
tokens. "Remove anyway" names the exact item to revoke by hand. Refused while a
session runs on it.
- Placeholder labels are a persisted fact (`label_is_generated`), not a prefix guess
that clobbered names like "Account manager".
Silent failures now surfaced: unreadable CURRENT account (auto-switch was a silent
no-op), frozen usage after a terminal error, identity-capture failure, rename failure,
failed account-list read, unknown account typed as `UnknownAccount` instead of a
network error retried forever. The default-account preference no longer drops when
the list is unloaded (seeded at boot; unloaded ≠ empty). Remote (SSH) conversations
refuse an account the launcher cannot carry. The orphaned-account chip stays visible.
Threshold inputs commit on blur, reject out-of-range values, and the ceiling has its
own control; the hysteresis invariant now holds on every clamp path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Multiple Claude accounts with manual selection, per-account rate limits in Settings → Accounts, and opt-in auto-switch near the usage limit. Conflict: src/ipc/bindings.ts (generated) — both sides added exported types at the same alphabetical position (dev: SpendBucket/SpendReport/SubagentBaseline/ SubagentRouting; feature: SpawnFlags). Resolved by regenerating from the merged Rust source rather than hand-editing, so every type from both sides is present. Integration: dev's new Settings search indexed the Accounts tab with a single generic entry, so none of the feature's new rows (auto-switch, default account, thresholds, per-account limits) were findable. Added them to SETTINGS_INDEX. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…g band The page had grown into one ~340px ConnectionCard per account (gradient tile, coral glow and border, provider line repeated, "Connected" shown twice, monospace name field), so two accounts filled the screen before any setting. Chosen from three mocked directions: tiles for the accounts, the switching band and the numbered sign-in from the compact-list direction. - Claude accounts: a two-column grid of tiles, each with a ring gauge per usage window (5h / 7d, amber past 80% like the context ring) and model-scoped caps on a quiet line. Rename / Make default / Sign out / Remove move into a portalled ⋯ menu; sign out and remove confirm first. A dashed tile adds an account. - Sign-in: numbered steps inside the tile. Step one is only marked done once the browser really opened; if it failed, the manual link fallback shows instead. - Switching: default account as a segmented control, the auto-switch toggle (disabled with its reason while there is nowhere to switch to), and a 0–100% band — the hysteresis gap hatched, amber past the trigger — with every account placed at its busiest window. Accounts closer than 12 points share one marker (their labels were printing over each other); the merged marker sits at its busiest member. Thresholds are steppers that commit on blur/Enter and never ratchet the ceiling. - Codex: the same tile language. No behaviour dropped: the account-bound login, superseded-login notice, refused removal with "Remove anyway", identity-capture warning, typed usage errors, and the shared-profile-cache guard on the default tile all carry over. ConnectionCard is untouched for the TOSSE tab. The settings search index follows the new titles, and the mock now starts added accounts signed out so the sign-in flow is exercisable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t after every tool call Armand changed a setting mid-conversation and the banner read: "The voice agent rejected a setting: Conversation already has an active response in progress: resp_ENz8mbt061hUCNaoQzB5k. Wait until the response is finished before creating a new one." That line was wrong twice, and it hid a real bug. Wrong label: the error came from a `response.create`. A `session.update` never creates a response, so no setting was refused. The previous commit had pushed EVERY server error into the banner under that one prefix, whatever caused it. Wrong audience: a response id means nothing to a person. Zero silent errors does not mean dumping the server's wording on screen. The real bug underneath: `runToolCall` fired `response.create` the moment a tool finished. A tool call is reported by `response.output_item.done` while the response CONTAINING it is still open, and app-control tools run locally in milliseconds, so the request routinely arrived first and was refused. That refusal was the only trace: the agent carried out the action and never said so. It now waits for the server to be quiet first (the same bounded wait announcements already use). The tool output goes in BEFORE the wait, so any response that starts in the gap still carries it. Errors are now sorted by cause, which the API reports exactly: every client event can carry an `event_id` and the error echoes it back (checked against the current client-events reference). - One of our tagged setting updates was refused → a plain note inside the Microphone card: the change applies the next time the agent starts. That is true, because a new session reads the stored preferences on connect. - The active-response race → absorbed. After the wait above, it can only happen when the server starts its own response from the user speaking, and that response already carries the tool output. Logged, not shown. - Anything else → a plain sentence in the banner. The server text stays in the log. No "applies on the next start" notice appears on every change, as was briefly considered: the same reference says `session.update` applies live for every field except voice and model, and voice already has its own note. Showing that notice on every change would be false. The classifier is a pure module (`serverErrors.ts`) with tests, including one that keeps ids and server wording out of anything user-facing. tsc --noEmit clean · vitest 1657. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…atus The attention tint (input/review/error) replaced the .on background, so the open conversation was indistinguishable from other rows with the same status. The selected row now gets a deeper wash of its status colour plus a thin ring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An account is its email address — two accounts cannot share one — so the tiles, the default-account selector and the switching band all show the address. The generated "Account 3" label only stands in until an address is known, and the rename action is gone with the name it edited (`claude_account_rename` removed from the IPC surface rather than left dead). The default account needed real work to say this honestly: its address cannot be read live once a second account exists, because `claude auth status` answers from a profile cache every account SHARES and would name whichever signed in last — which is why the old card hid it. Its identity is now captured at its OWN sign-in and stored in `meta` (`claude_default_identity`), the same treatment the added accounts already got in their row. Absent that capture (signed in outside the app) the tile keeps the "Claude" fallback and says why, instead of showing an address that may belong to another account. Also: Codex is named by its address too (its status is its own, never shared), and a marker near either end of the band anchors its label to that edge — a full address centred on a marker at 8% hung outside the card. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each account's email is read from GET /api/oauth/profile with that account's own token (claude_account_identity), instead of the CONFIG-dir profile cache the CLI shares across accounts. The address now names accounts in the Settings tiles, the default-account selector, the threshold band, the composer account chip and its menu, and the auto-switch notices. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…"Other" A pending AskUserQuestion is itself a can_use_tool, so the voice agent announced it as "blocked on a permission prompt for the AskUserQuestion tool" and relayed dictated answers with send_message, which only queued behind the open question. - Classify attention on the tool, not on "anything pending" (shared attentionFields for the OS ping, the journal and the announcement), and carry the question text. - Expose get_pending_request + answer_request to the voice agent. - answer_request: questions need no opt-in; the executor builds updated_input.answers from a loose `answers` payload (free text lands as the "Other" choice) and refuses answers that match no question. Real permission prompts stay behind Settings -> Control. - Extract the framework-free questionnaire helpers into questionnaire.ts. - Voice brief: questions are not permissions, never answer with send_message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ts search
- Search no longer offers results in a hidden tab: picking a Claude Code result while no
Claude account is signed in used to bounce the user back to General.
- The Claude Code tab now leaves the rail after a successful "signed out" read; the
latch only survives a FAILED read, as it was meant to.
- A search highlight no row claims (a page heading, a tile, a label) disarms itself after
2.5 s instead of flashing that row out of the blue whenever it appears later. Results that
name a label inside a card ("Switch at", "Target below") flash the card via a new `flash`
redirect, cross-checked by a unit test.
- git: `path_is_ignored` had been inserted between `normalize_remote_url`'s doc block and
its `fn`, stealing its documentation. Moved above it; pure move.
- CHANGELOG v2.2.0: add the voice turn-detection change and the clean-output width fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t and the Claude Code tab Multi-account - Remote (SSH) conversations: always seeded on the default account (create, worktree, reopen, fork), refused a non-default one in the store, skipped by the auto-switch, and excluded from `claude_handles_for` so the default account's ring never shows the SERVER account's quota. - Account restart only at a real boundary: not busy, no `activity`, no queued message, no pending permission, no background task, no spawn in flight — then a 3 s settle and a re-check against fresh state. `!busy` alone killed work the CLI takes back after a result. - Auto-switch: transient read failures keep the last good figure; terminal ones are reported once per ACCOUNT into live conversations only; nothing is evaluated until every probe has answered once; only live conversations are moved; a user pick pins the conversation against the policy; told markers are voided by any account change. - New non-terminal `UsageError::TokenExpired`: an idle added account's lapsed access token no longer reads as a revoked sign-in that stops polling forever. - Account removal refuses while a sign-in for it is awaiting or redeeming its code, and an account-lifecycle RwLock closes the check-then-delete window against spawn / sign-in / identity capture. - The default slot removes an inherited CLAUDE_SECURESTORAGE_CONFIG_DIR; a corrupt stored default identity is an error shown in Settings, not a silent null; the usage ring follows the account the live process actually runs on. - Docs: `claude auth login/logout` DO rewrite the shared profile cache in ~/.claude.json (verified in the 2.1.270 bundle); the CLI refetches it at the next session start. Claude Code tab - Spend: fold the CLI's per-content-block lines by (message.id, requestId), keeping the final output count — totals were inflated ~4.5x. - The budget lock follows the baseline (and clears with it); the forced model is shown. - Effort edits send the file's own `model:` (new `defined_model`), never the baseline. - Plugin agents are read-only with a visible reason; drift ignores locked and `inherit` rows and is worded as an observation; `settings.json` read failures surface as `SubagentBaseline.error`; unpriced models can be priced and never chart as $0; scan warnings and unparsed lines are shown; atomic saves write THROUGH symlinks; a failed CLAUDE.md save keeps the draft. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release v2.2.0. Voir les commits de dev depuis la dernière release.
🤖 Generated with Claude Code