Skip to content

Release v2.5.0 - #57

Merged
Alex375 merged 164 commits into
mainfrom
dev
Sep 20, 2026
Merged

Alex375 merged 164 commits into
mainfrom
dev

Conversation

@Alex375

@Alex375 Alex375 commented Sep 20, 2026

Copy link
Copy Markdown
Owner

Release v2.5.0. Voir les commits de dev depuis la dernière release.

🤖 Generated with Claude Code

clousty8 and others added 30 commits August 24, 2026 20:11
Server app for the always-on remote-workstation project (TOSSE fa059422).
Design brief in docs/CADRAGE.md; milestone M0 = simplest remote (SSH into a
local container running claude, stream it back).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The M0 SSH transport is now wired into the Flight Deck app (tosse-code branch
feat/remote-ssh). These are the flightdeck-server-side helpers to drive it:

- scripts/gen-ssh-config.sh: write a self-contained ssh_config (alias flightdeck-m0
  -> 127.0.0.1:2222, M0 key) under .secrets/, so the app's SSH transport reaches the
  container via TOSSE_SSH_CONFIG without touching ~/.ssh/config.
- scripts/open-flightdeck-remote.sh: refresh the container's Claude creds, gen the
  ssh config, and launch the dev build seeded + pointed at the container.
- docs/M0-APP-REMOTE.md: the 60-second demo, how to connect a remote repo, the
  first-cut scope + caveats, and the auth-refresh note.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A design proposal for the human-facing onboarding on top of the SSH transport:
- Parcours A — pair a fresh remote device (one-command paste, Flight-Deck-generated
  dedicated key with only the PUBLIC key leaving the Mac, claude login on the server,
  return ticket + host-key TOFU). Forward-compatible with the CADRAGE daemon bootstrap.
- Parcours B — add a repo that lives on the remote (machine picker → detect/browse/
  type a remote path → RepoRecord{ssh_target,path}), with the mechanics explained.
- 3-bis — the manual equivalent you can follow TODAY (ssh config Host + seed env) to
  connect any box before the UI exists.
- What exists vs to-build, forward-compat, out-of-scope, and 5 decisions to settle.

No code — proposal to iterate on together.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ine model

- open-flightdeck-remote.sh: seed the container as a "machine" via explicit SSH coords
  (host/port/user + the throwaway key) — the transport is self-contained now, no
  ssh_config needed. App opens with a working remote conversation AND the pairing UI.
- docs/M0-TEST-PAIRING.md: step-by-step to (1a) test the seeded remote conversation,
  (1b) pair a "fresh" server from the UI (generate key → paste one-liner on the box →
  Test & pair → open a repo), and (2) drive that same remote conversation from the
  phone via Remote Flight Deck (relay → Mac → SSH → container). Plus the auth-refresh
  note and the ship step.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…iner caveat

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…+ detached sessions)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…elay node

Sessions survive client disconnects: per-session actor owns the claude
process (stream-json, persistent), seq-numbered replay ring (64MB), attach
plane over a unix socket for SSH clients (fd_attach/fd_detach/fd_stop,
epoch+cursor replay), pending-permission tracking re-emitted on attach.
Relay client presents the daemon as a node on the existing Railway relay
(authorize_phone, _cid echo, turn_completed/needs_attention events) and
answers the phone RPC catalog (list/read/send/create/interrupt/stop/
pending/answer/browse) from its SQLite registry + claude transcripts.

Proven locally (Mac, isolated HOME): mid-turn client kill -> turn finishes
unattended -> reattach replays from exact cursor; mock-phone through the
production relay lists + drives the same session (PHONE_DIRECT_OK).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…hone cut via relay)

m1-daemon/: multi-stage image (rust builder + m0 base + flightdeckd),
entrypoint running sshd + the daemon, up.sh (build, start, inject SSH key +
Claude creds from the Keychain, seed /work/demo, print the pairing link).

tests/detach_test.py — attach through REAL ssh, SIGKILL mid-turn, turn
finishes unattended, reattach replays from the exact cursor.
tests/phone-cut-test.mjs — phone creates a conversation via the production
relay, drops the socket mid-turn, reconnects: full history there, Mac off.
Both pass against the container.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…eck container

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ndings)

Concurrency/lifecycle:
- session creation fully serialized under the sessions lock (resolve + spawn
  + insert as one critical section) — no more check-then-spawn double-spawn;
  entries carry a generation token so an exiting actor can only remove itself
- claude stdin gets its own writer task: the actor queue-accepts and can never
  block on a full pipe; all actor round-trips bounded (status 3s, RPC acks 15s)
  so one wedged claude fails its RPC instead of hanging the daemon
- exit finalization waits for BOTH the exit status and stdout EOF — a racing
  wait() can no longer drop claude's final lines
- attach client queue byte-budgeted (128MB): a stalled half-open ssh link gets
  dropped and reattaches from its cursor instead of ballooning memory
- one failed accept() no longer kills the attach server (and every session)

Protocol conformance (PWA contract):
- event detail fields flattened onto the event object (attention_cleared.reason
  / request_id etc. — the PWA reads them top-level); event text clipped 1500
- read_conversation trims oldest turns to stay under the relay's 256KB frame cap
- relay backoff now really resets after an established connection; 90s read
  deadline detects silently-dead TCP links

Semantics:
- fd_attach now carries the daemon's busy state + pending permission ids (the
  Mac resyncs a stuck busy flag and drops stale permission cards on reattach)
- cold start with --resume-session prefers the existing registry row over a
  client-minted conversation id (no duplicate rows after a daemon restart);
  attaching to an archived conversation un-archives it
- new 'flightdeckd stop --conversation X' (socket op) so the Mac's explicit
  Stop works even when its attach link is already gone
- ensure_args no longer panics on a bare trailing --resume

13 unit tests green; ssh detach test + phone-cut test + both app live tests
re-run green against the rebuilt container.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…turns only

The fixed 16s wait raced the cold-start + turn duration, and the sentinel
regex also matched the user's own turn. Now: reconnect, then poll
read_conversation until DONE_PHONE appears in an ASSISTANT turn (90s bound).
Strongest observed run: status 'running' at reconnect (session alive mid-turn
after the cut), live turn_completed event on the reconnected socket, reply
landed. Doc: mention flightdeckd stop + the adversarial review pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
frames::fd_status builds the status reply with version =
CARGO_PKG_VERSION (frames::DAEMON_VERSION) so a client can detect version
skew over the attach socket itself; distinct from `flightdeckd --version`
(the binary on disk). Adds a cfg(test) testutil module (temp socket +
in-memory manager) and tempfile as a dev-dependency.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The attach writer did an unbounded write_all/flush: a degraded-but-alive
link (Wi-Fi cut under a half-open ssh) blocked it forever, the client
stayed attached and the daemon queued every line for it (the byte budget
never trips on small stream deltas).

Every write is now bounded by session::ATTACH_WRITE_TIMEOUT (20 s, on
progress, not per line). On a stall: one best-effort 2 s write of the
torn line's remainder (never a spliced line) + fd_detach{stalled}, then
the pump ends — dropping the queue receiver, so the actor's next push
fails and it forgets the client (no session.rs state-machine change).
The fd_detach wire shape is unchanged; "stalled" documented in frames.rs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… parse-order loss

Extracts run_actor's inline FdDetach reason handling into a pure,
unit-testable reconnect_policy_for_reason(reason, exit_code, message),
the single table a new detach reason gets added to instead of a second
hand-edit of run_actor's match. Adds "stalled" as the only
reconnect-eligible reason (every other/unknown reason keeps today's
terminal behavior, verified table-driven).

Fixes reader_loop counting a replayable line as "seen" before it was
successfully parsed: a replayable-typed line that failed to parse was
counted anyway, so on the next reattach the daemon believed we already
had it and never resent it — permanently lost. The increment now only
happens on a successful parse. A line that still fails to parse after
the fix would otherwise be re-requested identically forever (the
daemon replays deterministically), so a bounded cross-reconnect
counter (malformed_replay_step) gives up after 3 consecutive
reconnects that each saw the same kind of failure: it forces the
cursor past those bytes and surfaces a one-time protocol_error notice,
trading a handful of wasted round trips for guaranteed forward
progress instead of an unbounded resend loop.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A1 — probe_remote now runs an ACCUMULATING remote script: independent
claude/flightdeckd checks instead of a mid-script `exit 3` that silently
skipped the flightdeckd check whenever claude was ALSO missing. Adds a typed
RemoteProbeResult, MIN_DAEMON_VERSION + version_at_least (pure, degrades
malformed versions to 0 instead of panicking), and hard-blocks pairing on a
missing/outdated flightdeckd with a message naming it specifically.
flightdeckd is looked for on PATH, then ~/.local/bin and /usr/local/bin,
mirroring how the daemon-attach command resolves it.

A3 — generate_machine_key now reuses a FIXED ssh_keys/pending(.pub) keypair
instead of minting a fresh {slug}-{uuid} pair on every call (7 keys for one
pairing before this). PENDING_KEY_LOCK makes concurrent calls race-free.
add_machine claims the pending key by renaming it to <machine_id> on
success, so the next generate_machine_key call mints a fresh pending pair.
delete_machine now reads the MachineRecord before deleting it and
best-effort removes both key files (Store::delete_machine is SQL-only, so
every removed server used to leak its key files). Front end clears genKey
immediately on a successful pair() and surfaces a specific "already used"
message when identity_file no longer exists on disk.

A4 — the pairing command is now joined with "; " instead of "\n": a serial
paste target submitting each line on its own Enter could leave an
unterminated quote/subshell open, silently swallowing the ticket line.
Adds Tailscale/LAN/hostname address discovery embedded in the ticket as
addresses: [{kind,value}] (host stays addresses[0].value for back-compat),
with the confirm screen letting the user pick among discovered candidates
and preferring a Tailscale name by default. parseTicket synthesizes a
single manual candidate from host for tickets printed before this change.

Rust: 22 new/updated unit tests in ipc::commands (version comparison, pure
probe-output parsing over captured stdout/stderr/exit-code triples, pending
key reuse/concurrency/claim/cleanup). Front: 9 new vitest cases for
buildServerCommand (single-line regression) and parseTicket (addresses
shape + old-ticket back-compat). Bindings regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e outage

detach_test.py: Attach gains pid/pause()/resume() (SIGSTOP/SIGCONT of the
local ssh child) and byte accounting; A/B/C become scenarios_abc(), D runs
on its own throwaway conversation (argv selects: abc / d / default all).
D fills the path (the ssh window buffers ~2.2 MB — measured) with image
Reads interleaved with small Bash calls, stalls the link until the turn
is done + WRITE_TIMEOUT + 5 s, resumes, and asserts the stream ENDS after
the buffered prefix (optional fd_detach{stalled} last), never carries the
turn's result, and that a reattach from the cursor replays the withheld
backlog up to the result.

Fixes the bug D surfaced: the attach bridge read stdin through
tokio::io::stdin, whose uncancellable blocking read stalls runtime
shutdown — after the daemon closed the stream the bridge lingered with
stdout open, so ssh never saw EOF. stdin is now read on a detached std
thread (bounded channel keeps backpressure). tests/attach_bridge.rs
covers it with the real binary.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The claude session id is the join key between the Mac's conversation ids
and the daemon's registry ids (they legitimately differ). Surfaced from
what is already tracked: the live StatusSnapshot wins over the registry
row (which can lag on a brand-new conversation); null until known.
testutil gains a fake claude script (init frame + held stdin).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…overed parse failure

The reattach cursor was computed from Transport::lines_seen() — a raw count
of successfully-parsed replayable lines this connection — treated as the
daemon's wire position. Those are not the same thing: if a replayable line
fails to parse and MORE replayable lines go on to parse fine later in the
SAME connection (the common case, not an edge case), lines_seen() silently
overtakes the failed line's true position. Reattaching with that cursor
told the daemon "I already have everything through here", so it never
replayed the failed line again — permanently and silently lost, exactly
the class of bug the parse-order fix was meant to close, while also
starving the new give-up/streak safety net (it only sees the failure when
the bad line happens to be the very last one a connection ever delivers).

Add Transport::first_unparseable_offset(): lines_seen's value at the
moment the FIRST unparseable replayable line hit a connection, frozen
there regardless of what parses afterward. Extract the cursor composition
into a pure, unit-tested reattach_cursor_delta(lines_seen,
first_unparseable_offset, force_advance): while still retrying, roll back
to first_unparseable_offset so the daemon resends starting right before
the loss (accepting a bounded re-delivery of already-displayed lines
instead of a silent, permanent drop); once malformed_replay_step's streak
bound forces a give-up, skip past everything the connection delivered as
before.

Tests: transport.rs OK-FAIL-OK-OK duplex-stream test asserting
first_unparseable_offset freezes at the pre-failure count instead of
following lines_seen; session.rs table tests for reattach_cursor_delta
covering the clean, still-retrying and give-up cases, including the
divergent lines_seen vs first_unparseable_offset case the bug lived in.

Also tightens the AttachPoint::cursor doc comment, which previously
described the cursor as a received-line count — the reading that led to
this bug.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…/B5)

Adds src-tauri/src/bootstrap/ with two pure, self-contained primitives for
the remote flightdeckd provisioning flow:

- templates.rs (B4): render_user_unit/render_system_unit/
  render_persistence_escalation, pure string generators for the two systemd
  unit files and the one sudo-scoped escalation script. shq() is reused for
  every real shell interpolation (render_persistence_escalation, and the
  home directory inside ExecStart=/Environment=, both of which honour
  shell-style quoting); User=<user> is left unquoted on purpose — verified
  live against the real reference server (josty-cc, systemd 245) that
  quoting a User= value is NOT shell-unquoted by systemd and breaks the
  common case (systemd-analyze verify: "Accepting user/group name
  ''josty''..."). render_system_unit matches that server's real unit file
  byte for byte.

- askpass.rs (B5): a per-call mkfifo-based SSH_ASKPASS relay so the
  bootstrap's first SSH connection (before a key exists) can prompt for a
  password from the Tauri GUI instead of hardcoding BatchMode=yes like
  every other ssh call in this codebase. AskpassGuard is an RAII guard
  (0700 temp dir + fifo + helper script) removed via Drop on every exit
  path; bootstrap_ssh_command() builds the ssh invocation, run_with_password
  () races delivering the password against ssh exiting on its own and
  classifies the result into a typed BootstrapError (WrongPassword/Timeout/
  HostUnreachable/HostKeyMismatch/Other) whose wording never carries the
  password.

Tests: golden-string + shell-metacharacter-username tests for all three
template renderers; AskpassGuard Drop-on-every-exit-path tests; a
password-never-leaks test mirroring tosse::session_gone_errors_...; a fast
non-ignored real-ssh test (closed local port -> HostUnreachable); and an
#[ignore]d live test against a throwaway Docker sshd (linuxserver/openssh-
server) proving wrong vs right password end to end - run manually via
`cargo test --lib -- --ignored --nocapture` and confirmed passing.

shq() in ipc/commands.rs is now pub(crate) so bootstrap::templates can
reuse it instead of growing a second escaping helper.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ct GO

cargo-zigbuild (zig as cross C compiler/linker) in a rust:1-alpine
builder links both targets statically with rusqlite(bundled) and
rustls/ring unchanged; ring needs no fallback on aarch64-musl.
x86_64 6.3 MiB, aarch64 5.8 MiB; ldd: not a dynamic executable.

flightdeckd/scripts/build-musl.sh builds both (Docker only, any host,
caller's uid, incremental); smoke-musl.sh runs --version, init, run
(attach socket + real TLS handshake) and status in stock Debian per
arch. Verdict, versions, timings and fallbacks: docs/B0-MUSL-SPIKE.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ix Tailscale regex

Applies the review findings on feat/pairing-reliability:

- transport::resolve_remote_daemon_bin makes build_remote_command (and
  run_remote_stop) search PATH -> ~/.local/bin -> /usr/local/bin for
  `flightdeckd`, the SAME order probe_remote's script already searched.
  Previously the probe searched all three but the actual attach bare-`exec`'d
  daemon_bin with no fallback, so a server with flightdeckd only under
  ~/.local/bin would pass pairing and then fail on the very first attach.
  Verified against a real POSIX shell with flightdeckd stubbed in each spot.

- add_machine's pending-key claim now goes through claim_pending_key_locked,
  which holds PENDING_KEY_LOCK around the rename. Two concurrent add_machine
  calls sharing the same not-yet-claimed pending key no longer race the
  rename into a raw OS error for the loser: a rename that fails because the
  source already vanished now maps onto the same "already used" message
  stale_identity_file_error produces.

- delete_machine_and_key no longer silently discards a key-file removal
  failure that isn't "already gone" — it's logged via eprintln!.

- ControlSection's Tailscale DNSName grep required no space after the colon,
  but real `tailscale status --json` (Go's json.MarshalIndent) always prints
  one — Tailscale discovery was dead code against any real install. Fixed
  the pattern to tolerate optional whitespace, and added a test that runs
  the ACTUAL extracted shell pipeline through /bin/sh against both real and
  compact JSON shapes.

- The pairing wizard's discovered `addresses` are now cleared when leaving
  the ticket-derived path (Back to the command step, or switching to manual
  entry), so a stale set from a previous ticket can't ride along to a
  different host.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
templates.rs (blocker/major, newline injection): render_user_unit/
render_system_unit now reject an embedded \n or \r in `user`/`home`
before interpolation. shq()'s shell quoting only defends the shell-lexer
hazard; a systemd unit file is parsed line by line, so a raw newline
starts an uncontrolled new KEY=VALUE directive regardless of quoting.
Both renderers now return Result<String, UnsafeUnitValue> and fail
closed; golden tests updated, new rejection tests added.

askpass.rs:
- blocker: run_with_password leaked a permanently-blocked blocking-pool
  thread when delivery's own internal timeout fired (fifo never opened)
  and ssh then exited gracefully on its own within the second wait —
  the one exit path that never called guard.poke_writer(). Now called
  unconditionally on entry to that branch.
- major: the same branch applied the caller's `deadline` a second time
  in full, letting a call block up to 2x the requested deadline. Now
  tracks a single start Instant and waits out only the remaining budget.
- minor: classify_output folded a signal-killed ssh process (status
  code() == None) into the "auth succeeded" Ok branch. Now classified
  as its own Other outcome.
- minor (TOCTOU): AskpassGuard's temp dir was created with default
  permissions then chmod'd to 0700 — a create-then-restrict window.
  Now created already-restricted via DirBuilder::mode(0o700).
- minor: the askpass helper script invoked `cat` via ambient PATH.
  Now uses the absolute /bin/cat.
- security: classify_output's catch-all Other branch forwarded ssh's
  raw stderr verbatim, untested for password leakage. Since the
  bootstrap flow's first connection is TOFU (accept-new) against an
  unverified host that legitimately receives the real password to
  authenticate it, that host could reflect the password back in a
  banner/diagnostic line. classify_output now takes the password and
  redacts any literal occurrence of it from the forwarded line; test
  coverage extended to exercise this branch.
- major: the live #[ignore]'d Docker test hardcoded port 12245 and
  relied on the developer's real, persistent ~/.ssh/known_hosts via
  StrictHostKeyChecking=accept-new, so a second run collided with the
  first run's pinned host key and failed with HostKeyMismatch instead
  of proving anything (reproduced; found stray port-12245 entries in
  the real known_hosts file, since removed). Fixed by scoping the test
  to an isolated, per-run known_hosts file — which in turn required
  splitting bootstrap_ssh_command into an option-building half
  (bootstrap_ssh_options) and the destination/remote-command append,
  since ssh only parses -o flags placed BEFORE the destination.
  Verified repeatable across three consecutive live runs.

cargo test --lib: 709 passed. Live Docker test run manually via
--ignored --nocapture (not part of the default suite): passed 3x in a
row. tsc --noEmit and pnpm test (vitest, 1777 tests): unaffected, run
clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
bootstrap-fixtures/: one Dockerfile (a stage per fixture, systemd as
PID 1, password-auth sshd, base ubuntu:20.04 = josty-cc) and fixture.sh
up|check|down|ssh with fixed ports 2231-2234 and documented throwaway
credentials; password ssh via SSH_ASKPASS_REQUIRE=force (no sshpass).
  A deploy + sudo (password)   B root password login
  C flightdeckd pre-installed as josty-cc's exact root-owned system unit
    (static musl binary from B0, unreachable relay)   D no sudo at all
check all: every fixture accepts its expected auth mode.

d-linger-experiment.sh answers D empirically (Ubuntu 20.04/24.04,
Debian 12, identical): without linger a --user unit dies ~10 s after
the session closes; setsid+nohup survives unless KillUserProcesses=yes;
a no-sudo user CAN enable-linger themselves (polkit set-self-linger:
yes), after which the --user unit survives logout and reboot.
docs/B6-FIXTURES.md records it with the installer fallback order.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… missing

Today, if flightdeckd isn't on the remote PATH, the SSH-spawned process exits
immediately (127) before ever completing the fd_attach handshake, so no
FdDetach is produced and the reconnect loop retries forever on a generic
"Connection to the server lost — reconnecting…" notice.

Add a narrow classifier, looks_like_missing_daemon(exit_code, stderr_tail),
requiring BOTH exit code 127 AND the last non-empty stderr line naming
flightdeckd (anchored to the command name, not a generic "command not
found"/"no such file" match, to avoid false-positiving on unrelated
remote-shell noise like MOTD or dotfile errors and permanently killing a
retryable blip). run_actor captures whether the just-closed transport ever
attached (had_attach) before reaping it, and on a match routes through D2's
reconnect_policy_for_reason table with a new "daemon_missing" entry
(terminal, with an actionable message), rendered as a process_exited notice
so the conversation visibly stops instead of spinning.

Adds a resolve_ssh_bin() test seam (mirrors the existing $TOSSE_CLAUDE_BIN
escape hatch) so an actor-level test can point a remote spawn at a fake
script without mutating $PATH, exercising the real spawn -> classify ->
single-terminal-notice path end to end.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…config lock

config.rs: Config::save (tmp + fsync + rename, mode 0600) and
Config::update (read-modify-write of the ON-DISK config) under
ConfigLock — an exclusive flock(2) on the sidecar config.json.lock, the
only thing that serializes the daemon against a separate
`flightdeckd init` process. init now takes it across its
exists-check and write (no more plain non-atomic fs::write).

SessionManager::new(cfg, registry, config_path) gains the live phone
access (tokens + tombstones) and relay_out. add_phone_token /
remove_phone_token persist first, then push authorize_phone /
revoke_phone when connected. Removed tokens are tombstoned in the config
(revoked_phone_tokens, capped at 16): the relay persists authorizations,
so a revoke sent while offline must be re-sent on every connect (C3).

registry: set_title_authoritative (unconditional, blank = no-op),
distinct from the fill-only set_title backfill.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
attach.rs: one-shot add_phone {token,label} / remove_phone {token}
(run via spawn_blocking — config file + cross-process lock) replying
fd_phone_added {ok,added} / fd_phone_removed {ok,removed}, or ok:false
+ error; replies never echo the token. AttachParams.title: once the
conversation is resolved and before the pump starts, a non-blank title
is recorded with set_title_authoritative. Client helpers
add_phone_client / remove_phone_client (an ok:false reply is an Err, so
the CLI exits non-zero).

main.rs: add-phone / remove-phone (--token - reads the secret from
stdin, keeping it out of the process list), whoami (config only, no
daemon, never the secret), attach --title.

Tests: socket-level (temp socket + attach::serve) for all one-shot
verbs and the title; tests/cli_e2e.rs drives the real binaries against
a real `flightdeckd run` with a fake claude.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…et_label

connect_once replays the manager's CURRENT phone access on every
(re)connect — revoke_phone per tombstone, authorize_phone per token —
then one {type:set_label, label}; all in one critical section under
the phones lock with the link's writer published in manager.relay_out,
so a concurrent add/remove lands in the burst or live on the link,
never in neither. A guard unpublishes the writer when the connection
ends (cancellation included) unless a newer link took over.

set_label is safe on a relay that predates it: unknown mac frames are
passed through to phones, and the PWA's handleFrame ignores unknown
types (flightdeck-remote web/app.js; PROTOCOL.md says as much).

Test: a local mock relay (tokio-tungstenite accept) — burst order,
set_label once per connect, add/remove mid-connection reach the relay
on the same link, a reconnect replays the updated state, relay_out is
cleared when the link ends.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
clousty8 and others added 28 commits September 19, 2026 02:58
…tach

Counter-verification of the R1 fix (CRM `1abfc028`) found a quick owner
Cancel-then-Start could reattach to the exact same (still dying) session_id —
`attach_or_reserve` hands back whatever is still registered while the backend
actor is mid-teardown. `selfCancelledSessionIdRef` was never cleared by `start()`/
`restart()`, so the OLD cancel's belated `ServerLoginResultEvent{reason:"cancelled"}`
for that reused id was wrongly swallowed as an echo of the prior intent, leaving
the newly (re-)attached view stuck on "Waiting for the sign-in link…" forever
instead of showing the neutral "cancelled from another panel" message. Fixed by
clearing the ref at the top of both `start()` and `restart()`, before their IPC
call — any earlier echo-suppression intent stops being relevant the moment this
instance establishes a fresh session view.

Also drops the now-dead `selfCancelledSessionIdRef` write from the unmount
cleanup (separate minor finding from the same counter-verification): the
listener effect that reads the ref tears itself down in the same unmount commit,
with no async gap in between, so no event could ever reach it after that write —
it implied an echo-suppression mechanism that could never actually run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…iring

Counter-verification finding (CRM `1abfc028`, test-honesty): the R2 regression
test for the host-rotation resume fix called `resolve_resume_machine` and
`sync_resume_request_to_machine` by hand instead of going through anything
`bootstrap_resume` itself calls — verified empirically that deleting
`bootstrap_resume`'s own one-line wiring call to the sync step left the full
`cargo test --lib bootstrap::orchestrator::` suite green, including this test.

A `#[tauri::command]`-level test remains out of reach (this crate's `tauri`
dependency has no `test` feature/dev-dependency override to build a
`mock_builder` app from, and faking every ssh round trip the full pipeline's
steps make would be a separate, much larger undertaking). Short of that, this
collapses `resolve_resume_machine` + `sync_resume_request_to_machine` into one
`resolve_and_sync_resume_machine`, called ONCE from `bootstrap_resume` — there
is no second call site left to silently drop, and the test now calls that same
merged function directly, which fails immediately if the sync step is ever
removed from it (re-verified the same deletion now fails the test).

Bindings regenerated to pick up the doc-comment reference change on
`bootstrap_resume`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Counter-verification residuals of the holistic-review fix wave: typed terminal event
for a cancelled shared Claude sign-in (R1), resume dials the machine's current address
after a rotation (R2), stale doc comments (R3) + their review fixes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…he shared bin resolver (B14)

PROBLEM 1: every remote `claude` lookup (pairing probe, B7 install-mode probe,
diagnose, sign-in) used a bare `command -v claude`/`claude ...` inside a
non-interactive ssh shell, whose PATH never includes `~/.local/bin` — exactly
where the official native installer puts `claude`. A genuinely-installed
server was reported as "claude is not installed" unless it had a hand-made
`/usr/local/bin/claude` symlink. Fixed by adding `resolve_claude_bin_expr`
(mirroring the existing `resolve_daemon_bin_expr`/`resolve_remote_daemon_bin`
resolver) and routing every remote claude invocation through it:
`ipc::commands::probe_script`, `bootstrap::connect::probe_script`,
`bootstrap::orchestrator::diagnose_script`,
`bootstrap::server_setup::claude_auth_status_cmd`/`claude_auth_login_cmd`.
Also gave the systemd USER unit template the same `Environment=PATH=` line
the system unit already had (missing before, so a user-level flightdeckd
could never spawn a `~/.local/bin`-only claude) and added a
`user_unit_missing_path` diagnosis fact so `InstallService`'s repair
re-renders a stale unit.

PROBLEM 2: the app never installed Claude Code for the user — the wizard
pipeline just failed into a generic "needs sign-in" state whether claude was
missing or merely signed out, and the legacy manual-pairing dead end told
users to run `curl ... | sh` (wrong: the official command pipes to `bash`).
Fixed by adding a new `StepId::InstallClaude` pipeline step right after the
probe (skipped once the shared resolver already finds a working claude,
fails the whole pipeline early otherwise — a server that can't run Claude
Code can't usefully continue), backed by
`bootstrap::server_setup::install_claude` (downloads
https://claude.ai/install.sh to a temp file with curl -fsSL, falls back to
wget, only runs it with bash on a complete download, verifies with the
resolved `claude --version` — see that function's doc for the citation
against the current Claude Code docs). `DiagnosisState` gains
`NeedsClaudeInstall`, distinct from `NeedsClaudeSignIn`, with its own
`RepairAction::InstallClaude` wired into `machine_repair` and
`repairSuggestionsFor`. `ClaudeSignInInline`'s gate (`claudeNeedsSignIn`) now
requires `claude_installed === true`, so it's never offered while claude is
missing. The legacy ticket path keeps its hard block but fixes the command
text and points at "Add a server" instead.

No new setting: installing Claude Code is part of the install flow the user
already started, mentioned via the step checklist alone.

Tests: resolver unit tests (PATH / ~/.local/bin / /usr/local/bin / absent /
space-in-HOME, real `sh` execution), a crate-wide regression test grepping
every remote script for a bare `command -v claude`, updated golden-string
and collapse_state tables, repairSuggestionsFor/claudeNeedsSignIn coverage,
a wizard step-list test, and a live Docker-fixture test (installs Claude Code
for real on fixture A, re-probes via the resolver with no symlink, diagnoses
NeedsClaudeSignIn, and checks the user unit's PATH via `systemctl --user show`).
All 21 bootstrap:: live tests pass; full suites green (cargo test --lib,
tsc --noEmit, vitest, test:scripts). Bindings regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, not sh)

The one-line pairing ticket's NOTE hint told the user to run
`curl -fsSL https://claude.ai/install.sh | sh` — the official native
installer pipes to `bash`, not `sh` (see B14,
bootstrap::server_setup::install_claude's own citation). Non-blocking text
fix only; this ticket line runs in the user's own interactive terminal
(unlike our ssh probes' hard pairing gate), so `command -v claude` is left
as-is here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…install log capture, dead InstallService repair)

- Blocker: RepairAction::InstallService for a confirmed pre-B14 user unit
  (missing PATH) was silently routed through install_service's generic
  conflict-detecting entry point, which adopts-and-does-nothing on any
  already-init'ed server (every real target). Add a dedicated
  install::repair_user_unit_path that goes straight to install_user_unit,
  and dispatch to it from repair()'s InstallService arm when diagnose
  confirms installed_as=User && user_unit_missing_path. Proven live against
  fixture A (simulated pre-B14 unit -> repaired -> diagnose + systemctl
  agree).
- Major: probe/diagnose only checked that `claude` exists and is
  executable, never that it actually runs — a broken/corrupted install
  was misdiagnosed as "installed, needs sign-in" instead of triggering a
  reinstall. PROBE_SCRIPT_BODY (ipc::commands + connect, both copies) and
  DIAGNOSE_SCRIPT_BODY now require `claude --version` to exit zero with
  non-empty output before treating it as present.
- Major: install_claude's installer stdout was unconditionally discarded
  (`>/dev/null`), so a failure inside the compiled `claude install`
  subcommand (as opposed to the wrapper script's own stderr-disciplined
  checks) left no diagnostic at all. Capture stdout+stderr and surface a
  joined tail of it on failure instead.
- Minor: fix a stale step_add_machine doc comment describing the pre-B14
  model where "claude missing" was non-blocking; it is now owned by the
  blocking StepId::InstallClaude step.
- Minor: mock install_service repair case now clears
  user_unit_missing_path, matching how install_claude already flips
  claude_installed.

Tests: cargo test --lib (1089 passed), targeted bootstrap:: run (220
passed), live Docker-fixture tests re-run for every touched script
(connect fixture c/d probe, orchestrator diagnose fixture a/c,
install_claude end-to-end on fixture a, install_service fixture a/c
regressions, and the new repair_user_unit_path live test) all green.
tsc --noEmit clean, pnpm test (2040 passed), pnpm test:scripts (5
passed). No #[tauri::command] signatures changed; bindings.ts
regenerated via export_bindings_regenerates_ts_client with no diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
B14 fix round 2 (CRM be1d47fc): the verification pass found that
repair_user_unit_path rewrote a pre-B14 user unit's Environment=PATH=
line and ran `daemon-reload && enable --now`, but that combination is a
restart-wise no-op on a unit that is already active — which is the ONLY
real-world case this repair ever fires for. Verified empirically against
Docker fixture A: the PID and /proc/<pid>/environ were unchanged while
`systemctl --user show -p Environment` already reported the new value,
so the repair reported success and diagnose said "fixed" while the live
daemon kept its stale, PATH-less environment and kept failing to spawn
`claude`.

repair_user_unit_path now forces an actual `systemctl --user restart
flightdeckd` after the reload, by reusing orchestrator::restart_daemon
verbatim (made pub(crate)) instead of hand-rolling a second restart
path. This gets the same busy-conversation guard the RestartDaemon
repair already relies on for free: if a fresh diagnosis can't confirm
zero busy conversations, restart_daemon returns
BootstrapError::DaemonBusy, which surfaces to the front as a normal
error (never a silent success) — the corrected unit file is still
written unconditionally (never touches the live process on its own),
so a busy server is left ready for the very next retry. Fresh-install
behavior (install_user_unit called directly, never through this
function) is unchanged.

Updated the live test
(live_repair_user_unit_path_rewrites_a_pre_b14_unit_the_generic_path_would_silently_adopt)
to assert the daemon's PID actually changed across the repair and that
the NEW process's own /proc/<pid>/environ carries the corrected PATH,
alongside (not instead of) the existing diagnose/systemctl-show checks
that the verification pass showed cannot tell a real restart apart from
a no-op.

Tests: cargo test --lib (1089 passed), the two named B14 live tests,
live_install_service_fixture_a, and the full bootstrap:: live suite
with --test-threads=1 (22 passed) against Docker fixtures (no josty-cc
contact); tsc --noEmit clean; pnpm test (2040 passed); pnpm test:scripts
(5 passed). No #[tauri::command] signature changed, so bindings.ts is
unchanged (confirmed via export_bindings_regenerates_ts_client + git
status).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…clamp

Whole-branch re-review of feat/claude-install (B14) surfaced five more
findings on top of fix round 2:

- repair_user_unit_path (blocker): the on-disk unit file was rewritten
  BEFORE the busy-conversation check that gates the actual restart. On
  a busy daemon this permanently erased the only diagnosis signal
  (user_unit_missing_path) that routes a repair call here at all and
  makes the front offer the repair — so a retry after conversations
  finished silently fell through to the generic install_service entry
  point instead of re-running this fix. Now diagnoses and bails out
  with DaemonBusy BEFORE writing anything when busy; the unit file
  (and therefore the diagnosis signal) is left untouched until the
  daemon is confirmed idle.

- install_claude (blocker): had no timeout anywhere in its ssh round
  trip, so a hang in curl/wget, the installer, or the final `claude
  --version` check could freeze the step (and this machine's
  ServerLocks slot) forever with no way to cancel. Wrapped in a 5
  minute tokio::time::timeout (generous for a real download+install),
  plus curl --max-time/wget --timeout on the download itself.

- connect::probe / ipc::commands::probe_remote (major): same missing-
  timeout gap as orchestrator::diagnose already guards against (B11) —
  both now share diagnose's own SSH_ROUND_TRIP_TIMEOUT (made
  pub(crate)), so a wedged remote shell (including one stuck in the
  B14 broken-install check's own `claude --version`) can't stall
  step_probe/repair/the "Add a server" pairing flow forever.

- install_claude's captured installer-failure log (major): capped to
  5 lines by the remote script but not to any byte length, and never
  run through this module's own strip_ansi() before reaching the UI —
  a single long or escape-sequence-laden line from the black-box
  `claude install` subcommand would reach BootstrapError unbounded/raw.
  Now stripped and clamped to 2000 chars.

- DiagnosisState::NeedsClaudeInstall UI copy (minor): said "not
  installed" even when collapse_state routes a present-but-broken
  claude binary into the same state. Reworded to "isn't working on
  this server" (and the matching repair reason), accurate for both.

Tests: cargo test --lib (1089 passed); tsc --noEmit clean; pnpm test
(2040 passed); pnpm test:scripts (5 passed); the two named B14 live
tests, live_install_service_fixture_a, and the full bootstrap:: live
suite with --test-threads=1 (22 passed) against Docker fixtures (no
josty-cc contact) — fixtures brought back down afterwards. No
#[tauri::command] signature changed, so bindings.ts is unchanged
(confirmed via export_bindings_regenerates_ts_client + git status).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
B14 (CRM be1d47fc): the server wizard installs Claude Code when it is missing
(official native installer, as the SSH user, never sudo), one shared remote resolver
for the claude binary (PATH -> ~/.local/bin -> /usr/local/bin), user unit PATH line
+ restart-on-repair, distinct 'Claude Code not installed/not working' diagnosis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
First real novice run (19/09, fresh Ubuntu server): the wizard failed at
'Set up the background service' with a crash-looping unit — the pipeline
started the service BEFORE RunInit, and `flightdeckd run` refuses to start
without the ~/.flightdeckd/config.json that init writes. The live tests never
caught it because each one ran init by hand before install_service.

RunInit now runs right after UploadDaemon (the probe still runs first, so our
own fresh config is never seen as a pre-existing conflict). The order lives in
one PIPELINE_ORDER const that build_pipeline is checked against, with tests
pinning upload < init < service and probe < init; the front's STEP_ORDER and the
mock sequence follow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… code, orphans, feedback

First real novice run (19/09, claude 2.1.278 over Tailscale): the sign-in failed
with 'the claude CLI changed its sign-in output' although the CLI output was the
expected one.

- The URL line AND the 'Paste code here' prompt arrived in ONE read; feed() made
  one transition per read, stopped at UrlReady, and the CLI then blocks on stdin so
  no further read ever came. feed() now advances to a fixpoint, and the driver
  remembers the URL itself (the state can jump past UrlReady).
- A code pasted before the prompt was recognized was silently dropped
  (submit_code only fires from AwaitingCode); it is now held and written as soon
  as the prompt shows up.
- With no pty, killing our local ssh never stopped the remote claude auth login
  (one orphan per attempt; it also survives SIGTERM). The login command now kills
  a stale login and records its own PID; cancel/failure/supersede kill exactly
  that PID (checked against /proc/<pid>/cmdline, SIGKILL, never by name).
- The panel showed no sign of life while waiting for the link or while the CLI
  exchanged the code: spinner + 'Checking the code with Claude…' state.

Tests: one-read URL+prompt fixpoint, URL remembered across reads, login command
shape; live (Docker fixture A): a code pasted the instant the URL shows reaches
the CLI (rejected by the CLI itself, never the recognition timeout) and no remote
login process is left behind after cancel or failure; front: loading states.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…-in'

Found in the novice run (19/09): after a successful Claude sign-in the server
card kept showing 'Needs Claude sign-in' + 'Claude signed in: Unknown', and the
Sign in button looped. `claude auth status --json` is PRETTY-PRINTED (12 lines
on 2.1.278) while every diagnose marker is read as one line, so only '{' was
parsed. The diagnose script now flattens it (JSON strings never hold a raw
newline). Test runs the real diagnose script against a fake claude that
pretty-prints — it fails without the fix.

Also: the server card showed the status chip twice (next to the name and again
above the facts); DiagnosisSummary gains showHeadline, off in the card.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…s first message

Found on the first real remote conversation (19/09): the very first message of a
fresh remote conversation flashed 'The server has no turn running — your last
message may not have been delivered' half a second before its turn ran. The
message is written on the same link right behind the daemon's fd_attach, whose
busy:false is simply older than it. The core now tracks whether a user message
was written on the CURRENT link (reset when a reconnect swaps the link) and only
treats an idle daemon as 'message lost' for messages written on an earlier link.
Tests cover both sides (the first one fails without the fix).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The legacy ticket/paste flow (Settings → Control → Remote servers → "Use a
command instead") used to end by calling the old `addMachine` IPC, which
hard-blocks with "claude is not installed…" on a fresh server — a real
novice's Ubuntu box dead-ended there while the primary password path
installed everything fine.

Once a ticket is confirmed, `LegacyPairing` now hands the connection details
to `PrimaryBootstrap` via a new `onInstall` prop instead of pairing on its
own. `PrimaryBootstrap` seeds its form from that prefill and auto-starts
`bootstrap_server` with no password on mount (guarded against StrictMode's
double-invoked effects) — the pasted command already authorized this Mac's
pending key on the server, and `bootstrap_server`'s own `InstallKey` step
finds that key already works and skips straight past it, so the SAME full
install (Claude Code, the daemon, persistence, sign-in) runs as the guided
path. If the key turns out not to work, the pipeline's own error shows and
the primary form's password field lets the user retry.

Also renames "Test & pair" to "Install" (confirm + manual-entry stages),
adjusts the command-stage text to describe the full install rather than
just pairing, and drops the now-dead `addMachine` call and "done"/
matched-existing stage from `LegacyPairing` (the `addMachine` store action
and Rust command are untouched — still used elsewhere).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The paste-a-command server path now hands off to the full install pipeline
(no password: the pasted command authorized the pending key) instead of the
legacy add_machine that dead-ended on a fresh server.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gather the conversation's state into a panel at the far right so the
header carries actions only:

- TOSSE task card: status menu + the CRM's status ladder, assignee
  picker (new tosse_set_task_assignee command, PATCH /api/v1/tasks/:id),
  collapsible tickable subtasks; clicking the card opens the task.
- Goal, todo list and artifacts (moved out of the composer / todo bar).
- Session footer: stream (on/restart/off) and worktree (+ Manage),
  moved out of the header.
- Closed: a one-line Goal · Todo summary above the composer.
- Toggle in the header and on ⌘I, with the chord always shown.
- Docks when there is room, steps aside otherwise (floats when asked).
- Display pref "Conversation side panel" (on) restores the old layout.
- The layout-orientation button now also shows while an artifact or a
  TOSSE task occupies the side region.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# Conflicts:
#	src/features/conversation/ConversationPane.tsx
#	src/ui/kit.tsx
…oves into the composer

The sidebar's conversation rows drop the leading status dot: the row itself IS the state,
as a tinted pill — green running, green with violet washing in on its right for background
work, amber needs-you, blue to review, red error; idle and stopped rows stay plain (stopped
a touch dimmer). A second line carries the thread's working dots and a counter: ticking
while the agent works (centiseconds under a minute, then "3m 07s", then "1h 02m"), frozen
on how long the turn ran once it stopped on a state, gone once the row is calm again. Where
"mark as seen" cannot apply — a questionnaire or a permission, which are answered in the
thread — the row shows a non-interactive glyph (? / key) instead of the check. Selected is
deliberately louder than hover: hover lifts the base one step, selected fills the pill in
the state's colour with a near-solid outline, an outer ring and a white name.

The review / question / error / background status leaves its full-width bar above the
composer and becomes a header band INSIDE the composer card, whose border takes the state
colour (a soft halo, breathing for a question), keeping "Mark as seen" (Cmd+Enter) and
"Continue".

Both are display preferences, ON by default, so the previous look is one toggle away:
Display -> Appearance ("Tinted conversation rows", "Status inside the composer"), plus
Display -> Durations ("Time on sidebar rows") to hide the counter alone.

Supporting pieces: ui/liveElapsed.ts (pure formatters + ONE shared clock that only runs at
25 fps while a counter still shows centiseconds, and pauses when the window is hidden),
agent/rowTiming.ts (pure live/paused/none rules), agent/status.ts::questionExcerpt (the
question itself, not the message's opening lines), and the timing the store now keeps
(lastTurnStartedAt / lastTurnEndedAt / awaitingSince, plus firstSeen in the background task
store so background work counts from its own start, not from a later turn).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…in a native webview

A typed artifact's page belongs to its TYPE, hosted on claude.ai; the artifact itself
only carries data files (canvas.json, .dc.html). Rendering one of those as a page is
what showed a broken screen. They are now detected and shown hosted.

The hosted view is a native CHILD webview (`artifact_host/` + `artifactHost.ts`) laid
over the side panel: claude.ai refuses to be framed (X-Frame-Options + a Cloudflare
challenge) and a private artifact needs the user's session. It also becomes the route
for any artifact whose local temp file is gone, and for a link to another
conversation's artifact — all of it behind a new "Show hosted artifacts in Flight Deck"
setting (ON), so the previous browser behaviour stays one click away.

Load-bearing details:
- ⚠️ a child webview makes `get_webview_window("main")` return None — the UI zoom and
  the Dock bounce now go through `get_window`/`get_webview`;
- a native view paints above all HTML, so the host hides on any overlay (portal
  intersection + hit-test grid), and follows the panel's box and the UI zoom;
- wry reports no failed navigation, so a load watchdog is the only failure detector;
  a late page load overrules its verdict;
- pop-ups stay in-app only while a sign-in is under way (`Allow`, never `Create` —
  building a Tauri window in WebKit's callback aborts the app), non-web schemes are
  refused and reported, and per-page link volume is capped;
- passkeys-on-this-Mac and password AutoFill need Apple's browser entitlement, which a
  self-signed app can't hold: the sign-in strip says so and takes a pasted sign-in link.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# Conflicts:
#	src-tauri/src/lib.rs
#	src/features/settings/settingsSearch.ts
#	src/ipc/bindings.ts
#	src/ipc/mock/mockBindings.ts
…the three TOSSE tasks

An adversarial review of the drag-and-drop, failed-background-task and
conversation↔TOSSE-task commits (2184af6, 71f6118, 59935e1) confirmed 12
defects. All are closed here.

CRITICAL — `link_tosse_task` turned an agent-supplied `task_id` into an
arbitrary authenticated GET on the CRM. Nothing validated the id, and
`api_get` built its URL by concatenation, so the `url` crate resolved `..`
at parse time: `../clients` reached /api/v1/clients carrying the human's
`tosse:app` Bearer, and the first 300 characters of the body came back to
the agent inside the error. Closed on both sides — `api_url` now builds
every URL with `Url::parse` + `path_segments_mut` and refuses a `.`/`..`
segment (this also covers `api_write`, same defect class on the PATCH
paths), `task_detail` refuses a non-UUID id before a token even exists, the
MCP schema declares the pattern, and the front validates + normalizes the id
once and no longer echoes the response body. This one shipped in v2.4.0.

HIGH
- `read_conversation` surfaced only `task_failed` and dropped every fatal
  notice, so a conversation polling a crashed one saw the prompt with no
  answer and concluded it was still thinking. Every error-bearing notice is
  reported now, through `NOTICE_ERROR_HEADINGS`, and system lines no longer
  spend the caller's `max_turns` budget.
- A read that never settled held a conversation's send lock forever, with a
  silently dead Enter key. Reads are bounded per file and per batch, and the
  composer says it is reading.
- A Finder drop built its mentions against `conv.cwd`, the --resume anchor,
  not the live worktree cwd: after an EnterWorktree the agent resolved a
  DIFFERENT file with no error. Drop and the "+" picker both use
  `effectiveCwd`, resolved after the native dialog rather than at render.

MEDIUM / LOW
- `link_tosse_task` could blow the hub's 30 s deadline and still write the
  link the agent was told had failed; the CRM reads are bounded first.
- A batch finishing after its conversation was deleted resurrected its
  attachments and rewrote a persisted draft; callers pass a liveness
  predicate and only this batch's own entries are cleaned up.
- Overlapping batches overwrote each other's errors; they merge.
- A dropped path containing a space reached the agent as several tokens; it
  is now one inline-code token. Widening `parseFileMention` to take it back
  was tried and reverted: its segment class is ASCII-only so the real
  "Capture d'écran … à 14.03.21.png" still failed, a user turn renders as
  plain text so no chip was at stake, and it flagged `/usr/bin/ls -la` as a
  file. The wire is what this fixes, and only the wire.
- An agent-sourced re-link overwrote a CRM-verified status with null; ids
  are compared on a canonical form, and unverified input no longer erases
  verified state.
- `onPaste` was a third divergent copy of the attach pipeline.
- The unlink × was unreachable by keyboard, as were the row actions one row
  up (`.rowActs:focus-within` can never match inside `display: none`), and
  focus fell to <body> after every unlink.
- A `task_failed` notice landing at a turn boundary now advances
  `replayAnchor`, so a message sent from the phone afterwards renders below
  the failure instead of above it.

Deliberately unchanged: `task_failed` stays a soft notice and keeps folding
into clean output. That was reported twice as a silent error and is a
product decision — a background task failing is benign, Claude handles it
through its own <task-notification>.

tsc clean, 2253 front tests, 1109 Rust tests, no bindings drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gitguardian

gitguardian Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

️✅ There are no secrets present in this pull request anymore.

If these secrets were true positive and are still valid, we highly recommend you to revoke them.
While these secrets were previously flagged, we no longer have a reference to the
specific commits where they were detected. Once a secret has been leaked into a git
repository, you should consider it compromised, even if it was deleted immediately.
Find here more information about risks.


🦉 GitGuardian detects secrets in your source code to help developers and security teams secure the modern development process. You are seeing this because you or someone else with access to this repository has authorized GitGuardian to scan your pull request.

@Alex375
Alex375 merged commit 8815c90 into main Sep 20, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants