Skip to content

fix(firecracker): a guest that dies fails its agent-ready waiters at once, with the console tail - #1610

Merged
nikhilunni merged 1 commit into
mainfrom
fix/fc-dead-vm-wait
Oct 8, 2026
Merged

nikhilunni merged 1 commit into
mainfrom
fix/fc-dead-vm-wait

Conversation

@nikhilunni

Copy link
Copy Markdown
Contributor

Summary

A Firecracker guest that dies before agentd dials ready now fails its waiters immediately, with the console tail in the error, instead of sitting out the 180 s deadline with a generic message.

The problem

Reproduced on a KVM host while debugging the swap snapshot test: the capture guest panicked 3 s after boot, the supervisor pruned the sandbox within 5 s and deleted the jail, but the ready-dial listener task kept the watch sender alive. wait_agent_ready saw neither "ready" nor "closed", waited the full 180 s, and failed with "agentd did not dial ready port". The panic text was only recoverable by polling the jail before the prune.

The change

  • The ready-dial listener is owned by the LiveSandbox (AgentReadiness: receiver, task handle, death note).
  • Prune and destroy read the last 4 KiB of firecracker.log before removing the jail, store it as the death note, log it at WARN on an unexpected death, and abort the listener so the watch closes.
  • A closed watch maps to guest died before agentd dialed ready port with the console tail. The ready fast path and the 180 s backstop for a hung guest are unchanged. start_agent and base capture get the behaviour through their existing calls.
  • read_tail is UTF-8 tolerant (a tail that starts mid-character no longer drops the rest).

Tests

  • Unit (no KVM): two pending waiters receive the panic marker after the supervisor's real teardown callback; the same with no console file; explicit destroy; mid-character tail.
  • Pooled: base capture with a dead guest fails within one second and still destroys the VM.
  • KVM (lifecycle.rs, already in the unprivileged FC lane): a fixture whose init exits immediately fails wait_agent_ready within 15 s with "Kernel panic" in the message.

Validation

just check green (2766 tests); Linux cross clippy for firecracker and host-agent clean; Codex adversarial review approved with no findings.

🤖 Generated with Claude Code

https://claude.ai/code/session_01UHxL6gj4o8EvxYpgtwEWaM

…once, with the console tail

`wait_agent_ready` sat out its full 180 s deadline when the guest died
first: the supervisor pruned the sandbox and deleted the jail, but the
ready-dial listener kept the watch sender alive, so waiters saw neither
"ready" nor "closed". A base capture whose guest panicked at boot hung
with no diagnosis, and the console log was gone before anyone read it.

The listener is now owned by the sandbox. Prune and destroy read the last
4 KiB of `firecracker.log` before the jail is removed, store it as the
sandbox's death note, log it at WARN on an unexpected death, and abort
the listener. A closed watch becomes "guest died before agentd dialed
ready port" with the console tail attached. The ready fast path and the
180 s backstop for a hung guest are unchanged.

Tests: two pending waiters receive the panic marker after the supervisor's
real teardown callback; the same without a console file; explicit destroy;
a tail that starts inside a UTF-8 character; pooled base capture propagates
the death within one second and still destroys the VM; a KVM lifecycle test
boots a fixture whose init exits and asserts the failure names the kernel
panic within 15 s.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHxL6gj4o8EvxYpgtwEWaM
@nikhilunni
nikhilunni force-pushed the fix/fc-dead-vm-wait branch from c489a7d to 7dda8c2 Compare October 8, 2026 02:47
@nikhilunni
nikhilunni merged commit e244a4a into main Oct 8, 2026
30 checks passed
@nikhilunni
nikhilunni deleted the fix/fc-dead-vm-wait branch October 8, 2026 03:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant