Skip to content

feat(coord): the durable teleport state machine; TeleportSession and RetireHost plan it (ADR 0123 B) - #1591

Merged
nikhilunni merged 1 commit into
mainfrom
feat/teleport-machine
Oct 6, 2026
Merged

nikhilunni merged 1 commit into
mainfrom
feat/teleport-machine

Conversation

@nikhilunni

Copy link
Copy Markdown
Contributor

Problem

Issue #1549's control-plane half (ADR 0123 findings 4, 7, 8, 9, 10, and the UI teleport). Three modules each owned part of a move with unfenced writes, in-memory phase state, detached tasks and a give-up path to Idle: evacuation.rs wrote host, sandbox and state in three separate statements after the restore; evac_resumer.rs burned a 20-attempt budget confirming a source destroy on a host that no longer existed and dropped the session to Idle; live_migration.rs ran a synchronous verb with a detached finalize, a process-local host gate, a rollback that declared Active while the source stayed paused, and an illegal Evacuating → Failed edge. Destination capacity was checked, never reserved. The UI's EvacuateSession ignored its target, never attempted a live move, and hid errors behind an optimistic status flip.

Change

One durable machine. session_teleports (C1) is the journal and the reservation; OpKind::Teleport is driven by the existing session_ops executor (claim, fence, heartbeat, reclaim); teleport.rs is the one driver for every entry point. Phases admitted → captured → restored → committed → attached → done, with rolling_back → aborted and failed; every step ends in one fenced CAS and a successor resumes from the row.

  • Admit reserves the destination inside the same locked placement transaction create/queue use, with a per-destination open-teleport cap; NoFit leaves the session Active on the source (the don't-strand guard, now durable). Live is chosen when every candidate can serve post-copy; a host refusal (ADR 0112 swap) downgrades the row to snapshot in place. No feature flag.
  • Capture holds the source paused: new snapshot_hold on the host keeps the VM paused after capture and does not re-arm swap; resume re-arms once; destroy clears the hold. Without this the source kept emitting after the capture point and its events collided with the destination's replays under the same epoch and seq. A live move persists its full presetup result (live_payload) in the same CAS as export_id, so a successor never calls presetup twice.
  • Commit is one statement: host, sandbox, binding_epoch + 1, phase. No tombstone (a live source is still the page server).
  • Attach binds at the new generation, starts the agent, waits for the first event of that generation (C5), then Evacuating → Active. A deterministic spawn failure rests at Created and the move still releases the source.
  • Release is a destroy ack or a dead/retired source host. No budget, no Idle fallback. Rollback releases a live export through migration_abort and a held snapshot through resume, and stays rolling_back until the source answers.

Entry points. RetireHost plans teleports for its residents before evaluating the grant (the C1 TODO). AdminDrainHost cordons as admin and plans without awaiting moves. TeleportSession{session_id, target_host?} replaces EvacuateSession: synchronous admission under the op claim, honors the target under the full 2D, cordon and capability checks, returns {teleport_id, kind, dest_host_id}, FAILED_PRECONDITION with the reason otherwise. teleport_finished is a session event; the web shows it instead of flipping status optimistically.

Deleted. evacuation.rs, evac_resumer.rs, live_migration.rs, the evacuation types, the attempt budget, the teleport pin columns (migration 0122), MIGRATION_GATE, walk_back_to_active, parachute_or_kill, the drain JoinSet and placement preview, the unlocked reserved pick, the Evacuating target of the evict pipeline, ENGRAM_LIVE_TELEPORT.

Test

Conformance (sim + Postgres): admission reserves under lock (NoFit, SessionNotActive, Fenced, Conflict, cordoned dest, per-dest cap, pinned no-fit), phase CAS fenced and legal with payload preservation, one-statement commit (no tombstone, reservation moves), release entombs the source once, rollback keeps the dest reserved, dead-host orphan skips machine-owned sources, fail/abort fenced. Scripted scenarios on sim and live Postgres: end to end, crash at every phase resumed by a successor, presetup payload reuse, rollback pending until the source answers, live rollback through migration_abort, lost live export rolls back, live refusal downgrades, source dead at release, fenced step stops, no-fit stays Active, two per destination, retire grants only after an empty heartbeat then DeleteHost. Unit: attach waits for the generation, attach failure rests at Created and still releases, snapshot_hold stays paused without re-arm, resume re-arms once. DST drives the scanner with capacity and quiescence oracles. e2e (KVM lane): e2e_teleport_session_honors_target_host, e2e_retire_host_relocates_and_grants.

Local: clippy on 11 crates, 1320 unit tests, 72 PG-conformance, 160 live-PG (including the new teleport_live_pg), just check 2679, Linux cross clippy, web tsc/tests/build, orchestrator and CLI typecheck. KVM e2e runs in CI.

Stacked on #1590 (C3). Host follow-up noted in the ADR: make the export lifetime teleport-driven instead of TTL-driven.

🤖 Generated with Claude Code

@engrams-agent

engrams-agent Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

👀 engrams is reviewing 97fbc2f.

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

The latest Buf updates on your PR. Results from workflow CI / buf (pull_request).

BuildFormatLintBreakingUpdated (UTC)
✅ passed⏩ skipped⏩ skipped❌ failed (9)Oct 6, 2026, 5:49 PM

@engrams-agent engrams-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Engrams review

Verdict: 5 findings included in this summary.
Severity: Critical 0 · High 0 · Medium 4 · Low 1
Categories: 🩺 Stability & Availability: 2 · 📐 Maintainability & Code Quality: 2 · 🗄️ Data Integrity & Integration: 1

View the full engrams review

Findings on the review page

crates/engram-coordinator/src/api/admin.rs:L406 — Plain admin drain silently leaves Active sessions on the cordoned host, with no retry and a false "planned" report

WHAT: admin_drain_host_core only cordons the host (via set_host_cordon, which does not set retire_requested_at) and then fire-and-forgets per-session teleport ops through plan_host_teleports(AdminDrain); a session that cannot be relocated is still reported in DrainHostResponse.planned yet never actually moves, and nothing ever retries it.

WHEN: Two routes reach the same stranded outcome, sharing one root cause — a cordon-only drain has no durable backstop:

  1. Fleet at capacity (no cancellation needed). plan_host_teleports enqueues a deferred teleport op for every Active session and counts each in report.planned. When the op runs, drive_inner calls steps::admit(None, AdminDrain); if placement returns TeleportAdmitOutcome::NoFit, drive_inner (teleport.rs:260-267) returns OpOutcome::Done. The op completes "successfully", no session_teleports row is created, the session stays Active on the cordoned host, and the operator sees it listed under planned — indistinguishable from a real in-flight move.
  2. Request cancelled mid-plan. plan_host_teleports_with_limit is awaited inline in the gRPC/HTTP handler (the old code ran this fan-out on a detached task — the deleted comment: "This guard is preserved exactly" — but the new code dropped it). If the client disconnects mid-loop, sessions after the cut-off never get an op enqueued.

Neither case self-heals: unlike RetireHost (which sets retire_requested_at, so run_once's list_retiring_hosts sweep re-plans the host every tick), a plain drain leaves retire_requested_at NULL, so list_retiring_hosts (WHERE retire_requested_at IS NOT NULL) never re-sweeps it, and reconcile deliberately excludes Evacuating/resident Active mid-move from strike accounting. The operator believes the host is drained and removes it, killing the still-resident sessions.

Trigger likelihood: routine

crates/engram-coordinator/src/teleport.rs:L228 — Teleport ops have no attempt-budget ceiling, so a persistently-failing move retries forever and permanently consumes a destination slot

WHAT: The Teleport verb never bounds its retries: drive turns every Err from drive_inner into OpOutcome::Retry (teleport.rs:223-227), and the step bodies return RetryAfter unconditionally (e.g. rollback's host.resume/migration_abort failure → 5 s retry at 1147-1152; release's source-destroy failure → 10 s retry at 1064-1069). The only attempts check in the whole file is the narrow ctx.op.attempts <= 3 on one restore-error class (535). Every sibling verb has an explicit ceiling that routes to a terminal fallback — RESUME_MAX_ATTEMPTS=60, EVICT_MAX_ATTEMPTS=20, CREATE_BOOT_MAX_ATTEMPTS=30, QUARANTINE_EVICT_MAX_ATTEMPTS=3 (session_verbs.rs).

WHEN: A source or destination host that is reachable enough to stay out of Dead/Retired but whose resume/migration_abort/destroy RPC keeps failing (a wedged-but-heartbeating host-agent, a stuck sandbox) drives the row into RollingBack (or leaves it at Attached) and keeps failing:

  1. The session is stuck in Evacuating indefinitely — the user cannot use it — because rollback can only finish by first resuming the source back to Active, which keeps failing.
  2. The session_teleports row stays in a non-terminal phase, and teleport_admit's cap query counts every phase NOT IN ('done','aborted','failed') row against max_open_per_dest (default 1). So a single stuck row permanently consumes the destination host's teleport-receive slot, silently blocking all future teleports to that host, with no metric or terminal state to signal it.

source_gone() only rescues the row if the source host independently transitions to Dead/Retired; a host that merely misbehaves never triggers that.

Trigger likelihood: plausible-fault

web/src/test-utils.tsx:L74 — adminDrainHost test mock still returns the old (removed) AdminDrainHostResponse shape

WHAT: The regenerated AdminDrainHostResponse dropped evacuating/failures and now has planned/descended/skipped (fleet.proto + web/src/gen/.../fleet_pb.ts), but the adminDrainHost mock in testTransport still returns { hostId: "", evacuating: [], failures: [] }. The sibling evacuateSession → teleportSession mock on the next line was migrated correctly; this one was missed.

WHEN: The mock is an object literal returned from a typed connect-es (@connectrpc/connect v2 / @bufbuild/protobuf v2) service implementation, so evacuating/failures are excess properties not on the new message init shape and planned/descended/skipped are absent. The most likely outcome is a tsc -b excess-property error that fails pnpm build in the web CI lane; if TS does not flag it (e.g. a looser inferred return type), it is instead a silently-wrong fixture that any future test reading .planned/.descended/.skipped will get undefined from. Either way the response contract is now duplicated between the proto and this stub with only one side updated.

Trigger likelihood: routine

deploy/migrations/0122_teleport_machine.sql:L5 — Sessions mid-evacuation under the old scanner are permanently stranded in Evacuating after this deploy

WHAT: This PR removes the only driver for status='evacuating' sessions (deletes evac_resumer.rs, and the list_evacuating_sessions/bump_evac_attempts trait methods, and drops evac_attempts + idx_sessions_evacuating), replacing it with a scanner that drives relocation solely from session_teleports rows (run_once → list_open_teleports). Any session already sitting in Evacuating at deploy time with no matching session_teleports row has nothing left to advance or fail it.

WHEN: At the base revision, evacuations were driven by evac_resumer, which swept sessions WHERE status='evacuating' directly and had no session_teleports row for that session. If the new binary takes over while a session is mid-evacuation (a rolling coordinator deploy overlapping a host drain / dead-host relocation — both routine operations), that session is left Evacuating forever:

  • the new run_once only enqueues sessions it finds via list_open_teleports, and this one has no open teleport row;
  • dead_host no longer routes into Evacuating and does not recover sessions already in it;
  • reconcile deliberately excludes Evacuating from missing-sandbox strike accounting ("mid-move"), so it is never flipped to HostLost.

The session is stuck and unusable until manual intervention; the dropped evac_attempts column means the old logic cannot even be resurrected. I found no pre-deploy drain gate or boot reconciler for orphaned Evacuating rows in the diff; if the team enforces "zero evacuating sessions before applying 0122" operationally, that is the mitigation, but it is not visible in the change.

Trigger likelihood: plausible-fault

crates/engram-coordinator/src/teleport.rs:L536 — Bundled low: misleading failure label and stale comments left by the evac→teleport rename

WHAT: A bundle of trail/comment items from the refactor, each of which misleads a future reader or operator (grouped per the one-low-finding rule):

  1. teleport.rs:535-536 — after the restore-retry budget, finish_live_restore discards the real error and records the fixed literal "dest_lost_after_blackout" regardless of the actual cause (OOM, corrupt manifest, backend bug). This label is then surfaced in SessionDiagnostics and the teleport_finished.error field, so an operator debugging a failed teleport reads a specific-but-false root cause. The old live_migration.rs preserved the real error text end to end.
  2. session_verbs.rs:31 — the CheckpointFinalize/Teleport comment still points recovery at "the ADR 0018 parachute / evac scanner for a torn teleport"; that scanner (evac_resumer) is deleted in this PR.
  3. crates/engram-host-agent/src/pooled_backend.rs:812 — the doc block describing inflight_snapshots ("Cleared by commit_snapshot/abort_snapshot; … orphan dir leak …") now sits directly above the newly-inserted held_swap_policy field, so it attaches to and mis-describes held_swap_policy, while inflight_snapshots is left undocumented.
  4. crates/engram-protocol/proto/engram/app/v1/fleet.proto:45 — the TeleportSession RPC still carries the comment // ADR 0018 async evacuation (POST /api/admin/sessions/:id/evacuate).; that HTTP route and evacuate_session_core were deleted (the CLI now calls the gRPC RPC directly).

WHEN: The next operator or maintainer trusts a label/comment that no longer matches the code — misdiagnosing a teleport failure (1), looking for a deleted recovery path (2), reasoning about the wrong field's lifecycle (3), or hitting a removed route (4).

Trigger likelihood: routine

nikhilunni added a commit that referenced this pull request Oct 6, 2026
…ttles orphaned Evacuating rows

Review findings on #1591:

- `AdminDrainHost` was a cordon plus a one-shot plan: a resident that did
  not fit was reported as planned and never retried. It is now the
  retirement request with `owner = admin`, shared with `RetireHost`
  through `request_host_retirement_core`; the scanner re-plans it every
  tick and grants it when the host is empty. `UncordonHost{admin}` cancels
  the request before the grant.
- Migration 0122 settles sessions the retired evacuation scanner left in
  `evacuating` with no `session_teleports` row: tombstone the bound
  sandbox, then Idle with a recoverable snapshot and Dead without one.
- The restore-budget failure keeps the real error behind the
  `dest_lost_after_blackout` label.
- Stale comments: the Teleport verb's recovery note, the `inflight_snapshots`
  doc block, the `TeleportSession` proto comment (generated code follows).
- The web test mock returns the current `AdminDrainHostResponse` shape.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01965DMBwLXzE9baCmj1Wp8Q
@nikhilunni

Copy link
Copy Markdown
Contributor Author

Addressed in b937d23 and 9baba34:

  • Admin drain strands sessions (MEDIUM): AdminDrainHost is now the retirement request with owner = admin, sharing request_host_retirement_core with RetireHost. The scanner re-plans the host every tick, a resident that does not fit blocks the grant as a visible bound_sessions blocker, and UncordonHost{admin} cancels the request. The one-shot cordon-plus-plan path is gone.
  • No attempt ceiling (MEDIUM): by design (ADR 0123 B, invariant D): a rollback never declares Active until the source acks, and a release never gives up on a live source. The destination slot a stuck row holds is the honest state of a wedged source host; the row and its blocker are visible on GetHost.retirement and teleport_finished never fires early. 9baba34 tightens the one escape that was wrong: a consumed export (migration_abort → NotFound) now still requires a resume ack, and a destroyed source sandbox fails the move with source_lost_during_rollback.
  • Web mock shape (MEDIUM): fixed.
  • Orphaned Evacuating rows at deploy (MEDIUM): migration 0122 now settles sessions the retired scanner left in evacuating with no session_teleports row: tombstone the bound sandbox, Idle with a recoverable snapshot, Dead otherwise.
  • Low bundle: the restore-budget failure keeps the real error behind dest_lost_after_blackout: …; the Teleport verb comment, the inflight_snapshots doc block and the TeleportSession proto comment are corrected (generated code regenerated).

nikhilunni added a commit that referenced this pull request Oct 6, 2026
…RetireHost plan it (ADR 0123 B)

Rebuilt onto the rebuilt readiness branch after the squash merges of #1583-#1586; content unchanged. Includes the review fixes for #1591.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01965DMBwLXzE9baCmj1Wp8Q
@nikhilunni
nikhilunni force-pushed the feat/teleport-machine branch from 72069fa to dd69420 Compare October 6, 2026 17:01
@nikhilunni

Copy link
Copy Markdown
Contributor Author

Adversarial review round (my pass + three Codex Astra passes) landed in dd69420, on top of the branch rebuilt onto the squashed main:

  • Every teleport write locks the session row under its current_epoch first (lock_teleport_session_at_epoch), so a reclaimed driver can no longer land a phase CAS through a statement-snapshot epoch read.
  • teleport_settle replaces teleport_fail: session terminal state, source tombstone, failed row and both events in one fenced transaction. fail_move and the lost-source settle share settle_target (dead-host predicate; Failed instead of the illegal Created→Dead).
  • The source keeps the move's budget in committed|attached (source_reserving_phases).
  • The heartbeat's stably-unbound cleanup never entombs a sandbox an open move names, nor an unbound sandbox on a host a move is restoring into; the host's ownership answer says owned for both endpoints of an open move, so the export TTL sweep cannot destroy a post-commit source under a draining destination.
  • A deleted host row counts as gone; a consumed export on commit proceeds to the destroy ack; deterministic HarnessSpawn failures share resume's classifier.
  • Host: snapshot_hold persists a held-source role (a host-agent restart keeps the VM frozen); resume clears any source role; the detached swap re-arm runs under the capture lock; the abort's post-consumption cleanup runs in a detached task; a failed post-restore relocation destroys the VM before its NBD state is dropped.

ADR 0123's divergence log carries the full list. Local gate: just check 2704 passed, live-PG lane 2822 passed, Linux cross-clippy on host-agent.

@nikhilunni
nikhilunni added this pull request to stack #1595 October 6, 2026 17:38
@nikhilunni
nikhilunni force-pushed the feat/teleport-machine branch from dd69420 to eb37d82 Compare October 6, 2026 17:40
Base automatically changed from feat/harness-generation-readiness to main October 6, 2026 17:41
…RetireHost plan it (ADR 0123 B)

Rebuilt onto main after the squash merge of #1590; content unchanged. Includes the review fixes (fenced teleport writes, one-transaction settlement, source reservation, move-owned sandboxes, held-source role).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01965DMBwLXzE9baCmj1Wp8Q
@nikhilunni
nikhilunni force-pushed the feat/teleport-machine branch from eb37d82 to 97fbc2f Compare October 6, 2026 17:42
@nikhilunni
nikhilunni merged commit 815e965 into main Oct 6, 2026
3 checks passed
@nikhilunni
nikhilunni deleted the feat/teleport-machine branch October 6, 2026 17:44
nikhilunni added a commit that referenced this pull request Oct 6, 2026
Rebuilt onto main after the squash merge of #1591; content unchanged.

Remove the blind sandbox setter from MetadataStore, both stores, and all
mocks. Use the surviving fenced binding writes in test fixtures. Make
reconcile compare the exact current binding, including None. Keep
assign_session_host for dead-host cleanup. Preserve strike reset and
live-manifest cleanup in fenced assignment; extend conformance coverage.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01965DMBwLXzE9baCmj1Wp8Q
nikhilunni added a commit that referenced this pull request Oct 6, 2026
Rebuilt onto main after the squash merge of #1591; content unchanged.

Remove the blind sandbox setter from MetadataStore, both stores, and all
mocks. Use the surviving fenced binding writes in test fixtures. Make
reconcile compare the exact current binding, including None. Keep
assign_session_host for dead-host cleanup. Preserve strike reset and
live-manifest cleanup in fenced assignment; extend conformance coverage.


Claude-Session: https://claude.ai/code/session_01965DMBwLXzE9baCmj1Wp8Q

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant