Repository navigation
fix(resume): an unusable memory image recovers by disk-only cold boot, not Dead - #1613
Merged
Merged
Conversation
…, not Dead A memory snapshot taken before the ADR 0112 phase 2b roll has no swap manifest, so every host refuses to restore it. The resume verb counted each refusal as a terminal failure and set the session Dead after five, although its disk was intact. A host-agent that adopts a running guest from an older host-agent writes snapshots of the same shape, and those can reference swap pages that no snapshot holds. PG cannot tell the two kinds apart. - engram-core: new SandboxError::MemoryImageUnusable. The refusal is deterministic, and the disk is unaffected. - host-agent: the three swap restore refusals return it. It crosses gRPC as a failed_precondition marker (the HarnessSpawn precedent), so WIRE_VERSION does not change. - coordinator: resume_from_fc_snapshot catches it and does a disk-only cold boot on the newer of the live and snapshot disk lineages. The guest loses its processes and keeps its disk. A new counter, engram_session_resume_memory_image_fallback_total, counts each case. - engram-dst: a sim host can refuse memory images and records each create's root disk. The new memory_image_fallback test fails without the coordinator change. - ADR 0112: a dated addendum line records the gap and the fallback. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…refusals Before this branch, a host refused a pre-roll memory snapshot with an untyped Snapshot error. The new coordinator reads that as a generic failure and counts it toward the five-failure Dead budget, and snapshot affinity keeps sending the resume to that same host. The exact-match wire gate now fences such hosts out of placement during the roll, so resumes wait for upgraded hosts instead of failing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… rolls back finish_live_restore sent every untyped error through three retries and then fail_move, which destroys the source guest. A memory image refusal is deterministic, so a retry cannot help. Handle it like postcopy-never-loaded: abort the export and roll back, so the source keeps running. The new scenario fails without this branch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cesses were lost - resume_disk_only_cold_boot passes the image's memory and CPU budgets to placement. Before, it passed none, so a cold boot could land on a host without room. - After a successful fallback, the coordinator appends a fenced resumed_from_disk event. The web renders it as a marker that says the files are intact and running processes stopped. The orchestrator frame taxonomy lists the new kind. - The fallback counter increments only after the cold boot succeeds. - The sim host counts refusals, and the DST test asserts one refusal and one resumed_from_disk event. The cold boot can no longer come from the resume skipping the snapshot. - ADR 0112: the addendum line covers the wire bump, the live teleport rollback, and why the refused row stays recoverable (demoting it would unpin the disk the cold boot can mount). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
After the ADR 0112 phase 2b roll, an Idle session with a memory snapshot from before the roll could not resume. That snapshot has no swap manifest. Every host refuses it (
swap restore has no manifest; refusing a stale device). The resume verb counted each refusal as a terminal failure, and after five (RESUME_FAILURE_STREAK_BUDGET) it set the sessionDead("fork it to continue"). The disk was intact the whole time.A host-agent that adopts a running guest from an older host-agent writes snapshots of the same shape. The guest still uses a raw swap file, and the new capture runs no
swapoff, so those memory images can reference swap pages that no snapshot holds. The two kinds look the same in PG. A migration or a host-side rule cannot separate the safe ones from the unsafe ones.Fix
Refuse the memory image, keep the disk:
engram-core: newSandboxError::MemoryImageUnusable. The refusal is deterministic (every host answers the same), and the disk is unaffected.engram-host-agent: the three swap restore refusals inpooled_backend.rsreturn the new variant.grpc_server.rssends it as afailed_preconditionmarker (theHarnessSpawnprecedent), andengram-protocolmaps it back.WIRE_VERSION34 → 35. An older host refuses the same snapshot with an untyped error, which would still count toward the Dead budget. The exact-match gate keeps such hosts out of placement during the roll, so resumes wait for upgraded hosts.engram-coordinatorresume:resume_from_fc_snapshotcatches the variant and callsresume_disk_only_cold_booton the newer oflive_disk_manifestand the snapshot'sdisk_manifest. The cold boot now carries the image's memory and CPU budgets, so placement checks that it fits. After success, the coordinator appends a fencedresumed_from_diskevent, andengram_session_resume_memory_image_fallback_totalincrements.engram-coordinatorteleport: a live move whose destination refuses the memory image aborts the export and rolls back. Before, it retried three times and thenfail_movedestroyed the source guest. A captured teleport already rolled back.web:resumed_from_diskrenders as a marker that says the files are intact and running processes stopped. The orchestrator frame taxonomy lists the new kind.The guest loses its processes and keeps its disk. An operator who upgrades from v0.10.0 needs no data migration: pre-upgrade memory snapshots resume by disk-only cold boot.
Not changed:
recoverable. Demoting it would unpin, from chunk GC, the snapshot disk that the cold boot mounts when the session has no live manifest.live_disk_manifest. A retry after that point boots the snapshot's older disk. This is older than the PR and also affects the existing disk-only branch. It is left for a separate change.Test
crates/engram-dst/tests/memory_image_fallback.rs: the sim host refuses the memory image once. The resume reachesActivewith exactly one cold boot on the live disk and oneresumed_from_diskevent. The test fails with the coordinator arm disabled.teleport_live_pgscenariorefused_memory_image_rolls_back_live_move: one restore attempt, the export is aborted, the source is not destroyed, and the session isActive. It fails without the teleport arm.engram-protocolandengram-host-agent; the wire golden pins 35.just check: 2767 passed, 317 skipped.cargo clippy --target aarch64-unknown-linux-musl -p engram-host-agent --all-targets -D warnings: clean.format:check,lint,buildMessages.test.ts(79 passed),build. orchestrator:typecheck,frame-taxonomy.test.ts(5 passed).🤖 Generated with Claude Code