Force recovery test failover with SIGSTOP - #8506
Open
Amaury Chamayou (achamayou) wants to merge 2 commits into
Open
Amaury Chamayou (achamayou) wants to merge 2 commits into
Amaury Chamayou (achamayou) wants to merge 2 commits into
Conversation
Suspend the original primary after the final recovery share commits so it cannot win another election after a stale successor is rejected. Preserve the recovery and single-opening assertions, remove obsolete SIGTERM-only setup and diagnostics, and retain bounded cleanup of suspended processes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot started reviewing on behalf of
Amaury Chamayou (achamayou)
October 5, 2026 18:07
View session
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The focused test-harness change resolves the observed race while preserving recovery invariants and cleanup.
Review effort: Balanced
Findings: None
What changed in this PR
Updates the recovery election test to eliminate re-election of the original primary.
Changes:
- Uses
SIGSTOPto force failover. - Removes obsolete SIGTERM diagnostics/configuration.
- Preserves recovery assertions and bounded cleanup.
Custom instructions used
.github/copilot-instructions.md.github/instructions/reviewing.instructions.md.github/skills/testing/SKILL.md
| File | Description |
|---|---|
tests/recovery.py |
Forces deterministic recovery failover and updates diagnostics and cleanup comments. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Max (maxtropets)
approved these changes
Oct 5, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Fix the recovery-election race shown in the failing Coverage Virtual recovery-test log. A SIGTERM stop notice nominates a successor but leaves the original primary eligible for re-election. In this failure, the nominated backup's committable log was behind, both peers correctly rejected it, and the original primary won again. The test then timed out waiting for a different primary despite successful recovery.
Implementation summary
Node.suspend()helper.ignore_first_sigtermsetup and SIGTERM-specific step-down diagnostic. Report the verified single opening and elected primary instead.The immediate variant does not wait for the backups to catch up or for the service to open before suspension. It still exercises interrupted recovery, but now a lagging backup can lose an election without allowing the original primary to win again.
Why suspension removes the race
This illustrates the observed stale-candidate ordering. The winning backup can vary; the invariant is that the suspended original primary cannot campaign or send heartbeats, while the two surviving backups retain a quorum.
sequenceDiagram participant T as Recovery test participant P as Original primary participant B as Lagging backup participant U as Up-to-date backup T->>P: Submit final recovery share P-->>T: Recovery initiated, final share committed alt Before: SIGTERM only requests a handover T->>P: SIGTERM (ignore_first_sigterm=true) P->>B: Nominate successor B->>P: Request vote (view 5, log 4.94) P-->>B: Reject: local committable log is 4.96 B->>U: Request vote (view 5, log 4.94) U-->>B: Reject: local committable log is 4.96 P->>U: Request vote (view 6, log 4.96) U-->>P: Vote granted Note over P: Original primary wins again P-->>T: Same primary, higher view Note over T: Waiting for a different primary times out else After: SIGSTOP forces primary failure T->>P: SIGSTOP Note over P: Paused until SIGKILL, cannot campaign or send heartbeats Note over B,U: Two surviving backups retain a quorum B->>U: Request vote with stale committable log U-->>B: Reject stale candidate U->>B: Request vote with up-to-date committable log B-->>U: Vote granted U-->>T: Different primary in a higher view Note over B,U: Recovery completes, exactly one service opening per ledger T->>P: SIGKILL and reap endValidation
Local build: Clang 21.1.8, Debug,
COVERAGE=ON,LONG_TESTS=ON,WORKER_THREADS=1, on a 10-core-affinity Ubuntu host.open_service_test,raft_test, andraft_enclave_testpassed 10 consecutive repetitions each.scripts/ci-checks.shpassed without auto-fix, including ASCII, formatting, lint, types, and the fresh CMake CI-bucket inventory check.The primary repeated test and full unit/check commands were:
Stress limitations, not hidden or treated as passes:
8a152f602reproduces the same late-opening assertion with a controlled 1.25-second scheduling delay before suspension.recovery_corrupt_ledgercase hitHistorical range for idx 9 not available after 3s. The same corrupt-ledger case from the unmodified baseline at8a152f602passed in isolation.These separate timing-sensitive cases are not changed by this PR. Their timeouts and assertions remain intact; this is not a full recovery-suite stress pass.
Safety and compatibility
Test-harness-only change: no production code, API, ledger format, consensus voting rule, or mixed-version behavior changes. The final recovery-share transaction is already committed before the original primary is suspended. The survivors still have to elect a different primary in a higher view, complete recovery, remain healthy, and commit exactly one service opening per ledger. The after-backups-recovered variant still requires that opening to belong to the new primary's view.
The test now waits for the ordinary election timeout rather than requesting an immediate nomination. No timeout or assertion was relaxed.