Repository navigation
fix(consensus): fail over within one operation after an unplanned leader loss - #1046
Draft
VerifiedOrganic wants to merge 34 commits into
Draft
VerifiedOrganic wants to merge 34 commits into
VerifiedOrganic wants to merge 34 commits into
Conversation
Add the documented leader-loss bounds to the fixed timing profile and a regression that kills the elected leader of a real three- and five-process projected-mTLS fleet with SIGKILL while one persistent V2 consumer streams sequential epoch-fenced writes through a surviving voter. The pinned Openraft engine leases a committed leader's vote for election_timeout_max, then waits one sampled election timeout, and checks its election timer only on a tick of heartbeat * 3 / 2. The new profile helpers expose that chain: the first-campaign bound, the replacement-election bound allowing one further campaign, and the documented write stall, which adds two cold connections and three AppendEntries rounds. Each streamed write renews one lease and updates the exact generation committed before it, so a write applied twice would break its successor. The test checks every write's exact retained receipt and requires the outage and every write stall to stay inside the documented bound, itself inside one operation timeout. Both checks fail on the current profile: - the derived write-stall bound is 39 s against a 10 s operation timeout; - three voters: outage 15,630 ms, the in-flight write stalled 15,634 ms across 21 attempts, one campaign to term 2; - five voters: outage 15,629 ms, in-flight write stall 15,632 ms. Refs: #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The pinned Openraft engine leases a committed leader's vote for election_timeout_max, waits that lease plus one sampled election timeout before a follower campaigns, and checks its election timers only on a tick of 1.5 heartbeats. With the former 2,000 ms heartbeat and [5,000 ms, 8,000 ms) elections, a surviving voter first campaigned 13 to 19 seconds after its last contact with a lost leader, so a write in flight at an unplanned loss outlived the 10-second operation timeout. The fixed profile now uses a 500 ms heartbeat/AppendEntries/read-index ceiling, a 1,000 ms Vote ceiling, [1,000 ms, 1,800 ms) elections and a 500 ms contained cold-connect cap. Every existing validator relation is kept; the cold cap follows the heartbeat because it must stay within the smallest family ceiling. InstallSnapshot, forwarded mutation, read barrier, operation and listener ceilings are unchanged. Profile validation additionally requires the documented unplanned leader-loss write stall (re-election allowing one further campaign, two cold connections and three AppendEntries rounds) to stay below the operation timeout: 9,400 ms here. The issue's example of [1,500 ms, 3,000 ms) elections would give 12,000 ms and is rejected; 2,000 ms as the maximum would leave no margin. The derived qualification envelopes follow the profile. The traffic schedule advances to v11, and its availability-recovery envelope becomes the larger of the two-election transition plus one operation and the reconciliation's two operations plus one retry, which it previously covered only implicitly. The frozen v6 profile is now checked against its own historical timing, and the rotation-fault checker pins the new values. Tests that encoded the former timing now derive it from the profile and keep their assertions: the transport cold-budget, soft-TTL reserve, long-family and follower-restart tests, the lib cooldown and budget tests, and the stale V2 warm-hint test. That test holds an all-voter scope fault for its full 10-second rejected call, during which voters now campaign; ticker elections stay disabled for exactly that window so its vote/log witness keeps proving that only the rejected call ran. The unplanned leader-loss regression now passes: SIGKILL of the leader during a fenced write stream gives a 3,606 ms outage for three voters and 3,611 ms for five, each write committed once with its exact receipt. Refs: #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Pin every thread of a three-process projected-mTLS fleet, and two spinning threads per CPU, to two CPUs for four first-campaign windows (17.4 seconds) while one consumer streams fenced writes through a follower. The fleet must keep committing and every voter must end in the starting term under the starting leader: the shorter detection window must not turn scheduling contention into an election. Two runs committed about 470 writes each; the slowest write took 159 ms and 326 ms against about 12 ms uncontended, and all voters stayed in term 1. The testkit enables rustix's thread feature for the affinity calls; no dependency is added and the lock file is unchanged. Refs: #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A crashed leader process refuses connections, but a failed node or a partition can black-hole an established connection instead. A follower that forwarded a write to such a leader then waits for its whole call deadline, even after the surviving voters elect a successor that the follower already knows. The regression makes one follower's route to the leader accept calls without answering, cuts every other path to and from the leader, and writes through that follower. The write must reroute to the successor and commit within the documented unplanned leader-loss write stall. It fails before the fix: the write ends OperationOutcomeUnavailable after 10,001 ms, the full operation deadline. Refs: #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A crashed leader process refuses connections, so a follower re-routes its forwarded write at once. A failed node or a partition can instead black-hole an established connection: the forward, read barrier, exact V2 status ticket, capability activation or expiry preflight then waited for the caller's whole deadline, about 10 seconds, even after the surviving voters had elected a successor that the follower knew. Every leader-routed call is now bounded by the replica's own leader view. Once its engine names a leader other than both the target and the leader it named when the call began, the call keeps the profile's 500 ms cold-connect allowance for an answer already in flight, such as a planned handoff's not-leader reply, and is then abandoned. Abandonment never proves that the request stayed undelivered, so it reports AfterTransmission exactly like a deadline: callers retry only the same request identity or report an unknown outcome. The write-stall bound's stale-attempt term already reserves that allowance. The configuration store keeps its existing routes; its limitation is documented. The black-holed leader regression now commits through the successor in 3.5 to 4.3 seconds across five runs instead of ending OperationOutcomeUnavailable after 10 seconds. All opc-session-store and opc-session-net tests pass: 58 test binaries, 3,093 tests. Refs: #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Pin the fork revision 5ab83b63bb199651f0b0a58a21acf89e24d77c27, still 0.9.25. In it a follower campaigns after the longer of its leader lease and its sampled election timeout instead of their sum, and the lease is the minimum election timeout. AppendEntries can have its own deadline, an optional Pre-Vote round keeps a voter that cannot win from raising its term, and a leader rejects other candidates while a quorum acknowledges it. The SDK does not enable the new options yet. The revision is the head of an unmerged fork pull request and must be re-pinned to its merge commit before this lands. The manifest, lockfiles, metadata checks, publish-order gate and documentation bind it; the frozen HA profiles keep their original revisions. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The pinned engine can run a Pre-Vote round before a voter raises its term, but only through a network that carries it. Add the PreVote RPC family (wire discriminant 9, Vote's payload bound and deadline), carry it through the session and configuration Raft adapters, and enable Pre-Vote in the shared durable Openraft configuration. - A PreVote request is admitted exactly like a Vote: from current voters, or from a staged successor only after voting admission, and only while persistence would accept the vote itself. - A failed, refused or timed-out PreVote call is never a grant, so a voter that is cut off from the others keeps its term. - Test harnesses that block or count election traffic treat PreVote as election traffic. The new consensus_openraft regression cuts one voter of a three-voter cluster off for two election timeouts while the others keep committing, then admits it again. The voter keeps its term and the leader keeps leading. With the engine's default network behavior instead of the PreVote family, the same test fails: the voter reaches term 3 and the cluster is left leaderless in that term after it returns. The leader now also rejects candidates while a quorum acknowledges it. An async-persistence race test forced a healthy leader out with a manual election; it now hands leadership to the same successor with a planned transfer, which still elects it in a higher term by an ordinary election, and its assertions are unchanged. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Advance the pin to 252ee9589db998b5d80e2a339398892d84c5d5f7, the head of the same unmerged fork pull request. It adds configuration validation that the minimum election timeout outlasts the engine tick on which a leader sends heartbeats, and fixes two fork examples. The SDK profile already satisfies the new rule. The pin must still move to the fork merge commit before this lands. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…udgets With the consumed fork rules, a follower campaigns after the longer of its lease (the minimum election timeout) and its sampled election timeout, AppendEntries has its own deadline, and Pre-Vote is on. The timing profile therefore returns to the 2,000 ms AppendEntries/read-index, 5,000 ms Vote and 1,500 ms contained cold-connect budgets that earlier commits on this branch had shortened, adds a separate 200 ms heartbeat interval, and samples elections from [5,000 ms, 6,500 ms). - The first successful campaign after a leader loss starts within election_timeout_max plus one 300 ms engine tick: 6,800 ms. - The documented write stall adds six round trips answered within one heartbeat interval, the stale-route grace and one new connection: 9,700 ms, below the 10,000 ms operation timeout. Validation still requires that, and now also that a follower lease spans two engine ticks and two AppendEntries ceilings. - The session store abandons a call to a black-holed lost leader one heartbeat interval after it observes the successor, instead of after the cold-connect allowance. - The session-net transport module and transport tests, the qualification envelope formulas and the V2 stale-hint test return to their main-branch form. The unchanged formulas follow the shorter maximum election timeout: the envelopes move from 26 to 23 seconds, and the traffic and candidate schedule digests and the rotation evidence checker with them. - The frozen v6 timing comparison records that only the heartbeat interval and the maximum election timeout moved. On this host the three- and five-process SIGKILL leader-loss streams recover in 5.6 s and 5.4 s with every write committed once, a healthy fleet on two contended CPUs keeps its leader for 20.4 s, the black-holed reroute commits in 6.5 s, and the consensus_openraft, fenced V2 qualification, consensus_transport, protected-roster and session-store library suites pass. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Restate the timing contract for the current profile and engine rules: 2,000 ms AppendEntries/read-index, a separate 200 ms heartbeat interval, 5,000 ms Vote and PreVote, [5,000 ms, 6,500 ms) elections and the 1,500 ms contained cold-connect cap. RFC 004 and ADR 0003 keep their transport values; the earlier shortened values on this branch are withdrawn. - The opc-consensus README derives the 6,800 ms first-campaign bound and the 9,700 ms write stall from the overlapping lease, Pre-Vote and the 300 ms tick, states their assumptions and the CPU requirement, and describes the PreVote family. - The operator runbook's section 2.4, the HA design and operator readiness give the same bounds and the effect of Pre-Vote on a voter that is cut off or restarted. - The session-store README describes the one-heartbeat stale-route grace and the PreVote admission rule. - Profile-derived qualification envelopes read 23 seconds instead of 26, with the dependent restart, recovery and settlement stages. - ADR 0019 records why detection took 13 to 19 seconds before. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…e envelope Two paused-clock tests advanced a literal 25 s, one second before the 26 s traffic availability recovery envelope. The shorter maximum election timeout moved that envelope to 23 s, so the advance passed the deadline before the probe started, the probe was never entered and both tests waited forever. Advance to one second before the envelope constant instead. Every assertion is unchanged: the pending probe still ends exactly at the original episode deadline. Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…nvelope Three harness arithmetic tests pinned the member recovery progress checkpoint at 13 s and its coverage and availability recovery deadlines at 26 s. Both follow the maximum election timeout through unchanged formulas, so the shorter election window moved them to 11.5 s and 23 s, and the margin a fresh pulse leaves above one complete operation from 3 s to 1.5 s. Pin the new values. Every relationship the tests check is unchanged: a refreshed pulse still cannot hide the residual coverage deadline, a later phase deadline still never replaces the rolling bounds, and a completion just after the initial checkpoint is still admitted once the availability allowance applies. Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…indow The write-budget contract delays and counts every empty AppendEntries that reaches a follower before the measured write's own entries, to prove that the write issues no separate leadership-confirmation round. A leader's periodic idle heartbeat is an empty AppendEntries too. When one fell inside the window it was delayed on both followers' ordered replication streams, so the write could not commit within its budget and the test failed although the write path was unchanged. With the 300 ms engine tick this happens often enough to fail CI. Waiting just over one tick inside the window fails the test every time. Add a test-control hook that pauses a leader's periodic heartbeats, and pause them before the binding write that precedes the measurement. Every voter applying that write proves that any heartbeat sent before the pause has crossed its follower's stream. Confirmation rounds do not use periodic heartbeats: a leadership-confirmation round injected into the consumer write path still fails the test with heartbeats paused. Every assertion is unchanged. Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
This was referenced Oct 2, 2026
A rolling upgrade, or its rollback, replaces one voter at a time, so for a while some voters run the previous release beside this one. The multi-process qualification fleet can now start each voter from its own node binary and roll a voter in place on its database. A new ignored module takes the previous release's binary from OPC_SESSION_QUORUM_NODE_PREVIOUS_RELEASE and runs, with 3 and 5 voters: - a leader loss while the previous-release voters lag behind this release's majority; - a rolling upgrade and a rolling rollback under a stream of fenced writes, killing the leader once while it still runs the outgoing release and once when the voter set is mixed. Each requires a successor within the mixed-release bound, every write applied exactly once, and every voter following one leader afterwards. The lagging-voter test fails on this revision against the SDK main node binary. A voter of this release sends Pre-Vote to previous-release voters, which cannot decode it, so it never campaigns. The previous-release voters cannot win with their stale logs, and no leader is elected: after 60 seconds the current voter is still in term 1, and the previous-release voter has reached term 3. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…rule Advance the pin to 38c6524a405a55e2d6808b136329ff62ba4e492d, the head of the same unmerged fork pull request. It adds two engine rules for voters that cannot answer Pre-Vote, such as voters of a release without it: - A voter that rejects a candidate for its stale log, while it has no leader, campaigns without Pre-Vote one election timeout later, in a term above that candidate's, unless a leader appears. - A Pre-Vote response never makes the requester adopt a vote for itself. No SDK election, vote or quorum logic changes. The pin must still move to the fork merge commit before this lands. With this engine, the mixed-release lagging-voter qualification elects the up-to-date survivor of this release after the stale previous-release candidate's campaign. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A voter of this release sent Pre-Vote to every voter. A voter of the previous release cannot decode the PreVote family: it drops the connection, and with it any other call in flight on that lane. A regression against a server that offers only the previous consensus protocol shows the client sending the Pre-Vote anyway. Each consensus connection now negotiates its protocol through TLS ALPN. Clients and servers of this release offer opc-session-consensus/3, which carries Pre-Vote, ahead of opc-session-consensus/2. A peer of the previous release negotiates /2. On such a connection the client sends no Pre-Vote, and the server refuses one that arrives. ConsensusPeer gains call_pre_vote. A transport that knows its voter runs this release forwards the request. Otherwise it reports the voter as unable to answer, without sending anything. The default reports that, so a transport that cannot tell never sends Pre-Vote. RemoteSessionConsensusPeer decides from the negotiated protocol; in-process peers forward with forward_pre_vote, and wrappers delegate. The session and configuration Raft adapters count a voter that cannot answer as rejecting the Pre-Vote, with no vote to adopt. Previous- release voters then lead an election they are needed for, with the classic vote. The engine's stale-candidate rule makes a voter of this release campaign when their logs are behind. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The runbook gains a rolling-upgrade section. It covers the per-connection ALPN negotiation, how a leader loss is resolved while a previous-release voter is still needed for a majority, and the 30-second bound. It also covers a write in flight at such a loss, which ends ambiguous at least once and still applies exactly once. The opc-consensus README describes ConsensusPeer::call_pre_vote and the same rules, and moves its unplanned-leader-loss section after the general timing material it had been placed before. The CHANGELOG replaces the stop-and-upgrade note with rolling-upgrade compatibility. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Add a required lane that checks out the previous release beside the tree under test: the pull request's base, or the previous main commit on a push. It builds that revision's quorum node exactly as the test shards build theirs, then runs the ignored mixed-release qualification against it with 3 and 5 voters: - a leader loss with lagging previous-release voters; - a rolling upgrade and a rolling rollback under fenced writes, with leader kills. A change that breaks compatibility with the previous release then fails in CI instead of during a rolling upgrade. The workflow header now counts the jobs it actually requests. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The server-side Pre-Vote refusal test serves two connections, and each ends in an idle retirement. Running beside the connection-outcome tests, which count that process-wide metric exactly, it moved their count by two and failed one of them. Take the same metrics lock those tests hold. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The qualification fleet records the tree under test, untracked files included, as candidate evidence. The previous release's checkout sat inside the workspace as an untracked nested repository, which that capture rejects as not a bounded regular file. Every mixed-release test therefore failed before starting a fleet. Move the checkout to the runner's temporary directory, and build its quorum node into this tree's ignored, cached target directory. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…before forwarding Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…d path The lost leader's black-holed path also carries the follower's Raft calls, so counting every call made the transmission check depend on election timing. Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A forwarded call to a lost leader that black-holes its connection is abandoned once this replica's engine names a successor. The wrapper compared later views with the one it saw when the call began. The caller revalidates source authority between choosing the route and starting the call, and a successor learned during that wait was taken as the starting view, so the call to the lost leader consumed the caller's whole deadline and ended with an unknown outcome. Each route now carries the leader view it was chosen under: the target itself for a route taken from the local view, and the local view at the time of a peer's redirect for a redirected route, which may lead away from a leader the local view still names. A route that the local view already supersedes is not sent at all and reports BeforeTransmission, so the caller refreshes its route without an ambiguous outcome; one superseded while in flight keeps the stale-route grace as before. Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…re-Vote Consume the Openraft fork at 71f7ff7ed2a0a1783362264d597926a243d62b37. Its Pre-Vote RPC now reports, per voter, that the voter is not known to answer Pre-Vote, and the engine then runs the classic election for that campaign. The revision also tells Pre-Vote rounds apart by an identifier, retries a round rejected by a running lease after the election-timeout window, gives up an in-flight AppendEntries when its replication stream closes, and drops the stale-candidate rule. Both Raft adapters counted a voter that the transport could not send a Pre-Vote to as rejecting it. A transport that does not implement call_pre_vote reports every voter that way, so after an unplanned leader loss no survivor could ever reach a Pre-Vote quorum, and only initialization, which campaigns without Pre-Vote, ever elected a leader. The adapters now report such a voter as unsupported. The default call_pre_vote still sends nothing, so a transport written before Pre-Vote keeps automatic failover with the classic election. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…Pre-Vote retry The bound claimed that the survivor that heard from the leader last campaigns last, after every other lease has expired. A survivor that heard from the leader slightly earlier can campaign first, while a lagging survivor's lease still runs, and the engine held that rejected Pre-Vote round for another sampled election timeout, about 11.5 s after the loss with these values. The consumed engine retries a rejected round after the width of the election-timeout window. Every lease runs out within the minimum election timeout of the loss, so the survivor with the most up-to-date log starts a round that no lease rejects within the minimum election timeout, the window and one tick: the unchanged 6,800 ms. The profile test now states that arithmetic. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…lection A voter of this build now runs the classic election whenever it reaches a previous-release voter, so a mixed voter set elects as the previous release does. The bound is the longer of two paths: previous-release voters only (their first campaign within 19 s of their last leader contact, and one split vote of 11 s more), or a voter of this build that can win (its first campaign after at most one cold connection, a retry after a previous-release lease rejected it, 15.1 s). Its value stays 30 s. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…views Pre-Vote is used only while every voter that can be reached is positively known to answer it; otherwise that campaign runs the classic election, so voters of a release without Pre-Vote and transports that do not implement it never block an election. Derive the 6,800 ms first successful campaign bound from the Pre-Vote retry after the election-timeout window, recompute the mixed-release bounds for classic elections during a roll, and describe how a leader route is judged against the view it was chosen under. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The fork head only drops the doc comment of the Pre-Vote round matcher that round identifiers replaced; the code is unchanged from 71f7ff7e. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
… again Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…observed A route that a peer's redirect sent to B, while this replica's view still named A, was never superseded by A. If the view advanced to B while the caller revalidated its authority, and B was then lost while the surviving voters elected A again, a call that B black-holed waited for its deadline although the replica knew the successor. The route now records each view the watcher observes, at its start included: once the view names the target, the leader the redirect was chosen under no longer exempts the route, and A leading again supersedes it like any other leader. Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…rs lag Lose a leader of this build while this build's other voters miss the last entries, so that only a previous-release voter can win, by its own slower timers. This build's survivors run the classic election because a previous-release voter cannot answer Pre-Vote; each such campaign votes for itself in its next term, where a previous-release campaign can meet it. The rolling-upgrade lane runs the 3- and 5-voter cases against the previous release's node binary. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The fork keeps a Pre-Vote round open until its replies are due, so voters that answer more slowly than the election-timeout window still reach a quorum and a late unsupported report still starts the classic election. A campaigning node that refuses a candidate with a more up-to-date log only because of its own vote now defers its next campaign by the greater-log timeout, so that candidate's next campaign finds no new self-vote. Neither changes the engine's public API. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The first successful campaign after an unplanned leader loss starts within 6,800 ms when every surviving voter answers a Pre-Vote within the 1,500 ms election-timeout window and every engine tick runs on time: each round is then rejected or granted within that width. A Pre-Vote round now stays open until its Vote deadline, so a slower voter delays the election by the excess and never prevents it, and a lost leader that never answers does not delay a round the survivors answer. While a previous-release voter is a member, a successor is elected within 30 s when the leader is the only voter lost, round trips take at most one heartbeat interval and no two voters campaign at once. Each voter whose log is behind refuses the winner at most once: a voter of this build defers its next campaign by its greater-log timeout after such a refusal, and a previous-release voter that has seen a greater log campaigns at most once per 21 s. The winner therefore needs at most one more campaign. Without the deferral, a three-process fleet whose survivor of this build lagged elected no successor within 30 s. Refs #1037 Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Status
Draft. This consumes the unmerged dependency openpacketcore/openraft#10, pinned at
20f4f3123168907d5a00ec3772c78872e9a47820. The fork is still a draft; its current head452df5e55b18840bea1311acd26cb5faafea88abdiffers from this pin only by comment formatting. Before merge, the dependency needs completed qualification and a reviewed integration revision, followed by an SDK re-pin and fresh CI.The fork's delayed-network qualification currently fails the election-retry test's expected-winner assertion. Its fixture checks connectivity before a delayed send, then reads the AppendEntries quota afterward. Restoring the quota immediately after isolating the leader can therefore admit an already-delayed append and invalidate the assumption that only node 1 has the newest log. This needs resolution in the fork before treating its timing qualification as complete.
Refs #1037. Faster detection, check-quorum step-down, and configuration-store route rediscovery remain tracked separately in #1047, #1048, and #1049.
Problem and behavior
Previously, the leader lease and a full sampled election timeout ran consecutively, and timers were checked every 3 seconds. An unplanned leader loss could therefore delay the first campaign for 13–19 seconds, beyond the default 10-second operation budget.
This change shares one timing profile across both durable adapters:
[5,000, 6,500)msThe consumed engine overlaps the leader lease and election timeout, separates the heartbeat interval from the AppendEntries deadline, and adds Pre-Vote with bounded retries. A leader rejects candidates while a quorum still acknowledges it. A campaigning voter that rejects a newer-log candidate only because of its own vote defers its next campaign by the greater-log timeout.
The SDK carries Pre-Vote through both Raft adapters and adds session-store route retirement: once the forwarding replica observes a successor, an unanswered call to the lost leader is abandoned after a 200 ms grace. A route already superseded before transmission is not sent. Possibly transmitted calls retain their ambiguous classification so callers retry only the same request identity. A redirect's exemption is retired once its target has been observed as leader, including leadership returning to the original source.
Compatibility and bounds
Consensus connections negotiate
opc-session-consensus/3ahead of/2through TLS ALPN. Pre-Vote is sent only to peers known to support it. A reachable peer that cannot answer Pre-Vote makes that campaign use the classic election; custom transports inherit this fallback until they implementcall_pre_vote. An unreachable peer grants neither request and does not itself force classic voting.For a fleet entirely on this release, the documented first successful campaign starts within 6,800 ms and the write-stall bound is 9,700 ms. These are conditional bounds: the majority must be reachable, engine ticks must run on time, processes must not be suspended or throttled, and the write bound assumes each round trip including disk sync completes within one heartbeat interval. The campaign bound requires surviving voters to answer Pre-Vote within the 1,500 ms retry window. Split votes are outside the bound. The derivation and assumptions are in the
opc-consensusREADME and consensus operator runbook.Mixed-release elections retain the documented classic-election envelope of 30 seconds under its stated assumptions: only the leader is lost, processes keep running, and campaigns do not overlap within one round trip. A rolling upgrade or rollback can therefore outlast one 10-second operation; ambiguous writes must be retried with the same identity. Pre-Vote protection resumes once every reachable voter supports it.
Coverage and validation
The PR adds or extends coverage for isolated-voter Pre-Vote behavior, transports without Pre-Vote,
/2negotiation, stale and redirected routes, three- and five-process leader loss, CPU contention, and eight mixed-release leader-loss/rolling-upgrade/rollback scenarios. Existing qualification envelopes retain their formulas and follow the shorter maximum election timeout; frozen historical profiles retain their original values.Current published SDK head:
bdce5c68f3e9d4ca36ce52c93f322f9a1b3b673f. Its CI run is not green. The i686 ordinary lane has two failing tests; the i686 and Rust workspace aggregate checks propagate that failure. The contracts lane passes.[10, 9, 10]to[10, 10, 10]; convergence must precede the baseline while preserving the exact sequence and read-only assertions.The earlier green result on
6f4076587is historical evidence, not qualification of the current head. Main integration, the i686 correction, the reviewed dependency pin, and full remote CI must be completed on the final candidate before this draft is ready for independent review.