Skip to content

fix(consensus): fail over within one operation after an unplanned leader loss - #1046

Draft
VerifiedOrganic wants to merge 34 commits into
mainfrom
fix/1037-consensus-failover-timing
Draft

VerifiedOrganic wants to merge 34 commits into
mainfrom
fix/1037-consensus-failover-timing

Conversation

@VerifiedOrganic

@VerifiedOrganic VerifiedOrganic commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Status

Draft. This consumes the unmerged dependency openpacketcore/openraft#10, pinned at 20f4f3123168907d5a00ec3772c78872e9a47820. The fork is still a draft; its current head 452df5e55b18840bea1311acd26cb5faafea88ab differs from this pin only by comment formatting. Before merge, the dependency needs completed qualification and a reviewed integration revision, followed by an SDK re-pin and fresh CI.

The fork's delayed-network qualification currently fails the election-retry test's expected-winner assertion. Its fixture checks connectivity before a delayed send, then reads the AppendEntries quota afterward. Restoring the quota immediately after isolating the leader can therefore admit an already-delayed append and invalidate the assumption that only node 1 has the newest log. This needs resolution in the fork before treating its timing qualification as complete.

Refs #1037. Faster detection, check-quorum step-down, and configuration-store route rediscovery remain tracked separately in #1047, #1048, and #1049.

Problem and behavior

Previously, the leader lease and a full sampled election timeout ran consecutively, and timers were checked every 3 seconds. An unplanned leader loss could therefore delay the first campaign for 13–19 seconds, beyond the default 10-second operation budget.

This change shares one timing profile across both durable adapters:

Setting Value
Heartbeat interval / engine tick 200 ms / 300 ms
Randomized election timeout [5,000, 6,500) ms
AppendEntries and read-index ceiling 2,000 ms
Vote and Pre-Vote ceiling 5,000 ms
Contained cold-connection allowance 1,500 ms
Complete operation default 10,000 ms

The consumed engine overlaps the leader lease and election timeout, separates the heartbeat interval from the AppendEntries deadline, and adds Pre-Vote with bounded retries. A leader rejects candidates while a quorum still acknowledges it. A campaigning voter that rejects a newer-log candidate only because of its own vote defers its next campaign by the greater-log timeout.

The SDK carries Pre-Vote through both Raft adapters and adds session-store route retirement: once the forwarding replica observes a successor, an unanswered call to the lost leader is abandoned after a 200 ms grace. A route already superseded before transmission is not sent. Possibly transmitted calls retain their ambiguous classification so callers retry only the same request identity. A redirect's exemption is retired once its target has been observed as leader, including leadership returning to the original source.

Compatibility and bounds

Consensus connections negotiate opc-session-consensus/3 ahead of /2 through TLS ALPN. Pre-Vote is sent only to peers known to support it. A reachable peer that cannot answer Pre-Vote makes that campaign use the classic election; custom transports inherit this fallback until they implement call_pre_vote. An unreachable peer grants neither request and does not itself force classic voting.

For a fleet entirely on this release, the documented first successful campaign starts within 6,800 ms and the write-stall bound is 9,700 ms. These are conditional bounds: the majority must be reachable, engine ticks must run on time, processes must not be suspended or throttled, and the write bound assumes each round trip including disk sync completes within one heartbeat interval. The campaign bound requires surviving voters to answer Pre-Vote within the 1,500 ms retry window. Split votes are outside the bound. The derivation and assumptions are in the opc-consensus README and consensus operator runbook.

Mixed-release elections retain the documented classic-election envelope of 30 seconds under its stated assumptions: only the leader is lost, processes keep running, and campaigns do not overlap within one round trip. A rolling upgrade or rollback can therefore outlast one 10-second operation; ambiguous writes must be retried with the same identity. Pre-Vote protection resumes once every reachable voter supports it.

Coverage and validation

The PR adds or extends coverage for isolated-voter Pre-Vote behavior, transports without Pre-Vote, /2 negotiation, stale and redirected routes, three- and five-process leader loss, CPU contention, and eight mixed-release leader-loss/rolling-upgrade/rollback scenarios. Existing qualification envelopes retain their formulas and follow the shorter maximum election timeout; frozen historical profiles retain their original values.

Current published SDK head: bdce5c68f3e9d4ca36ce52c93f322f9a1b3b673f. Its CI run is not green. The i686 ordinary lane has two failing tests; the i686 and Rust workspace aggregate checks propagate that failure. The contracts lane passes.

  • The aborted-terminal status test records its read-only baseline before a preceding quorum write has reached every follower. The failing observation changes from [10, 9, 10] to [10, 10, 10]; convergence must precede the baseline while preserving the exact sequence and read-only assertions.
  • The tenant-isolation terminal test currently combines several distinct failure outcomes in one panic. The recorded failure does not establish which outcome occurred or whether it is a product regression; focused diagnostics and another validation run are required.

The earlier green result on 6f4076587 is historical evidence, not qualification of the current head. Main integration, the i686 correction, the reviewed dependency pin, and full remote CI must be completed on the final candidate before this draft is ready for independent review.

Add the documented leader-loss bounds to the fixed timing profile and a
regression that kills the elected leader of a real three- and
five-process projected-mTLS fleet with SIGKILL while one persistent V2
consumer streams sequential epoch-fenced writes through a surviving
voter.

The pinned Openraft engine leases a committed leader's vote for
election_timeout_max, then waits one sampled election timeout, and
checks its election timer only on a tick of heartbeat * 3 / 2. The new
profile helpers expose that chain: the first-campaign bound, the
replacement-election bound allowing one further campaign, and the
documented write stall, which adds two cold connections and three
AppendEntries rounds.

Each streamed write renews one lease and updates the exact generation
committed before it, so a write applied twice would break its
successor. The test checks every write's exact retained receipt and
requires the outage and every write stall to stay inside the documented
bound, itself inside one operation timeout.

Both checks fail on the current profile:

- the derived write-stall bound is 39 s against a 10 s operation
  timeout;
- three voters: outage 15,630 ms, the in-flight write stalled 15,634 ms
  across 21 attempts, one campaign to term 2;
- five voters: outage 15,629 ms, in-flight write stall 15,632 ms.

Refs: #1037
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The pinned Openraft engine leases a committed leader's vote for
election_timeout_max, waits that lease plus one sampled election timeout
before a follower campaigns, and checks its election timers only on a
tick of 1.5 heartbeats. With the former 2,000 ms heartbeat and
[5,000 ms, 8,000 ms) elections, a surviving voter first campaigned 13 to
19 seconds after its last contact with a lost leader, so a write in
flight at an unplanned loss outlived the 10-second operation timeout.

The fixed profile now uses a 500 ms heartbeat/AppendEntries/read-index
ceiling, a 1,000 ms Vote ceiling, [1,000 ms, 1,800 ms) elections and a
500 ms contained cold-connect cap. Every existing validator relation is
kept; the cold cap follows the heartbeat because it must stay within the
smallest family ceiling. InstallSnapshot, forwarded mutation, read
barrier, operation and listener ceilings are unchanged. Profile
validation additionally requires the documented unplanned leader-loss
write stall (re-election allowing one further campaign, two cold
connections and three AppendEntries rounds) to stay below the operation
timeout: 9,400 ms here. The issue's example of [1,500 ms, 3,000 ms)
elections would give 12,000 ms and is rejected; 2,000 ms as the maximum
would leave no margin.

The derived qualification envelopes follow the profile. The traffic
schedule advances to v11, and its availability-recovery envelope becomes
the larger of the two-election transition plus one operation and the
reconciliation's two operations plus one retry, which it previously
covered only implicitly. The frozen v6 profile is now checked against
its own historical timing, and the rotation-fault checker pins the new
values.

Tests that encoded the former timing now derive it from the profile and
keep their assertions: the transport cold-budget, soft-TTL reserve,
long-family and follower-restart tests, the lib cooldown and budget
tests, and the stale V2 warm-hint test. That test holds an all-voter
scope fault for its full 10-second rejected call, during which voters
now campaign; ticker elections stay disabled for exactly that window so
its vote/log witness keeps proving that only the rejected call ran.

The unplanned leader-loss regression now passes: SIGKILL of the leader
during a fenced write stream gives a 3,606 ms outage for three voters
and 3,611 ms for five, each write committed once with its exact
receipt.

Refs: #1037
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Pin every thread of a three-process projected-mTLS fleet, and two
spinning threads per CPU, to two CPUs for four first-campaign windows
(17.4 seconds) while one consumer streams fenced writes through a
follower. The fleet must keep committing and every voter must end in
the starting term under the starting leader: the shorter detection
window must not turn scheduling contention into an election.

Two runs committed about 470 writes each; the slowest write took 159 ms
and 326 ms against about 12 ms uncontended, and all voters stayed in
term 1. The testkit enables rustix's thread feature for the affinity
calls; no dependency is added and the lock file is unchanged.

Refs: #1037
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A crashed leader process refuses connections, but a failed node or a
partition can black-hole an established connection instead. A follower
that forwarded a write to such a leader then waits for its whole call
deadline, even after the surviving voters elect a successor that the
follower already knows.

The regression makes one follower's route to the leader accept calls
without answering, cuts every other path to and from the leader, and
writes through that follower. The write must reroute to the successor
and commit within the documented unplanned leader-loss write stall.

It fails before the fix: the write ends OperationOutcomeUnavailable
after 10,001 ms, the full operation deadline.

Refs: #1037
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A crashed leader process refuses connections, so a follower re-routes
its forwarded write at once. A failed node or a partition can instead
black-hole an established connection: the forward, read barrier, exact
V2 status ticket, capability activation or expiry preflight then waited
for the caller's whole deadline, about 10 seconds, even after the
surviving voters had elected a successor that the follower knew.

Every leader-routed call is now bounded by the replica's own leader
view. Once its engine names a leader other than both the target and the
leader it named when the call began, the call keeps the profile's 500 ms
cold-connect allowance for an answer already in flight, such as a
planned handoff's not-leader reply, and is then abandoned. Abandonment
never proves that the request stayed undelivered, so it reports
AfterTransmission exactly like a deadline: callers retry only the same
request identity or report an unknown outcome. The write-stall bound's
stale-attempt term already reserves that allowance. The configuration
store keeps its existing routes; its limitation is documented.

The black-holed leader regression now commits through the successor in
3.5 to 4.3 seconds across five runs instead of ending
OperationOutcomeUnavailable after 10 seconds. All opc-session-store and
opc-session-net tests pass: 58 test binaries, 3,093 tests.

Refs: #1037
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Pin the fork revision 5ab83b63bb199651f0b0a58a21acf89e24d77c27, still
0.9.25. In it a follower campaigns after the longer of its leader lease
and its sampled election timeout instead of their sum, and the lease is
the minimum election timeout. AppendEntries can have its own deadline, an
optional Pre-Vote round keeps a voter that cannot win from raising its
term, and a leader rejects other candidates while a quorum acknowledges
it. The SDK does not enable the new options yet.

The revision is the head of an unmerged fork pull request and must be
re-pinned to its merge commit before this lands. The manifest, lockfiles,
metadata checks, publish-order gate and documentation bind it; the
frozen HA profiles keep their original revisions.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The pinned engine can run a Pre-Vote round before a voter raises its
term, but only through a network that carries it. Add the PreVote RPC
family (wire discriminant 9, Vote's payload bound and deadline), carry
it through the session and configuration Raft adapters, and enable
Pre-Vote in the shared durable Openraft configuration.

- A PreVote request is admitted exactly like a Vote: from current
  voters, or from a staged successor only after voting admission, and
  only while persistence would accept the vote itself.
- A failed, refused or timed-out PreVote call is never a grant, so a
  voter that is cut off from the others keeps its term.
- Test harnesses that block or count election traffic treat PreVote as
  election traffic.

The new consensus_openraft regression cuts one voter of a three-voter
cluster off for two election timeouts while the others keep committing,
then admits it again. The voter keeps its term and the leader keeps
leading. With the engine's default network behavior instead of the
PreVote family, the same test fails: the voter reaches term 3 and the
cluster is left leaderless in that term after it returns.

The leader now also rejects candidates while a quorum acknowledges it.
An async-persistence race test forced a healthy leader out with a
manual election; it now hands leadership to the same successor with a
planned transfer, which still elects it in a higher term by an ordinary
election, and its assertions are unchanged.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Advance the pin to 252ee9589db998b5d80e2a339398892d84c5d5f7, the head of the same unmerged fork pull request. It adds configuration validation that the minimum election timeout outlasts the engine tick on which a leader sends heartbeats, and fixes two fork examples. The SDK profile already satisfies the new rule. The pin must still move to the fork merge commit before this lands.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…udgets

With the consumed fork rules, a follower campaigns after the longer of its
lease (the minimum election timeout) and its sampled election timeout,
AppendEntries has its own deadline, and Pre-Vote is on. The timing profile
therefore returns to the 2,000 ms AppendEntries/read-index, 5,000 ms Vote
and 1,500 ms contained cold-connect budgets that earlier commits on this
branch had shortened, adds a separate 200 ms heartbeat interval, and
samples elections from [5,000 ms, 6,500 ms).

- The first successful campaign after a leader loss starts within
  election_timeout_max plus one 300 ms engine tick: 6,800 ms.
- The documented write stall adds six round trips answered within one
  heartbeat interval, the stale-route grace and one new connection:
  9,700 ms, below the 10,000 ms operation timeout. Validation still
  requires that, and now also that a follower lease spans two engine
  ticks and two AppendEntries ceilings.
- The session store abandons a call to a black-holed lost leader one
  heartbeat interval after it observes the successor, instead of after
  the cold-connect allowance.
- The session-net transport module and transport tests, the qualification
  envelope formulas and the V2 stale-hint test return to their main-branch
  form. The unchanged formulas follow the shorter maximum election
  timeout: the envelopes move from 26 to 23 seconds, and the traffic and
  candidate schedule digests and the rotation evidence checker with them.
- The frozen v6 timing comparison records that only the heartbeat interval
  and the maximum election timeout moved.

On this host the three- and five-process SIGKILL leader-loss streams
recover in 5.6 s and 5.4 s with every write committed once, a healthy
fleet on two contended CPUs keeps its leader for 20.4 s, the black-holed
reroute commits in 6.5 s, and the consensus_openraft, fenced V2
qualification, consensus_transport, protected-roster and session-store
library suites pass.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Restate the timing contract for the current profile and engine rules:
2,000 ms AppendEntries/read-index, a separate 200 ms heartbeat interval,
5,000 ms Vote and PreVote, [5,000 ms, 6,500 ms) elections and the 1,500 ms
contained cold-connect cap. RFC 004 and ADR 0003 keep their transport
values; the earlier shortened values on this branch are withdrawn.

- The opc-consensus README derives the 6,800 ms first-campaign bound and
  the 9,700 ms write stall from the overlapping lease, Pre-Vote and the
  300 ms tick, states their assumptions and the CPU requirement, and
  describes the PreVote family.
- The operator runbook's section 2.4, the HA design and operator
  readiness give the same bounds and the effect of Pre-Vote on a voter
  that is cut off or restarted.
- The session-store README describes the one-heartbeat stale-route
  grace and the PreVote admission rule.
- Profile-derived qualification envelopes read 23 seconds instead of 26,
  with the dependent restart, recovery and settlement stages.
- ADR 0019 records why detection took 13 to 19 seconds before.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…e envelope

Two paused-clock tests advanced a literal 25 s, one second before the
26 s traffic availability recovery envelope. The shorter maximum
election timeout moved that envelope to 23 s, so the advance passed the
deadline before the probe started, the probe was never entered and both
tests waited forever.

Advance to one second before the envelope constant instead. Every
assertion is unchanged: the pending probe still ends exactly at the
original episode deadline.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…nvelope

Three harness arithmetic tests pinned the member recovery progress
checkpoint at 13 s and its coverage and availability recovery deadlines
at 26 s. Both follow the maximum election timeout through unchanged
formulas, so the shorter election window moved them to 11.5 s and 23 s,
and the margin a fresh pulse leaves above one complete operation from
3 s to 1.5 s.

Pin the new values. Every relationship the tests check is unchanged: a
refreshed pulse still cannot hide the residual coverage deadline, a
later phase deadline still never replaces the rolling bounds, and a
completion just after the initial checkpoint is still admitted once the
availability allowance applies.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…indow

The write-budget contract delays and counts every empty AppendEntries
that reaches a follower before the measured write's own entries, to
prove that the write issues no separate leadership-confirmation round.
A leader's periodic idle heartbeat is an empty AppendEntries too. When
one fell inside the window it was delayed on both followers' ordered
replication streams, so the write could not commit within its budget
and the test failed although the write path was unchanged. With the
300 ms engine tick this happens often enough to fail CI. Waiting just
over one tick inside the window fails the test every time.

Add a test-control hook that pauses a leader's periodic heartbeats, and
pause them before the binding write that precedes the measurement.
Every voter applying that write proves that any heartbeat sent before
the pause has crossed its follower's stream. Confirmation rounds do not
use periodic heartbeats: a leadership-confirmation round injected into
the consumer write path still fails the test with heartbeats paused.
Every assertion is unchanged.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A rolling upgrade, or its rollback, replaces one voter at a time, so for
a while some voters run the previous release beside this one. The
multi-process qualification fleet can now start each voter from its own
node binary and roll a voter in place on its database. A new ignored
module takes the previous release's binary from
OPC_SESSION_QUORUM_NODE_PREVIOUS_RELEASE and runs, with 3 and 5 voters:

- a leader loss while the previous-release voters lag behind this
  release's majority;
- a rolling upgrade and a rolling rollback under a stream of fenced
  writes, killing the leader once while it still runs the outgoing
  release and once when the voter set is mixed.

Each requires a successor within the mixed-release bound, every write
applied exactly once, and every voter following one leader afterwards.

The lagging-voter test fails on this revision against the SDK main node
binary. A voter of this release sends Pre-Vote to previous-release
voters, which cannot decode it, so it never campaigns. The
previous-release voters cannot win with their stale logs, and no leader
is elected: after 60 seconds the current voter is still in term 1, and
the previous-release voter has reached term 3.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…rule

Advance the pin to 38c6524a405a55e2d6808b136329ff62ba4e492d, the head
of the same unmerged fork pull request. It adds two engine rules for
voters that cannot answer Pre-Vote, such as voters of a release without
it:

- A voter that rejects a candidate for its stale log, while it has no
  leader, campaigns without Pre-Vote one election timeout later, in a
  term above that candidate's, unless a leader appears.
- A Pre-Vote response never makes the requester adopt a vote for
  itself.

No SDK election, vote or quorum logic changes. The pin must still move
to the fork merge commit before this lands.

With this engine, the mixed-release lagging-voter qualification elects
the up-to-date survivor of this release after the stale previous-release
candidate's campaign.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A voter of this release sent Pre-Vote to every voter. A voter of the
previous release cannot decode the PreVote family: it drops the
connection, and with it any other call in flight on that lane. A
regression against a server that offers only the previous consensus
protocol shows the client sending the Pre-Vote anyway.

Each consensus connection now negotiates its protocol through TLS
ALPN. Clients and servers of this release offer
opc-session-consensus/3, which carries Pre-Vote, ahead of
opc-session-consensus/2. A peer of the previous release negotiates
/2. On such a connection the client sends no Pre-Vote, and the server
refuses one that arrives.

ConsensusPeer gains call_pre_vote. A transport that knows its voter
runs this release forwards the request. Otherwise it reports the voter
as unable to answer, without sending anything. The default reports
that, so a transport that cannot tell never sends Pre-Vote.
RemoteSessionConsensusPeer decides from the negotiated protocol;
in-process peers forward with forward_pre_vote, and wrappers delegate.

The session and configuration Raft adapters count a voter that cannot
answer as rejecting the Pre-Vote, with no vote to adopt. Previous-
release voters then lead an election they are needed for, with the
classic vote. The engine's stale-candidate rule makes a voter of this
release campaign when their logs are behind.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The runbook gains a rolling-upgrade section. It covers the per-connection
ALPN negotiation, how a leader loss is resolved while a previous-release
voter is still needed for a majority, and the 30-second bound. It also
covers a write in flight at such a loss, which ends ambiguous at least
once and still applies exactly once. The opc-consensus README describes
ConsensusPeer::call_pre_vote and the same rules, and moves its
unplanned-leader-loss section after the general timing material it had
been placed before. The CHANGELOG replaces the stop-and-upgrade note
with rolling-upgrade compatibility.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Add a required lane that checks out the previous release beside the
tree under test: the pull request's base, or the previous main commit on
a push. It builds that revision's quorum node exactly as the test shards
build theirs, then runs the ignored mixed-release qualification against
it with 3 and 5 voters:

- a leader loss with lagging previous-release voters;
- a rolling upgrade and a rolling rollback under fenced writes, with
  leader kills.

A change that breaks compatibility with the previous release then fails
in CI instead of during a rolling upgrade. The workflow header now counts
the jobs it actually requests.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The server-side Pre-Vote refusal test serves two connections, and each
ends in an idle retirement. Running beside the connection-outcome
tests, which count that process-wide metric exactly, it moved their
count by two and failed one of them. Take the same metrics lock those
tests hold.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The qualification fleet records the tree under test, untracked files
included, as candidate evidence. The previous release's checkout sat
inside the workspace as an untracked nested repository, which that
capture rejects as not a bounded regular file. Every mixed-release test
therefore failed before starting a fleet. Move the checkout to the
runner's temporary directory, and build its quorum node into this
tree's ignored, cached target directory.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…before forwarding

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…d path

The lost leader's black-holed path also carries the follower's Raft
calls, so counting every call made the transmission check depend on
election timing.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A forwarded call to a lost leader that black-holes its connection is
abandoned once this replica's engine names a successor. The wrapper
compared later views with the one it saw when the call began. The
caller revalidates source authority between choosing the route and
starting the call, and a successor learned during that wait was taken
as the starting view, so the call to the lost leader consumed the
caller's whole deadline and ended with an unknown outcome.

Each route now carries the leader view it was chosen under: the target
itself for a route taken from the local view, and the local view at the
time of a peer's redirect for a redirected route, which may lead away
from a leader the local view still names. A route that the local view
already supersedes is not sent at all and reports BeforeTransmission,
so the caller refreshes its route without an ambiguous outcome; one
superseded while in flight keeps the stale-route grace as before.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…re-Vote

Consume the Openraft fork at 71f7ff7ed2a0a1783362264d597926a243d62b37.
Its Pre-Vote RPC now reports, per voter, that the voter is not known to
answer Pre-Vote, and the engine then runs the classic election for that
campaign. The revision also tells Pre-Vote rounds apart by an
identifier, retries a round rejected by a running lease after the
election-timeout window, gives up an in-flight AppendEntries when its
replication stream closes, and drops the stale-candidate rule.

Both Raft adapters counted a voter that the transport could not send a
Pre-Vote to as rejecting it. A transport that does not implement
call_pre_vote reports every voter that way, so after an unplanned
leader loss no survivor could ever reach a Pre-Vote quorum, and only
initialization, which campaigns without Pre-Vote, ever elected a
leader. The adapters now report such a voter as unsupported. The
default call_pre_vote still sends nothing, so a transport written
before Pre-Vote keeps automatic failover with the classic election.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…Pre-Vote retry

The bound claimed that the survivor that heard from the leader last
campaigns last, after every other lease has expired. A survivor that
heard from the leader slightly earlier can campaign first, while a
lagging survivor's lease still runs, and the engine held that rejected
Pre-Vote round for another sampled election timeout, about 11.5 s after
the loss with these values.

The consumed engine retries a rejected round after the width of the
election-timeout window. Every lease runs out within the minimum
election timeout of the loss, so the survivor with the most up-to-date
log starts a round that no lease rejects within the minimum election
timeout, the window and one tick: the unchanged 6,800 ms. The profile
test now states that arithmetic.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…lection

A voter of this build now runs the classic election whenever it reaches
a previous-release voter, so a mixed voter set elects as the previous
release does. The bound is the longer of two paths: previous-release
voters only (their first campaign within 19 s of their last leader
contact, and one split vote of 11 s more), or a voter of this build that
can win (its first campaign after at most one cold connection, a retry
after a previous-release lease rejected it, 15.1 s). Its value stays
30 s.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…views

Pre-Vote is used only while every voter that can be reached is
positively known to answer it; otherwise that campaign runs the classic
election, so voters of a release without Pre-Vote and transports that do
not implement it never block an election. Derive the 6,800 ms first
successful campaign bound from the Pre-Vote retry after the
election-timeout window, recompute the mixed-release bounds for classic
elections during a roll, and describe how a leader route is judged
against the view it was chosen under.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The fork head only drops the doc comment of the Pre-Vote round matcher
that round identifiers replaced; the code is unchanged from 71f7ff7e.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
… again

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…observed

A route that a peer's redirect sent to B, while this replica's view
still named A, was never superseded by A. If the view advanced to B
while the caller revalidated its authority, and B was then lost while
the surviving voters elected A again, a call that B black-holed waited
for its deadline although the replica knew the successor.

The route now records each view the watcher observes, at its start
included: once the view names the target, the leader the redirect was
chosen under no longer exempts the route, and A leading again
supersedes it like any other leader.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…rs lag

Lose a leader of this build while this build's other voters miss the
last entries, so that only a previous-release voter can win, by its own
slower timers. This build's survivors run the classic election because
a previous-release voter cannot answer Pre-Vote; each such campaign
votes for itself in its next term, where a previous-release campaign
can meet it. The rolling-upgrade lane runs the 3- and 5-voter cases
against the previous release's node binary.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The fork keeps a Pre-Vote round open until its replies are due, so
voters that answer more slowly than the election-timeout window still
reach a quorum and a late unsupported report still starts the classic
election. A campaigning node that refuses a candidate with a more
up-to-date log only because of its own vote now defers its next
campaign by the greater-log timeout, so that candidate's next campaign
finds no new self-vote. Neither changes the engine's public API.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The first successful campaign after an unplanned leader loss starts
within 6,800 ms when every surviving voter answers a Pre-Vote within the
1,500 ms election-timeout window and every engine tick runs on time:
each round is then rejected or granted within that width. A Pre-Vote
round now stays open until its Vote deadline, so a slower voter delays
the election by the excess and never prevents it, and a lost leader
that never answers does not delay a round the survivors answer.

While a previous-release voter is a member, a successor is elected
within 30 s when the leader is the only voter lost, round trips take at
most one heartbeat interval and no two voters campaign at once. Each
voter whose log is behind refuses the winner at most once: a voter of
this build defers its next campaign by its greater-log timeout after
such a refusal, and a previous-release voter that has seen a greater
log campaigns at most once per 21 s. The winner therefore needs at most
one more campaign. Without the deferral, a three-process fleet whose
survivor of this build lagged elected no successor within 30 s.

Refs #1037

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant