Skip to content

Fix: faster failover after an unplanned leader loss - #10

Draft
VerifiedOrganic wants to merge 36 commits into
sdk-0.9from
work/sdk-1037-failover-timing-20261001
Draft

VerifiedOrganic wants to merge 36 commits into
sdk-0.9from
work/sdk-1037-failover-timing-20261001

Conversation

@VerifiedOrganic

@VerifiedOrganic VerifiedOrganic commented Oct 2, 2026 •

Copy link
Copy Markdown

Draft. The latest commit, 38c448d71f1fd32e762b57026a76109098691e30, repairs only the t25 election-retry fixture. Production code is unchanged by that commit. Full remote qualification is running before independent review and merge.

Faster and less disruptive failover after an unplanned leader loss, compatible with a rolling upgrade from a release without Pre-Vote. Each change is preceded by a test that fails on the previous commit.

1. Overlap the follower lease with the election timeout

A follower with a committed vote campaigned only after its leader lease (election_timeout_max) and then a full sampled election timeout. Both windows start at the last leader contact. An unplanned leader loss therefore went undetected for at least election_timeout_max + election_timeout_min, plus one tick.

  • A follower now campaigns after the longer of the two.
  • The follower lease is election_timeout_min, the minimum election timeout of the Raft dissertation (section 4.2.3). Every sampled timeout covers it, so campaigns stay randomized. While the lease is valid, a follower still rejects every other candidate.
  • RED (elect/t20): with elections in [1500, 1600) ms and a 50 ms heartbeat, the first campaign started 3.18 s after the loss against a 2.275 s bound. It now starts after about 1.6 s.

2. Give AppendEntries its own deadline

Every AppendEntries, including heartbeats and the heartbeats that confirm leadership for a linearizable read, was bounded by heartbeat_interval. A short heartbeat therefore forced a short deadline on every replication RPC.

  • New Config::append_entries_timeout: Option<u64>. Unset, it remains heartbeat_interval.
  • RED (append_entries/t62): with a 50 ms heartbeat and a 1,000 ms deadline, every AppendEntries carried 50 ms. A follower that held one AppendEntries for 300 ms received it four times. It now receives it once.

3. Pre-Vote, and the leader's quorum-acknowledged lease

A voter that is cut off from the others campaigned on every election timeout and persisted each new term. When it could be reached again, or after a restart, the leader saw the higher term in an AppendEntries response and stepped down, although a quorum had served it throughout.

  • New Config::enable_pre_vote: Option<bool> (unset means disabled), RaftNetwork::pre_vote, PreVoteReply and Raft::pre_vote. The default pre_vote reports the target as unable to answer without contacting it, so a network without Pre-Vote runs the classic election (section 5). An error, including an unreachable peer, is never a grant.
  • Answering a Pre-Vote persists nothing and changes no term, lease or server state. A quorum of grants starts the real election. A rejection with a greater log delays the next attempt, and one with a strictly higher vote catches the voter up to it in non-committed form.
  • Any vote the voter accepts afterwards, including a heartbeat, ends an in-flight round. The tick-driven re-evaluation of the server state does not.
  • A leader never renews the lease on its own vote. It now rejects Vote and Pre-Vote while a quorum acknowledged it within the lease. A planned leadership transfer releases that lease, as before.
  • RED (elect/t21): a voter cut off for four election timeouts in a 3- or 5-voter cluster went from term 1 to term 5. After a reconnect or a restart from its stores, the leader was deposed. It now keeps term 1, and the leader keeps leading.
  • RED (elect/t22): a follower that can send but not receive got the leader's grant, and the leader stepped down after about 0.6 s. The leader now keeps leading.

The design follows the upstream Openraft Pre-Vote work on main, adapted to this branch: the Pre-Vote round, adopting a higher vote from a rejection, cancelling a round once its election is won, and gating votes by the quorum-acknowledged lease.

4. Require the minimum election timeout to outlast the heartbeat tick

A leader sends heartbeats on its engine tick, every three halves of heartbeat_interval. With rule 1, a minimum election timeout at or below that tick lets every follower campaign between two heartbeats of a healthy leader. Config::validate now rejects that configuration with the existing ElectionTimeoutLTHeartBeat error, whose message names the tick.

  • Evidence: hosted CI ran the raft-kv-rocksdb example cluster test with a 250 ms heartbeat and [299, 300) ms elections, and it failed when a write reached a follower with no known leader. The example now uses [800, 1200) ms. Two test configurations that never elect are adjusted to stay valid.
  • Two example chores make the single-threaded example pass current nightly clippy (items_after_test_module) and its own rustfmt check.

5. Use Pre-Vote only while every reachable voter can answer it

During a rolling upgrade some voters run a release without Pre-Vote, and some networks do not implement it. Such a voter can grant a vote but never a Pre-Vote. Counting it as rejecting could deadlock an election: in one case the survivors with the up-to-date log needed its grant while its own log was behind; in another the voter still believed it led, because nothing it sent arrived, and never campaigned.

  • RaftNetwork::pre_vote returns PreVoteReply. A network sends the request and returns Answered only for a target positively known to answer Pre-Vote: the capability was negotiated on the connection, or the peer is an in-process peer that passes it to Raft::pre_vote. For any other reachable target it sends nothing and returns Unsupported, and the voter runs the classic election for that campaign at once. An unreachable target is an error: it grants neither request, so it neither counts as a grant nor forces the classic election.
  • The default pre_vote returns Unsupported, so a network that does not implement it always runs the classic election.
  • Each campaign decides again, so a voter set without such voters uses Pre-Vote again.
  • A Pre-Vote response never makes the requester adopt a vote for itself. No voter holds such a vote for a term the requester never campaigned in.
  • RED (elect/t24): a voter that cannot answer Pre-Vote leads, is cut off, and the others elect a successor, which is then lost too. The old leader is reachable again but nothing it sends arrives, so it never learns of the later term. The remaining voter was never elected within ten election timeouts. It is now elected with the old leader's classic vote.
  • elect/t23: voters that cannot answer Pre-Vote and miss the last entries no longer block an election with 3 or 5 voters.
  • An earlier revision of this PR added a stale-candidate rule for the second case. With this gating, elect/t23 passes without it, so it is removed and the vote rules are again those of the base release.

6. Tell Pre-Vote rounds apart by an identifier

Rounds in one term propose the same hypothetical vote, and a response was matched to the in-flight round by that vote. A grant delayed in the notification queue could reach a later round in the same term. If its sender had been removed from the membership in between, counting it panicked the Raft core.

  • Each round carries an identifier through Command::SendPreVote and the response notification, and the engine ignores a response to any round but the one in flight.
  • A grant from a node outside a round's quorum set is ignored instead of panicking.
  • RED (engine unit test): a same-term AppendEntries removes a voter and ends a round, a new round starts with the same vote, and the removed voter's grant for the first round panicked with target not in quorum set. It is now ignored.

7. Retry a rejected Pre-Vote round after the election-timeout window

A round that had not reached a quorum was held for a whole sampled election timeout. After an unplanned leader loss, a survivor whose last leader contact came shortly before another survivor's meets that survivor's running lease in its first round, and the retry came up to about twice the maximum election timeout after the loss.

  • A round rejected by a running lease is retried after the width of the election-timeout window, election_timeout_max - election_timeout_min, on a later tick. A round awaiting replies remains open until its Vote deadline so a slow answer is still considered.
  • The first-successful-campaign bound of election_timeout_max plus one tick assumes surviving voters answer Pre-Vote within that retry window and ticks run on time. Slower replies can extend election time.
  • RED (elect/t25): with elections in [1000, 1150) ms, a 20 ms heartbeat, an up-to-date voter that stopped hearing from the leader 400 ms before a lagging one, the up-to-date voter was elected 1.68 to 1.87 s after the loss against a 1.18 s bound. It is now elected after 1.0 to 1.1 s.

8. Give up an in-flight AppendEntries when replication closes

A leader that grants a vote stops leading and joins every replication stream before it answers, so no stream still reads a log range that a later truncation removes. A stream waiting for an AppendEntries answer did not notice that it was closed. With the deadline of section 2 longer than the candidate's vote deadline, a learner that never answers blocked the handoff and the Raft core for that long.

  • Closing a stream drops a watch sender that an in-flight AppendEntries races, and the request is given up at once. The entries are read before the request is sent, and a snapshot child is still cancelled and joined, so every storage operation a stream owns is still awaited.
  • RED (append_entries/t63): with a 10 s AppendEntries deadline and a learner that never answers, the vote that deposes the leader was answered after 9.04 s. It is now answered in milliseconds. A replication unit test covers the same close.

t25 fixture repair

The delayed-network job exposed a fixture race: a request admitted before the old leader was isolated could finish its send delay after the append quota was restored, catch up the lagging voter, and let that voter win legitimately. The test then waited for the wrong expected winner without reaching its timing assertion.

Keep the AppendEntries quota at zero through the election, elapsed-time assertion and term assertion, then restore it. The expected winner, contact skew, stale-log premise, 1,180 ms bound and 120 ms scheduling allowance remain intact. Vote and Pre-Vote requests do not consume the append quota.

Compatibility

  • New Config fields with defaults that keep the old behavior, except the follower lease (rule 1). That affects every deployment: the first campaign after a leader loss comes up to election_timeout_max sooner.
  • RaftNetwork::pre_vote is new in this PR; its default reports Unsupported, which keeps classic elections when Pre-Vote is enabled on a network that does not implement it.
  • Rule 4 rejects configurations whose minimum election timeout does not exceed three halves of the heartbeat interval; they were accepted before.
  • Two existing integration cases relied on the old ordering and now control automatic elections explicitly (append_entries/t61, elect/t11). Their assertions are unchanged.

Validation of 38c448d7

  • Current-nightly t25 with OPENRAFT_NETWORK_SEND_DELAY=30, RUST_TEST_THREADS=2, debug logs and full backtraces: 1 passed. The trace retains the stale log, records rejected Pre-Votes, and elects node 1 after 1.159 s, within the unchanged allowance.
  • Pinned-toolchain cargo fmt --all -- --check and focused cargo clippy --no-deps --manifest-path tests/Cargo.toml --test elect -- -D warnings pass. Local current-nightly formatting could not run because that toolchain lacks rustfmt; hosted lint remains required.
  • Full push and pull-request CI, including openraft-test (nightly, 30), must pass on this head. Earlier green results at 47ba29a6 are historical and do not qualify this candidate.

…er loss

After an unplanned leader loss a follower waits for its leader lease
(election_timeout_max) and then for a full sampled election timeout before
it campaigns. The two windows start from the same last leader contact, so
the first campaign is held back for at least
election_timeout_max + election_timeout_min.

The new case isolates the leader of a three-voter cluster with a
[1500, 1600) ms election timeout and a 50 ms heartbeat. It expects a
survivor to be elected within election_timeout_max plus one tick and a
fixed slack, and never before election_timeout_min.

On this revision it fails: a survivor is elected after 3.18 s against a
2.275 s bound, in every repeated run, with and without
single-term-leader.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A follower that held a committed vote campaigned only after its leader
lease and then a full sampled election timeout, both measured from its
last leader contact. With the lease fixed at election_timeout_max, an
unplanned leader loss went undetected for election_timeout_max plus the
sampled timeout plus one tick.

The lease and the election timeout now run in parallel: a follower
campaigns once the longer of the two has expired. The follower lease is
election_timeout_min, the minimum election timeout of the Raft
dissertation (section 4.2.3), so every sampled timeout covers it and
campaigns stay randomized. While the lease is valid a follower still
rejects every other candidate.

The heartbeat_reject_vote case now keeps its followers from campaigning
while it waits for the lease to expire, because the first campaign is no
longer held back a further election timeout after expiry.

The leader-loss case from the previous commit now elects a survivor
after about 1.6 s, within its 2.275 s bound.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Every AppendEntries RPC, including heartbeats and the heartbeats that
confirm leadership for a linearizable read, is bounded by
heartbeat_interval. A deployment that wants frequent heartbeats must
therefore give every AppendEntries a deadline no longer than the
heartbeat, and a follower that needs longer than one heartbeat to answer
is abandoned and sent the request again.

Add Config::append_entries_timeout, documented as the deadline of one
AppendEntries RPC and defaulting to heartbeat_interval. It is parsed and
reported but not used yet. The test router records the hard deadline of
every AppendEntries RPC.

The new case uses a 50 ms heartbeat and a 1,000 ms AppendEntries
deadline. On this revision every AppendEntries carries 50 ms, and with
that assertion disabled, a follower that holds one AppendEntries for
300 ms receives it four times instead of once.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Log replication, heartbeats and the heartbeats that confirm leadership
for a linearizable read now use Config::append_entries_timeout as their
hard deadline instead of heartbeat_interval. Heartbeats are still
scheduled every heartbeat interval. Unset, the deadline remains the
heartbeat interval.

The deadline case from the previous commit now passes: every
AppendEntries carries the configured 1,000 ms deadline, and a follower
that holds one AppendEntries for 300 ms receives it once.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
With single-term-leader, the first survivor to campaign after a leader
loss can be rejected by a survivor whose last leader contact was one
heartbeat later and whose lease is therefore still valid. The two then
split the next term, and the leader is elected one election timeout
later. The overlap case now measures the start of the first campaign,
a survivor's term exceeding the lost leader's term, and separately waits
for a new leader. Restored to the engine before the fix, it still fails:
the first campaign starts after 3.18 s against a 2.275 s bound, with and
without single-term-leader.

elect_seize_leadership disables heartbeats, so every follower campaigns
as soon as its lease and election timeout expire and competes with the
triggered election. It now disables automatic elections and lets the
leases expire before it triggers the election.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A voter that is cut off from consensus traffic campaigns on every
election timeout and persists each new term. When it can be reached
again, or after it restarts from its durable state, the leader sees the
higher term in an AppendEntries response and steps down, although a
quorum served it throughout.

Add Config::enable_pre_vote, documented as running a Pre-Vote round
before a real election and treated as disabled when unset. It is parsed
but not used yet.

The new cases enable it and cut one voter of a three- and a five-voter
cluster off for four election timeouts while the others keep
committing. The voter then reconnects, or restarts from its stores
while still cut off and then reconnects. They expect every voter to
stay in the leader's term, the leader to keep leading, and the voter to
catch up. A further case expects a voter whose peers are all
unreachable to keep its term, and one checks that a survivor is still
elected after the leader is lost.

On this revision the cut-off voter reaches term 5 from term 1 in all
four rejoin cases and in the unreachable case; with the rejoin checks
ordered after catch-up, the leader was deposed and re-elected in
term 6 or 7. The leader-loss case passes.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
With Config::enable_pre_vote, a voter whose election timer fires first
asks the other voters whether they would grant it a vote at term + 1.
Answering persists nothing and changes no term, lease or server state; a
voter answers by the leader-lease, last-log-id and vote rules of a real
vote request. Only when a quorum would grant does the voter start the
real election. A voter that is cut off, restarted, or behind on the log
therefore keeps its term, and does not depose a healthy leader when it
can be reached again.

- RaftNetwork::pre_vote carries the round and Raft::pre_vote answers
  it. The default implementation reports a grant without contacting the
  target, so a network without Pre-Vote keeps the election behavior
  without it. An error, including an unreachable peer, is never a grant.
- A grant from a quorum starts the real election. A rejection with a
  greater log delays the next attempt, and one with a strictly higher
  vote catches the voter up to it in non-committed form.
- A real election, a won election, prepared shutdown and any vote this
  node accepts, including a heartbeat from the current leader, end an
  in-flight round. The tick-driven re-evaluation of the server state
  does not, because it accepts no vote.
- A round is retried after one newly sampled election timeout. A single
  voter grants its own round and elects at once.
- The design follows the upstream Openraft Pre-Vote, including its later
  fixes for adopting a higher vote and cancelling a round once its
  election is won, adapted to this branch.

The cut-off voter cases from the previous commit now pass: the voter
stays in term 1 and the leader keeps leading after a reconnect or a
restart, with three and five voters, with and without
single-term-leader, and with a 30 ms send delay.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A voter that can still send but no longer receives consensus traffic
stops hearing heartbeats, so its lease expires and it runs Pre-Vote
rounds that reach the other voters. The other follower rejects them
under its own lease. The leader, however, never renews the lease on its
own vote, which expired one minimum election timeout after it was
elected, so it grants the Pre-Vote and then the real vote, and steps
down although a quorum keeps acknowledging it.

The new case blocks every RPC to one follower of a three-voter cluster
with Pre-Vote enabled and expects the leader to keep leading in its
term, the other follower to stay in that term, and the cut-off voter
never to raise its term.

On this revision the leader steps down after about 0.6 s, in every run,
with and without single-term-leader.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A leader never renews the lease on its own vote: AppendEntries
acknowledgements renew it only on followers. Its own lease therefore
expired one minimum election timeout after it was elected, and it
granted any up-to-date candidate's Vote and Pre-Vote although a quorum
kept acknowledging it.

A leader now rejects Vote and Pre-Vote requests while a quorum
acknowledged it within the leader lease. A follower that acknowledged an
AppendEntries sent at t received it no earlier than t and rejects other
candidates until at least t plus the lease, so the leader only joins
that quorum's decision. A planned leadership transfer releases the
retiring vote's lease, and its successor is still granted. A voter also
no longer starts a Pre-Vote round while its own leader lease is valid.

This follows the upstream Openraft change that gates votes by the
quorum-acknowledged lease and Pre-Vote by the local lease, adapted to
this branch.

The one-way case from the previous commit now passes: the leader keeps
leading in its term while a follower that cannot receive runs Pre-Vote
rounds, with and without single-term-leader and with a 30 ms send
delay. The full integration suites pass in every tested feature mode.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Two Pre-Vote rounds in the same term carry the same hypothetical vote, so the round check compares only the vote. Record why that is enough: every request is abandoned after election_timeout_min, the next round starts at least one sampled election timeout later, and responses and ticks share one ordered channel.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A leader sends heartbeats on its engine tick, every three halves of the
heartbeat interval. Now that a follower campaigns once the longer of its
lease and its sampled election timeout expires, instead of after their
sum, a minimum election timeout at or below that tick lets every
follower campaign between two heartbeats of a healthy leader.

Config::validate now rejects such a configuration with the existing
ElectionTimeoutLTHeartBeat error, whose message names the tick. The
raft-kv-rocksdb example used a 250 ms heartbeat with [299, 300) ms
elections; its cluster test failed in hosted CI when a write reached a
follower with no known leader. It now uses [800, 1200) ms. Two test
configurations that never elect are adjusted to stay valid.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Current nightly clippy denies items after a test module (clippy::items_after_test_module), which fails the example's lint and test job. Move the unchanged module below the store implementation.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Moving the test module left a double blank line, which the example's own rustfmt check rejects.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A network that cannot ask a voter, such as a voter of a release without
Pre-Vote, has to answer that voter's Pre-Vote itself. If its rejection
echoes the requested next-term vote, the requester adopts it as a
strictly higher vote and so votes for itself in a term it never
campaigned in, without sending any vote request. No voter can hold such
a vote, so the requester must leave it alone.

This test fails on the current engine: it adopts the echoed vote.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A Pre-Vote rejection with a strictly higher vote catches the requester
up to that vote. A vote for the requester itself in a term it never
campaigned in cannot come from a voter, only from a network that
answers for a voter it cannot ask. Adopting it voted for the requester
without sending a vote request, so its term climbed on every round
while no election ran. Leave such a vote alone, and keep the round
waiting for real answers.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
During a rolling upgrade some voters run a release without Pre-Vote.
They campaign with the real vote alone, and a network can only answer
their Pre-Vote for them, as a rejection. The test router gains that
answer. When such voters miss the last entries and the leader is lost,
the survivors with the up-to-date log cannot reach a Pre-Vote quorum
without them, and the classic voters cannot win with their stale logs.

With 3 and 5 voters, no leader is elected within ten election timeouts
on the current engine.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A voter that runs Pre-Vote could wait forever when the voters it needs
cannot answer Pre-Vote, as voters of a release without it cannot. Such
a voter campaigns with the real vote, but when its log is behind it
cannot win either, and nobody else campaigns.

A voter that rejects a vote request because the candidate's log is
behind its own, while it has no leader, now records the candidate's
term. If no leader appears within one election timeout of the first
rejection, it campaigns without Pre-Vote, in a term above every
recorded term. The rejected candidate voted for itself in its own term,
so a campaign in that term could not win its vote. The wait lets that
candidate still win with the other voters' grants without being
disrupted. Hearing from a leader, granting a vote, leading or
campaigning clears the record. No vote rule changes: a voter still
grants only a candidate whose log is at least as up to date as its own.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Explain how a network answers for a voter that cannot decode a Pre-Vote
request, why that voter counts as rejecting, and how the stale
candidate rule keeps such voters from blocking an election. Note that a
Pre-Vote response never makes the requester adopt a vote for itself.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…from a non-voter

Rounds in one term propose the same hypothetical vote, so a response was
matched to the in-flight round by that vote. A grant delayed in the
notification queue could therefore reach a later round in the same term,
and if its sender had been removed from the membership in between,
counting it panicked the Raft core.

Each round now carries an identifier through SendPreVote and the
PreVoteResponse notification, and the engine ignores a response to any
round but the one in flight. Counting a grant from a node outside a
round's quorum set is ignored instead of panicking.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A voter that cannot answer Pre-Vote, such as a voter of a release
without it, was counted as rejecting. Pre-Vote then could not reach a
quorum that needed such a voter. When that voter also believed it still
led, because nothing it sent arrived, it never campaigned either, and
no leader was ever elected.

RaftNetwork::pre_vote now returns a PreVoteReply. A network sends the
request and returns Answered only for a target that is positively known
to answer Pre-Vote: the capability was negotiated on the connection, or
the peer is an in-process peer that passes it to Raft::pre_vote. For any
other reachable target it sends nothing and returns Unsupported, and the
node runs the classic election for that campaign at once. The default
implementation returns Unsupported, so a network that does not
implement Pre-Vote always runs the classic election. An unreachable
target is still an error: it can grant neither a Pre-Vote nor a vote, so
it neither counts as a grant nor forces the classic election.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The rule let a node that rejected a candidate for its stale log campaign
without Pre-Vote, in a term above that candidate's. It existed because a
voter that cannot answer Pre-Vote was counted as rejecting, so the
up-to-date voters could never reach a Pre-Vote quorum that needed it.

Such a voter now makes the node run the classic election, so nothing
waits on its Pre-Vote, and a stale candidate's term only delays the
classic campaign as it does in any classic election: its rejection
reports that term and the next campaign is above it. The regression for
voters that cannot answer Pre-Vote and miss the last entries passes
without the rule, so it is removed and the vote rules are again those of
the base release.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…ound

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A Pre-Vote round that had not reached a quorum was held for a whole
sampled election timeout. After an unplanned leader loss, a survivor
whose last leader contact came shortly before another survivor's meets
that survivor's running lease in its first round. Holding the rejected
round for another sampled timeout delayed the first successful election
up to about twice the maximum election timeout, beyond any bound that
the maximum election timeout implies.

A round is now held for the width of the election-timeout window and
retried on a later tick. Every lease of a lost leader runs out within
the minimum election timeout of the loss, so the survivor with the most
up-to-date log starts a round that no lease rejects within the maximum
election timeout and one tick of the loss, the same bound as its first
round. In the regression, the election after the loss drops from about
1.7-1.9 s to 1.0-1.1 s, within the 1.18 s bound.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…loses

A leader that grants a vote stops leading and joins every replication
stream before it answers, so that no stream still reads a log range a
later truncation removes. A stream waiting for an AppendEntries answer
from an unresponsive follower or learner did not notice that it was
closed, so the vote answer waited for the whole AppendEntries deadline.
Once that deadline was decoupled from the heartbeat interval and could
exceed the candidate's vote deadline, a hung learner could block a
handoff and the Raft core for that long.

Closing a stream now drops a watch sender that an in-flight
AppendEntries races, and the request is given up at once. The log
entries are read before the request is sent, and a snapshot child is
still cancelled and joined, so every storage operation a stream owns
is still awaited. In the regression, the deposing vote is answered in
milliseconds instead of after 9 s of a 10 s AppendEntries deadline.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
…indow

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A Pre-Vote round that had not reached a quorum was replaced after the
width of the election-timeout window, and a reply to the replaced round
was discarded. When voters answered more slowly than that width and a
tick, though within the Pre-Vote deadline, every reply reached a round
that was already replaced: with a 1 ms window and a 30 ms tick, 65 ms
replies; with a 1,500 ms window and a 300 ms tick, 1,900 ms replies.
No round ever reached a quorum and a reachable majority stayed
leaderless, and a late report that a voter cannot answer Pre-Vote was
dropped the same way, so the classic fallback never ran.

Rounds are now tracked while they may still collect replies, until one
Pre-Vote deadline (election_timeout_min) after they started. While the
latest round awaits replies that no voter has rejected, no new round
starts. Once a voter rejects it, a new round may start after the width
of the election-timeout window, as before, and the rejected round stays
open beside it, so a late grant still completes it and a late
unsupported report still starts the classic election. Accepting a vote,
campaigning or leading ends every open round.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
A candidate whose log is behind can never win against a voter with a
more up-to-date log, but it may still campaign with the classic vote,
for instance when that voter cannot answer Pre-Vote. Each of its
campaigns persists a vote for itself in its next term. A more
up-to-date candidate that campaigned in the same term was refused, and
when the stale candidate campaigned again before the next attempt, that
attempt met another self-vote in its term. Openraft does not adopt the
term of a vote request it refuses for its log, so the up-to-date
candidate never learned to skip ahead: with three voters, it won only
its third or fourth campaign.

A campaigning node that refuses a candidate with a greater log only
because of its own vote now defers its next campaign by the greater-log
timeout from that moment, as it would after seeing a greater log in a
vote response. The candidate's next campaign then finds no new
self-vote, and this node grants it. Only timers change; no vote or log
rule does. In the regression, the up-to-date classic candidate is now
elected on its second campaign, after 5.0 to 5.3 s instead of 7.6 to
10.5 s.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
The examples run nightly rustfmt, which wraps comments; stable rustfmt
leaves them as written.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Keep AppendEntries capped until node 1 is elected and its timing and
term assertions complete. A send admitted before leader isolation can
still be delayed in the test router; clearing the quota early lets
that send catch node 2 up and changes which candidate should win.

Signed-off-by: VerifiedOrganic <verifiedorganic@sent.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant