bug: detect zombie websockets and recover realtime connections - #606
Merged
Merged
Conversation
The /ws/web socket now expects the api's pong: once a connection has answered a ping, an unanswered ping past the timeout replaces the socket instead of pinging a dead TCP connection forever. Coming back to the tab or the network reconnects straight away, including after the 50-retry give-up, and late events from a replaced socket are ignored. The graphql-ws client pings every 15s, closes (4408) and terminates a socket that does not answer within 10s, and retries forever with a capped backoff so subscriptions come back after a long sleep.
graphql-ws only retries close events by default, and a reconnect that fails its handshake reports an error first, which left subscriptions dead after waking offline. The retry wait keeps the library's spread (1s + up to 3s) under the 30s cap. A tab that comes back to a zombie is replaced ~10s after the resume ping instead of at a later heartbeat. connect() now owns the teardown so events queue while a replacement opens, and the offline queue flush no longer spins forever when the socket closes before it runs.
…ry limit are fixed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
After a sleep, a network switch or an idle NAT timeout, chat, presence and GraphQL subscriptions could sit on a dead ("zombie") websocket until a reload; this detects that and reconnects.
/ws/web: once a connection has answered a ping, one left unanswered (>10s, last pong >45s old) replaces the socket and rejoins its rooms.onlinereconnects at once with a fresh backoff, even after the old 50-retry give-up, or pings a live socket to catch a zombie within ~10s.4408+terminate()), infinite retries that also cover failed handshakes, and retry waits capped at 30s.Merge/deploy: full effect once api#432 (pong reply) is deployed; safe to ship first, since the watchdog only arms after a pong.
Tests: the zombie socket, the offline-queue hang, the 50-retry give-up/visibility/online recovery, send-while-connecting, replaced-socket events, the heartbeat while down, the GraphQL zombie, handshake retries and the 5-retry limit each fail without the fix. Based on DEAFCS 0c7a24d, d0ab422