Skip to content

bug: detect zombie websockets and recover realtime connections - #606

Merged
lukepolo merged 3 commits into
mainfrom
bug/realtime-reconnect
Sep 29, 2026
Merged

lukepolo merged 3 commits into
mainfrom
bug/realtime-reconnect

Conversation

@lukepolo

@lukepolo lukepolo commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

After a sleep, a network switch or an idle NAT timeout, chat, presence and GraphQL subscriptions could sit on a dead ("zombie") websocket until a reload; this detects that and reconnects.

  • /ws/web: once a connection has answered a ping, one left unanswered (>10s, last pong >45s old) replaces the socket and rejoins its rooms.
  • Becoming visible or coming back online reconnects at once with a fresh backoff, even after the old 50-retry give-up, or pings a live socket to catch a zombie within ~10s.
  • Fixes a tab freeze: the offline-queue flush spun forever when the socket closed within 100ms of opening. Events sent while a replacement opens are queued, and a replaced socket's late events are ignored.
  • GraphQL: a 15s keepAlive with a pong watchdog (close 4408 + terminate()), infinite retries that also cover failed handshakes, and retry waits capped at 30s.

Merge/deploy: full effect once api#432 (pong reply) is deployed; safe to ship first, since the watchdog only arms after a pong.

Tests: the zombie socket, the offline-queue hang, the 50-retry give-up/visibility/online recovery, send-while-connecting, replaced-socket events, the heartbeat while down, the GraphQL zombie, handshake retries and the 5-retry limit each fail without the fix. Based on DEAFCS 0c7a24d, d0ab422

The /ws/web socket now expects the api's pong: once a connection has
answered a ping, an unanswered ping past the timeout replaces the socket
instead of pinging a dead TCP connection forever. Coming back to the tab
or the network reconnects straight away, including after the 50-retry
give-up, and late events from a replaced socket are ignored.

The graphql-ws client pings every 15s, closes (4408) and terminates a
socket that does not answer within 10s, and retries forever with a
capped backoff so subscriptions come back after a long sleep.
graphql-ws only retries close events by default, and a reconnect that
fails its handshake reports an error first, which left subscriptions
dead after waking offline. The retry wait keeps the library's spread
(1s + up to 3s) under the 30s cap.

A tab that comes back to a zombie is replaced ~10s after the resume
ping instead of at a later heartbeat. connect() now owns the teardown
so events queue while a replacement opens, and the offline queue flush
no longer spins forever when the socket closes before it runs.
@lukepolo
lukepolo merged commit b0173f0 into main Sep 29, 2026
2 checks passed
@lukepolo
lukepolo deleted the bug/realtime-reconnect branch September 29, 2026 01:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant