Skip to content

feature: clean up Kubernetes leftovers of removed game server nodes - #471

Closed
Flegma wants to merge 1 commit into
mainfrom
feature/node-removal-cleanup
Closed

Flegma wants to merge 1 commit into
mainfrom
feature/node-removal-cleanup

Conversation

@Flegma

@Flegma Flegma commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Problem

Removing a game server node in the panel only deletes its game_server_nodes row (the web calls delete_game_server_nodes_by_pk straight through Hasura). The Kubernetes objects the api created for that node stay behind:

  • its update Jobs (update-cs-server-<id>, update-csgo-server-<id>) and its gamedata validation Job
  • any game-streamer Job pinned to it. The Steam claim reconciler counts every unfinished app=game-streamer Job as live, so the Steam account it claimed stays claimed.
  • its local-storage volumes (demos-, serverfiles-, serverfiles-csgo-, steamcmd-<id>) and their claims
  • its NotReady Node entry

Fix

A new NodeCleanupService removes them. A node id counts as removed only when all of these hold:

  1. It has no game_server_nodes row. A host that is still running re-creates its row on its next ping (every 30 s) and gets its volumes back, so it keeps everything.
  2. Its Node is not a control plane node (node-role.kubernetes.io/control-plane or node-role.kubernetes.io/master). The panel's own host can run game servers too, so it is always skipped, with a warning.
  3. Its Node is missing, or has been NotReady for at least 10 minutes. A host that went down may only be restarting. Once its Node is deleted, a kubelet that is still running does not register it again until k3s-agent restarts, so a short outage must not cost it its Node.

For each removed id, in this order:

  1. The Node. Once it is gone, k8s deletes the pods it still lists on it, so the claims and volumes those pods hold can finish deleting.
    • 404: already gone.
    • 409: the uid precondition found a newer Node, so the host registered again. The id is skipped and counted in recently_ready, and nothing else is deleted.
    • 403, when the ClusterRole has no delete on nodes: node_delete_forbidden. Its Jobs are still deleted, but its claims and volumes are kept, because the pods on the Node would keep them terminating. A host that came back would then create no new ones, and lose them once its old pods stopped.
    • Any other error counts as failed and keeps the claims and volumes too.
  2. The Jobs the api runs for the node itself: both update Jobs, plus the validate-gamedata and game-streamer Jobs pinned to it. Match server Jobs are left alone.
  3. Each local-storage volume pinned to it through 5stack-id, when its reclaim policy is Retain: first its claims, then the volume. A claim counts only when it is in the api namespace and bound or pre-bound to that volume, and the claim that the volume's claimRef names must also have the uid it records. A volume with any other reclaim policy is skipped with a warning, because releasing it can delete or scrub its data. A volume whose claimRef is in another namespace is skipped with a warning too, with its claims, because that claim is not the api's. Files on the node's disk are kept.

Every delete sends a uid precondition, so it never hits a newer object that has the same name, and uses Background propagation. For Jobs, claims and volumes, 404 and 409 count as already gone. Objects that are already terminating are skipped. A volume whose claim could not be deleted is kept for the next run.

There are two ways in:

  • Automatic. A new delete event trigger, game_server_node_removed, queues a CleanupRemovedNode job on the node offline queue (jobId node-cleanup.<id>). HasuraController answers every event with a success, so an error thrown in the handler would never be retried. The job retries instead: up to 6 attempts with exponential backoff while deletes fail, or while the cluster objects or the node rows cannot be read (then nothing is deleted).
    • While the node is still Ready, or went NotReady less than 10 minutes ago, the job checks it again every 60 s, up to 20 times (about 20 minutes). The count is derived from the grace: twice its 10 minutes, so a node that goes NotReady within 10 minutes of its removal is still cleaned up on its own. After that the job logs a warning and leaves the node to the manual cleanup. The checks move the job back to delayed with DelayedError, so they do not log errors or use up attempts.
    • The handler removes an earlier waiting or failed job of the node before it queues, since a taken job id makes the add a no-op. A running job cannot be removed (Queue.remove returns 0), so the new removal then gets a job id of its own, node-cleanup.<id>.<timestamp>.
  • Manual. A new administrator action, cleanupRemovedNodes, sweeps every node id the cluster still has a labelled Node, a pinned Job or a local volume for. It returns nodes, jobs, volume_claims, volumes, failed, node_delete_forbidden and recently_ready, which the web shows. When the cluster objects or the node rows cannot be read, it deletes nothing and returns an error. This also covers removals from before this change. The sweep logs every removed node id it skips, and how many of the removed ids it cleaned up without a failure, a skip or a forbidden Node delete.

The cleanup deletes serverfiles-csgo-<id> with the cs2 volumes, so GameServerNodeService.updateStatus now also creates it again when a node without a row reports a CS:GO build. Without it, a host that came back after the cleanup would get only the cs2 volumes, while the panel showed its CS:GO install from the build it reports.

Scope / notes

  • Needs feature: let the api delete Kubernetes nodes of removed game server nodes 5stack-panel#638 for the Node delete. It adds delete on nodes to the api ClusterRole, which is a privilege increase. Without it, the Node delete gets a 403. The api then logs a warning, deletes the node's Jobs and keeps its Node, claims and volumes; the manual action also returns node_delete_forbidden: true, which the web shows. The automatic run does not try again, so nodes removed before the panel update need one manual cleanup afterwards. Deploy order does not matter: in the other order, nodes removed before the api update need that same run.
  • The ClusterRole only lands through the panel's update script (git pull && ./update.sh). The in-app Update only restarts the Deployments, so the Node delete keeps getting a 403 until the script has run. The web's message for node_delete_forbidden says so.
  • feature: clean up removed game server nodes from the server settings web#648 adds the button under Settings > Application > Servers, and a confirm step to Remove Node. Merge this PR first, because the web calls the new action.
  • A removed host that comes back registers as a new Node. Its ping restores 5stack-id and 5stack-network-limiter, and re-creates its row and volumes. The labels that were set by hand are lost: nvidia-gpu and 5stack-game-streamer (game-streamer.sh), and the ones from custom.sh and plugin.sh. Run those scripts for it again.
  • A host that registers again within about a minute of its cleanup gets no new volumes. Deleting the Node leaves its pods for the pod garbage collector, which removes them after about 40 s, and until then the claims and volumes they use stay terminating. createVolume takes a terminating object as present, so the ping that re-creates the row creates nothing. Update CS on that node creates the cs2 volumes again and Update CS:GO the CS:GO one, since updateCsServer creates the volumes of the game it updates. On a node with CS:GO, run Update CS first: the CS:GO update also mounts the steamcmd and demos volumes. Getting there that fast takes a k3s-agent restart on the host, because its kubelet does not register the deleted Node again on its own.
  • A removed host that kept running but was cut off from the control plane for more than 10 minutes loses its Node like one that went down. When the connection comes back, its kubelet does not register again, so restart k3s-agent on it.
  • Game-streamer Jobs are deleted directly, without the streamer teardown. Render batches already treat a deleted streamer Job as gone and fail the in-flight rows with "render pod no longer present (Job deleted)", and the Steam claim reconciler then frees the account. A live stream row and its Service stay until the match ends or the stream restarts, and a demo session row until its viewer stops pinging, the same as before this change.
  • Out of scope: match server, dedicated server and map-assets Jobs pinned to a removed node. A periodic sweep would also catch removals whose job gave up; that can follow.
  • Tests: unit tests for the service, the job, the controller and the volumes of a node that registers again (src/game-server-node), plus the node scheduling SQL test for the new constructor argument. The full unit suite passes. We also ran a read-only dry run of the selection rule on our test panel: it selected nothing, as expected. We had already removed that panel's leftovers by hand, and every other node on it is registered.

Removing a game server node only deleted its row, so its Node, its
update, gamedata validation and streamer Jobs, and its local volume
claims and volumes stayed in the cluster.

NodeCleanupService removes them for a node id that has no row, is not a
control plane node, and whose Node is missing or has been NotReady for
at least 10 minutes. The Node goes first, then its Jobs; its claims and
volumes are deleted only once the Node is gone. Every delete carries a
uid precondition.

A delete event trigger on game_server_nodes queues CleanupRemovedNode,
which retries failed deletes and re-checks a node that is still Ready.
Admins can also run the sweep with the new cleanupRemovedNodes action.

A node that registers again without a row now also gets its CS:GO
volume back when it reports a CS:GO build, since the cleanup deletes
that volume too.
@lukepolo

lukepolo commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Closing in favor of #485, which does the same cleanup from a job that runs every 5 minutes instead of a delete trigger, retry job and manual action.

@lukepolo lukepolo closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants