Conversation
Removing a game server node only deleted its row, so its Node, its update, gamedata validation and streamer Jobs, and its local volume claims and volumes stayed in the cluster. NodeCleanupService removes them for a node id that has no row, is not a control plane node, and whose Node is missing or has been NotReady for at least 10 minutes. The Node goes first, then its Jobs; its claims and volumes are deleted only once the Node is gone. Every delete carries a uid precondition. A delete event trigger on game_server_nodes queues CleanupRemovedNode, which retries failed deletes and re-checks a node that is still Ready. Admins can also run the sweep with the new cleanupRemovedNodes action. A node that registers again without a row now also gets its CS:GO volume back when it reports a CS:GO build, since the cleanup deletes that volume too.
This was referenced Oct 1, 2026
Contributor
|
Closing in favor of #485, which does the same cleanup from a job that runs every 5 minutes instead of a delete trigger, retry job and manual action. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Removing a game server node in the panel only deletes its
game_server_nodesrow (the web callsdelete_game_server_nodes_by_pkstraight through Hasura). The Kubernetes objects the api created for that node stay behind:update-cs-server-<id>,update-csgo-server-<id>) and its gamedata validation Jobapp=game-streamerJob as live, so the Steam account it claimed stays claimed.local-storagevolumes (demos-,serverfiles-,serverfiles-csgo-,steamcmd-<id>) and their claimsFix
A new
NodeCleanupServiceremoves them. A node id counts as removed only when all of these hold:game_server_nodesrow. A host that is still running re-creates its row on its next ping (every 30 s) and gets its volumes back, so it keeps everything.node-role.kubernetes.io/control-planeornode-role.kubernetes.io/master). The panel's own host can run game servers too, so it is always skipped, with a warning.For each removed id, in this order:
recently_ready, and nothing else is deleted.deleteon nodes:node_delete_forbidden. Its Jobs are still deleted, but its claims and volumes are kept, because the pods on the Node would keep them terminating. A host that came back would then create no new ones, and lose them once its old pods stopped.validate-gamedataandgame-streamerJobs pinned to it. Match server Jobs are left alone.local-storagevolume pinned to it through5stack-id, when its reclaim policy isRetain: first its claims, then the volume. A claim counts only when it is in the api namespace and bound or pre-bound to that volume, and the claim that the volume'sclaimRefnames must also have the uid it records. A volume with any other reclaim policy is skipped with a warning, because releasing it can delete or scrub its data. A volume whoseclaimRefis in another namespace is skipped with a warning too, with its claims, because that claim is not the api's. Files on the node's disk are kept.Every delete sends a uid precondition, so it never hits a newer object that has the same name, and uses Background propagation. For Jobs, claims and volumes, 404 and 409 count as already gone. Objects that are already terminating are skipped. A volume whose claim could not be deleted is kept for the next run.
There are two ways in:
game_server_node_removed, queues aCleanupRemovedNodejob on the node offline queue (jobIdnode-cleanup.<id>).HasuraControlleranswers every event with a success, so an error thrown in the handler would never be retried. The job retries instead: up to 6 attempts with exponential backoff while deletes fail, or while the cluster objects or the node rows cannot be read (then nothing is deleted).DelayedError, so they do not log errors or use up attempts.Queue.removereturns 0), so the new removal then gets a job id of its own,node-cleanup.<id>.<timestamp>.cleanupRemovedNodes, sweeps every node id the cluster still has a labelled Node, a pinned Job or a local volume for. It returnsnodes,jobs,volume_claims,volumes,failed,node_delete_forbiddenandrecently_ready, which the web shows. When the cluster objects or the node rows cannot be read, it deletes nothing and returns an error. This also covers removals from before this change. The sweep logs every removed node id it skips, and how many of the removed ids it cleaned up without a failure, a skip or a forbidden Node delete.The cleanup deletes
serverfiles-csgo-<id>with the cs2 volumes, soGameServerNodeService.updateStatusnow also creates it again when a node without a row reports a CS:GO build. Without it, a host that came back after the cleanup would get only the cs2 volumes, while the panel showed its CS:GO install from the build it reports.Scope / notes
deleteonnodesto the api ClusterRole, which is a privilege increase. Without it, the Node delete gets a 403. The api then logs a warning, deletes the node's Jobs and keeps its Node, claims and volumes; the manual action also returnsnode_delete_forbidden: true, which the web shows. The automatic run does not try again, so nodes removed before the panel update need one manual cleanup afterwards. Deploy order does not matter: in the other order, nodes removed before the api update need that same run.git pull && ./update.sh). The in-app Update only restarts the Deployments, so the Node delete keeps getting a 403 until the script has run. The web's message fornode_delete_forbiddensays so.5stack-idand5stack-network-limiter, and re-creates its row and volumes. The labels that were set by hand are lost:nvidia-gpuand5stack-game-streamer(game-streamer.sh), and the ones fromcustom.shandplugin.sh. Run those scripts for it again.createVolumetakes a terminating object as present, so the ping that re-creates the row creates nothing. Update CS on that node creates the cs2 volumes again and Update CS:GO the CS:GO one, sinceupdateCsServercreates the volumes of the game it updates. On a node with CS:GO, run Update CS first: the CS:GO update also mounts the steamcmd and demos volumes. Getting there that fast takes a k3s-agent restart on the host, because its kubelet does not register the deleted Node again on its own.src/game-server-node), plus the node scheduling SQL test for the new constructor argument. The full unit suite passes. We also ran a read-only dry run of the selection rule on our test panel: it selected nothing, as expected. We had already removed that panel's leftovers by hand, and every other node on it is registered.