Repository navigation
Stage Implementation Monitoring #538
Description
Activity
- changed the title
[-]Think about a way to see if a stage implementation is online or when it was last seen - Add some sort of basic healthcheck endpoint (see kubernetes /livez & /readyz pattern)? or can we check when a message was last taken from the queue by a stage?[/-][+]Stage Implementation Monitoring[/+]on Aug 22, 2024 Suggestion:
Consumers registers at the controller when coming online (project+stage, queuename). While online they periodically send "alive" (heartbeats) messages to the controller. Registration and heartbeats can be done via an HTTP API or using RabbitMQ.
The controller keeps track of all consumers.
Users can query this information using the controller API.Requirements for this to work:
- Compute clusters must be able to connect to RabbitMQ
- If heartbeats are via HTTP, then compute clusters must also be able to connect to the controller over HTTPS
Benefits of using messages for heartbeats: Consumers does not need access to the controller's API key. Downside: The messaging adds an overhead and a heartbeat message could be delayed in the queue.
- modified the milestones: Scale-up processing & data serving, Up for revision and reassignment
on Jun 13, 2025 - addedPSv2Async & distributed ML backend (PSv2): job state, NATS dispatch, result handling. Umbrella #515.Async & distributed ML backend (PSv2): job state, NATS dispatch, result handling. Umbrella #515.
on Jun 16, 2026 Claude says: Closing as completed. Processing-service / stage monitoring shipped: services self-register on startup, report liveness via a last-poll heartbeat (
mark_seen) on/tasksand/result, expose status, and are marked offline by the periodiccheck_processing_services_onlinetask (#1146). With PSv2 now the default platform-wide (#1353, deployed 2026-06-27), this is the live monitoring path. Finer per-individual-worker identity is tracked separately (#1112 / #1153 / in-flight #1194).
Think about a way to see if a stage implementation is online or when it was last seen - Add some sort of basic healthcheck endpoint (see kubernetes /livez & /readyz pattern)? or can we check when a message was last taken from the queue by a stage?