Skip to content

feat(observability): serve health and readiness beside /metrics - #378

Merged
mfw78 merged 2 commits into
mainfrom
feat/health-and-readiness
Aug 26, 2026
Merged

mfw78 merged 2 commits into
mainfrom
feat/health-and-readiness

Conversation

@mfw78

@mfw78 mfw78 commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

What

One axum listener now serves /metrics, /healthz and /readyz, replacing the exporter's own hyper server which served /metrics alone and accepted no further routes.

/healthz is liveness and is static: the process is answering. It reads no module state, so a module problem can never make a Kubernetes liveness probe restart the engine.

/readyz is readiness and carries the state: 200 when at least one module is dispatchable, 503 otherwise, with the per-module detail in the body so an operator sees what the probe flattens.

ModuleState is a new public projection with four variants, Alive, Backoff, Dead and Poisoned. LifecycleState stays entirely private and so does the backoff deadline the supervisor schedules against. State reaches the endpoint through a watch channel rather than a lock on the supervisor.

The axum exemption is gone from scripts/workspace-deps-lint.sh, and with it the whole EXEMPT mechanism, which existed only for this one dependency.

Why

Closes #147

There was no /healthz and no readiness endpoint, so a Kubernetes probe had nothing to hit. Health was pub(super) and LifecycleState entirely private, so an operator could not see which modules were quarantined or backing off without reading logs.

Ready means at least one module is dispatchable

Alive counts; Backoff, Dead and Poisoned do not. A single poisoned module must not pull an engine that is still serving every other module out of rotation.

Backoff is the interesting one: not dispatchable, but expected back. It reads as not-ready for that module while the process stays ready if anything else can serve.

/metrics does not change

This was the hard constraint, since every alert rule in docs/production.md reads that endpoint.

The route calls the same PrometheusHandle::render() the exporter's own server called, so the exposition is produced by the identical code path rather than by a reimplementation that happens to match. the_latency_histogram_renders_bucket_series still asserts the # TYPE line and the le="5" bucket series against a real recorder, so the bucket configuration is pinned rather than assumed.

Same path, same port, same bucket configuration. The [engine.metrics].enabled = false behaviour is preserved: the recorder installs so call sites stay live, and nothing binds.

Public surface

ModuleState is #[non_exhaustive], so adding a fifth state later is not a breaking change for a downstream match. It derives strum::IntoStaticStr rather than carrying a hand-written label match.

Five items are promoted in total: ModuleState, HealthSnapshot, HealthWatch, HealthPublisher and health_channel. That is more than one, but each is part of one seam, and the alternative was promoting LifecycleState wholesale, which would have exposed the backoff deadline and made every internal state change a public API change. #145 makes these permanent, so they are documented as such.

Review found a listener bug

The first revision refused the launch on some listener failures and not others. It now refuses on every one, which matters because a bind failure that is swallowed leaves an engine running with no observability endpoint and no indication why, and an operator reading supervisor ready would reasonably conclude the probe was working.

Testing

/readyz semantics are pinned by test for each case rather than left to the reader: no modules configured, every module in backoff, every module dead, and a mix.

Workspace clippy, rustdoc, doctests and nextest green; just build and just test-e2e green; workspace-deps-lint.sh passes with the exemption removed, which is the check that proves axum is genuinely inherited now.

Merge note

#377 is open in parallel and also adds a row to the metrics table in docs/production.md. The two conflict only there and in crates/nexum-runtime-metrics/src/lib.rs if both add an entry; whichever lands second takes a trivial rebase.

AI Assistance

Implementation: claude-opus-5. Red-team review: claude-opus-5. Verification: claude-opus-5. PR description: claude-opus-5.

mfw78 added 2 commits August 26, 2026 07:20
Closes #147.

The engine had no /healthz and no readiness endpoint, so a Kubernetes probe had nothing to hit, and the supervisor's Health was pub(super) with LifecycleState entirely private, so an operator could not see which modules were quarantined or backing off without reading logs.

Replace the exporter's own hyper listener with one axum server bound from the same [engine.metrics] section, serving /metrics, /healthz and /readyz on one port. The exposition path, port, format and enabled = false semantics are unchanged; the recorder still installs when no listener binds. This consumes the axum dependency that had sat in [workspace.dependencies] with no inheritor, so its exemption in scripts/workspace-deps-lint.sh goes with it, and metrics-exporter-prometheus drops the http-listener feature.

Ready means at least one module is dispatchable. Alive counts; Backoff, Dead and Poisoned do not, so a single poisoned module cannot pull an engine that is still serving every other module out of rotation. /readyz carries the per-module detail the aggregate flattens, and answers 503 with no module lines before the supervisor has published, which is how a probe tells starting apart from degraded.

Promote a payload-free ModuleState rather than LifecycleState wholesale: the backoff deadline the supervisor schedules against stays internal. The supervisor publishes a snapshot over a watch channel once per event-loop iteration and on the exit path, and that publication also writes nexum_runtime_module_state{module,state}. nexum_runtime_module_poisoned stays as its own series so the NexumModulePoisoned alert keeps reading one metric.

RunEnd now derives strum::IntoStaticStr and every return path counts nexum_runtime_run_end_total{reason}, which makes the abnormal stream_ended end readable from a scrape; #317 still owns the exit code.

AI Assistance: Claude Code used for implementation, tests and documentation.
…pin the gauge and reason labels

The reactor handoff ran inside the spawned task, so a `from_std` failure logged
"observability listener never started" and left a running engine with a dead
`/metrics`, `/healthz` and `/readyz`, which is the outcome the synchronous bind
exists to prevent. Do the handoff in `bind`, where it joins the parse and the
bind as a launch refusal, and cover the taken-port path with a test that
asserts the refusal precedes the recorder install.

The `nexum_runtime_module_state` invariant that a transition clears the state
it left had no test: the existing one asserts the watch change flag, not the
gauge. Assert the four series through `capture_metrics`, which the tree already
ships for exactly this. Pin the `nexum_runtime_run_end_total` reasons the same
way, since `NexumEventLoopDied` reads `reason="stream_ended"` and a renamed
variant would stop the alert firing without failing to compile.

Move the gauge writes out of the `send_if_modified` closure, which holds the
watch's write lock that every `/readyz` read waits on.

`metrics_still_renders_the_prometheus_exposition` passed on an empty body, so
it asserted nothing. Record and describe a metric first.

`MetricsSection` still told operators the exporter serves `/metrics` alone, and
the Kubernetes probe block in docs/production.md contradicts the loopback bind
the same section mandates, because the kubelet dials the pod address.

AI Assistance: Claude Code used for the adversarial review of
feat/health-and-readiness and for these fixes.
@mfw78
mfw78 force-pushed the feat/health-and-readiness branch from cb89f5b to 6e53da1 Compare August 26, 2026 07:29
@mfw78
mfw78 merged commit fd16db7 into main Aug 26, 2026
7 checks passed
@mfw78
mfw78 deleted the feat/health-and-readiness branch August 26, 2026 08:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

observability: expose health and readiness

1 participant