Repository navigation
Review how the state and progress of jobs are tracked #1285
Description
Activity
- added a commit that references this issue
on Apr 30, 2026 - changed the title
[-]Job state triangulation: unify DB + Redis + NATS into a diagnostic view[/-][+]Review how the state and progress of jobs are tracked[/+]on Apr 30, 2026 - addedPSv2Async & distributed ML backend (PSv2): job state, NATS dispatch, result handling. Umbrella #515.Async & distributed ML backend (PSv2): job state, NATS dispatch, result handling. Umbrella #515.
on Jun 16, 2026 Claude says: A production sighting from 2026-07-29 relevant to this issue's bug-class table, plus a distinct failure mode the table does not yet cover.
(Correcting my earlier version of this comment, which said "Redis was drained" and matched this to row 1. Both were wrong — the two halves of that sentence contradicted each other, and the correct row is a different one.)
Where it sits in the table. The signature was: the database said the job was incomplete, Redis still held exactly one outstanding entry, and the NATS consumer was fully drained and acknowledged with
ack_floor == deliveredandnum_pending = 0. That is not row 1 — row 1 requires Redis drained, and it was not; the single remaining entry is exactly why progress stopped ticking.It is closest to the lost-image stranding row, but differs in one field:
num_pendingwas zero rather than greater than zero. So the message bus had genuinely finished and the only thing outstanding was a leaked tracker entry. That looks like a new variant worth adding to the table, since "Redis non-empty but the bus is fully drained" is diagnosable and the existing rows do not name it.The concrete outcome: one large job was revoked at 99.997 percent, a single image short, because that one leaked entry stopped the progress tick and the fixed no-progress threshold fired on an effectively finished job.
A bug class the table does not cover. The same job was revoked twice earlier at zero percent while the log asserted that no workers had been seen for its pipeline in the past hour. That was false — a worker had logged 585 batches in that window.
_log_worker_availability(ami/jobs/tasks.py:165-200, warning emitted at:200, called fromrun_jobat:160) derives its claim fromProcessingService.last_seen, which is a heartbeat, not evidence of work. During a broker outage the heartbeat write path is itself a casualty, so the log reports a second symptom of the outage as though it were a fact about workers — and it points triage toward the workers, which were healthy and busy the whole time.That is distinct from the existing issues in this area, which all concern how service status is displayed. None of them says this log line can be actively wrong while looking authoritative. Worth a row of its own.
Related: #1276 and #1241 are reopened, though with narrower claims than I first made for them — see the corrections on each. The infrastructure cause underneath the incident is in #1382.
Summary
The platform tracks an async_api job's lifecycle in three independent
places —
Jobmodel in Postgres,AsyncJobStateManagerin Redis, andJetStream consumer state in NATS. Each has different failure modes, and
none alone catches every observed bug class. We should expose a unified
diagnostic view that triangulates all three.
This is a low-priority follow-up to #1276, which patched the immediate
"reaper REVOKES on clobbered
Job.progress" loop by reading Redisdirectly. Triangulation is the next layer up — forensic visibility and
catching adjacent bug classes that Redis alone can't see.
Why three sources of truth lie differently
progress.is_complete()Falseack_floor == delivered_update_job_progressnum_pending > 0, redelivery exhaustednum_redeliveredclimbingnum_ack_pending > 0for minutesDB alone catches none reliably. Redis catches most. Redis+NATS catches all.
Proposal
Three pieces, mostly independent:
Snapshot NATS state on terminal transition. On SUCCESS / FAILURE /
REVOKED, persist consumer counters into
Job.progress.diagnostics(ora new
Job.diagnosticsJSONB). Cheap. NATS consumer state is deletedpost-completion — without a snapshot the forensic trail is gone.
GET /api/v2/jobs/<id>/diagnosticsendpoint returning structuredtriangulation:
db— Job.status, Celery state, progress.is_complete(), per-stage progressredis— per-stage pending/total, failed set, all_tasks_processed, TTLnats— stream/consumer, delivered, ack_floor, num_pending, num_ack_pending, num_redelivereddivergence— server-side computed flags (DB-vs-Redis, DB-vs-NATS, Redis-vs-NATS mismatches with severity)Cache 5–10s. NATS info is an async socket round-trip; cheap but not free.
JobReconciler.diagnose()abstraction. Promote the tri-stateAsyncJobStateManager.all_tasks_processed()added in fix(jobs): fix dangling jobs from going to revoked #1276 to aricher diagnose() returning the full triangulation. Both the
diagnostics endpoint and the reaper guard call the same function.
Single source of triangulation logic.
UI diagnostic panel (admin / power-user). Keep default progress
bars unchanged. Add expandable section:
Plus header strip with consumer name, Redis TTL, divergence flags,
and a "Force reconcile" admin button that re-runs the reaper guard
on demand.
Tradeoffs
have these via the natsconn integration).
permission.
Snapshot-on-terminal mitigates this.
surface keeps cost bounded.
Recommended order
it; no UI needed yet.
JobReconciler.diagnose()— replaces inlinetri-state branch from fix(jobs): fix dangling jobs from going to revoked #1276.
Related issues
Reference
Full breakdown captured in
docs/claude/planning/job-state-triangulation.mdin the same PR series.