Repository navigation
Conversation
…yloads Local, CI and the production example env files pinned CELERY_RESULT_BACKEND to rpc://, which keeps a long-lived RabbitMQ result consumer per process. RabbitMQ drops it after the 30 s heartbeat window when idle, so the next status read or wait() fails (job creation 500s; sync ML jobs end FAILURE after all saves succeeded). Removing the pin lets the settings derive Redis DB 1 from REDIS_URL. process_nats_pipeline_result now sets ignore_result=True: nothing reads its state, and with RESULT_EXTENDED each stored result copied the full ML payload. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
✅ Deploy Preview for antenna-preview canceled.
|
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for antenna-ssec canceled.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two kinds of job failure come from the same setting. Starting a job on a stack that has been idle for a minute or two can return a server error. Long synchronous processing jobs can end as Failed even though every batch was saved. Both happen because the local, CI and example production settings pin Celery's result backend to
rpc://, which keeps task results on long-lived RabbitMQ connections. RabbitMQ closes those connections when they are idle past the heartbeat window, and the next status check runs into the closed connection.This change drops that pin. The settings already fall back to Redis database 1 when the variable is unset, which is what the deployment described in #1189 uses. It also stops one high-volume task from copying its whole ML payload into the result backend on every batch.
This is the configuration-level fix for the two bugs that #1437 and #1443 fix in code. Those two PRs are still worth merging, and the reasons are below.
List of Changes
CELERY_RESULT_BACKEND=rpc://from.envs/.local/.django,.envs/.ci/.djangoand.envs/.production/.django-example.config/settings/base.pythen derivesredis://…/1fromREDIS_URL.rpc://need the same line removed.ignore_result=Trueonprocess_nats_pipeline_result.TestProcessNatsPipelineResultStoresNoResult.Detailed Description
What was measured
On an isolated stack built from
main, with no mocks:rpc://ConnectionResetError: [Errno 104]fromJob.enqueue()readingAsyncResult(task_id).status(3 of 3)save_resultssub-task succeeded. The errors were[Errno 104]fromwait(), thenTimeoutErroron the next job (2 of 2).wait()calls without an error, including after 150 s idleami.jobsandami.mltestsRabbitMQ logged
missed heartbeats from client, timeout: 30sfor the dropped connections. Under Redis, RabbitMQ still closes idle connections, but the only remaining operation on them is publishing a task, which kombu retries. The status reads that failed now go to Redis.Why
process_nats_pipeline_resultin particularCelery task state is read in these places:
run_job: inJob.enqueue(), in the stale-job sweep, and inJob.update_status()when no status is passed.save_resultssub-tasks: in the synchronous ML job's wait.Those keep their stored results. Nothing reads the state of
process_nats_pipeline_result. The processing service receives its task id in the HTTP response but does not use it. BecauseCELERY_RESULT_EXTENDED = Truestores task arguments with each result, every stored result held a copy of the full batch payload. #1189 measured an average of 191 KB per key, with thousands of keys per async job kept for 72 hours.A side benefit: with
rpc://, a task's state cannot be read from another process, so it always reads asPENDING. The stale-job sweep therefore never sawrun_jobfinish. With Redis it sees the real state.Relationship to #1437 and #1443
Both remain useful with Redis as the result backend.
Job.enqueue()from asking the backend for the status of a task it has just sent. That status is always pending. The PR also removes a backend round trip from inside the request transaction.wait()inside a task with pollingready()and a stall limit. Celery documents callingwait()inside a task as deadlock-prone, and the old 60-second limit per batch could fail a job whose last save legitimately takes longer.How to test
python manage.py test ami.jobsredis-cli -n 1 --scan | headshows result keys forrun_jobandsave_results, but none forprocess_nats_pipeline_result.Open points
allkeys-lrueviction policy, an evicted result reads asPENDING. This is inferred, not measured. Separate instances or eviction policies are discussed in Evaluate Celery result backend and broker architecture #1189.🤖 Generated with Claude Code
https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8