Skip to content

Ensure that a processing job's status is always updated correctly #721

Description

@annavik

Summary

If something goes wrong with a job, we are seeing an error, but the job status still has the value "STARTED". Sometimes, we have the same problem for successful jobs (job is done, but the status is still "STARTED"). The problem affects the overall status, and also status for individual stages.

Detailed Description

Not getting the correct status makes it difficult to present correct information to users, for example adjust the status bar color and labels to highlight there is a problem.

How to Reproduce the Bug

Environment

  • OS and/or device: Mac OS
  • Browser Chrome
  • Version: Production

Screenshots

Image

Activity

  1. annavik commented on Feb 11, 2025

    @annavik
    MemberAuthor

    After doing a retry on the job above (https://antenna.insectai.org/projects/107/jobs/1244), things seem to go better, except the job never changed its status to "SUCCESS" and still looks "STARTED".

    Later, Michael manually updated the status for this job, so it now has status "SUCCESS". However last stage still has status "STARTED".

    Image
  2. annavik commented on Feb 17, 2025

    @annavik
    MemberAuthor

    This job (see https://antenna.insectai.org/projects/107/jobs/1246) also stopped due to problems, but did not had the status updated, so it still looks like it's running:

    Image
  3. mihow commented on Mar 5, 2025

    @mihow
    Collaborator

    This is related to #719

  4. self-assigned this
    on Mar 5, 2025
  5. changed the title [-]Job status is not updated correctly[/-] [+]Ensure that a processing job's status is always updated correctly[/+] on Aug 28, 2025
  6. mihow commented on Sep 4, 2025

    @mihow
    Collaborator

    This is addresses & fixed by the changes in #910

  7. annavik commented on Sep 4, 2025

    @annavik
    MemberAuthor

    I'm sorry, but I still have the same problems here. From a user perspective I wouldn't consider this as fixed. For example, this job I started this morning looks like it's running but is actually stuck. I have also recently seen jobs that has errors, but I can't understand from the status if still active or not. Can we keep this issue open until we have tested things a bit more?

    Image
  8. mihow commented on Sep 4, 2025

    @mihow
    Collaborator

    Sorry this wasn't suppose to be closed. I must have used a keyword in my PR. It will be fixed in #910

    The changes in #937 only affect the initial status being wrong, when the job is first created.

    The job status will never be updated once the background task worker disappears, so I believe that's the cause of this issue.

  9. reopened this on Sep 4, 2025
  10. annavik commented on Sep 4, 2025

    @annavik
    MemberAuthor

    Oh I see, that make sense, thanks for reopening!

  11. mihow commented on Sep 18, 2025

    @mihow
    Collaborator

    There are likely two reasons why the job status is not being updated:

    • The main celery background task for the job is disconnecting & disappearing, so the job status is never updated. I believe the fix for this is the watchdog task approach we are testing in [integration] Enable async and distributed processing for the ML backend using Celery #910. Where a scheduled "check status" task runs periodically to check on a job, or all running jobs. Rather than a single long-running task per job.
    • There are concurrency issues with the job logs & status. Some threads may be saving stale data to the job instance in the database.

    Both of these should be addressed. No. 2 may be a smaller change.

    Another observation is that the status may not be updated if a save_results subtask fails. Here is one example I just encountered. The logs show the job failing, but the status is still STARTED.

    [2025-09-17 20:09:38] ERROR Job #1686 "Hourly through sept 15" failed: Detection algorithm fasterrcnn_for_ami_moth_traps_2023 is not a known algorithm. The processing service must declare it in the /info endpoint. Known algorithms: []
    [2025-09-17 20:09:38] INFO Waiting for batch 1 to finish saving results (sub-task c8083799-4501-4a52-a3d3-cbe1ed96b465)
    [2025-09-17 20:09:38] INFO Checking the status of 6 remaining sub-tasks that are still saving results.
    [2025-09-17 20:09:38] INFO Processed 100% of images successfully.
    [2025-09-17 20:09:37] INFO Saving results for batch 6 in sub-task 79c77234-9a0e-4beb-baa0-57046ad416cd
    [2025-09-17 20:09:37] INFO Processed image batch 6 in 13.48s
    
  12. mihow commented on Sep 29, 2025

    @mihow
    Collaborator

    Plan: implement the watch dog task from #910 so that the job status can be updated even if the original background task disappeared or failed.

  13. added
    PSv2Async & distributed ML backend (PSv2): job state, NATS dispatch, result handling. Umbrella #515.
    on Jun 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

PSv2Async & distributed ML backend (PSv2): job state, NATS dispatch, result handling. Umbrella #515.backendbugSomething isn't working

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions