Skip to content

Fix null detections in exports & API. Don't mark images as processed too soon - #1312

Merged
mihow merged 11 commits into
mainfrom
fix/premptive-processed-marker
Jun 23, 2026
Merged

mihow merged 11 commits into
mainfrom
fix/premptive-processed-marker

Conversation

@mihow

@mihow mihow commented May 20, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fixes #1310.

Null detections (empty-bbox sentinels marking "image processed, nothing found") were being created before the downstream save steps inside save_results. Two consequences:

  1. Preemptive processed marker. If any downstream step raised, the image was already flagged as processed via the null marker. filter_processed_images would then skip it on retry, leaving the image permanently stuck as "processed, zero detections." Observed in production where several hundred captures had only null detections and no real ones.
  2. Phantom occurrences. create_and_update_occurrences_for_detections iterated every detection including nulls, so each null marker spawned an Occurrence with determination=NULL. Those leaked through OccurrenceQuerySet.valid() (which only excluded occurrences with zero detections, not occurrences whose only detection is a null).

Reviewer heads-up — silent semantic change to OccurrenceQuerySet.valid()

valid() changes meaning from "has any detection" to "has at least one Detection.objects.valid() row AND determination is not null." Three call sites pick this up without any line change at the call site:

  • OccurrenceViewSet.get_queryset (ami/main/api/views.py) — intended target of the fix. Phantom occurrences stop appearing in the list endpoint.
  • project summary stats occurrences_count (ami/main/api/views.py) — will silently decrease on any deployment that has accumulated phantoms. No-op on clean deployments.
  • DwC-A export (ami/exports/format_types.py) — null-determination occurrences will be excluded from exports. Probably correct (DwC requires taxonID) but not validated against an actual export run in this PR.

If any of those three are load-bearing in a way I'm missing, flag it.

Changes

  1. test(ml) — RED test for the broker-outage path: asserts the null marker is never persisted if create_detection_images.delay raises and filter_processed_images re-yields the image.
  2. fix(ml) — move null persistence to the absolute final step in save_results. Null markers now run after the source_image.save() loop, create_detection_images.delay(), update_calculated_fields_for_events, and Deployment.update_calculated_fields(save=True). Closes the silent-bug window the prior reorder left open.
  3. refactor(main) — null-marker abstraction on Detection.
    • Detection.NULL_BBOX = None — canonical sentinel value for new writes.
    • Detection.is_null_marker property — recognises both bbox=None and legacy bbox=[].
    • Detection.build_null_marker(source_image, detection_algorithm) classmethod — single construction point.
    • DetectionQuerySet.valid() — consumer default (excludes null markers).
    • DetectionQuerySet.null_markers() — narrow, for "has this image been processed?" checks.
  4. refactor(main) — sweep inline NULL_DETECTIONS_FILTER call sites to the new manager methods across ami/main/models.py, ami/main/api/views.py, ami/ml/models/pipeline.py, plus a null_detections_q(prefix) helper for relation-prefixed Q expressions.
  5. fix(main) — tighten OccurrenceQuerySet.valid() to require at least one valid detection AND a non-null determination. Closes the phantom-Occurrence leak. See the reviewer heads-up above for the consumers that pick up the new semantic.
  6. feat(main) — cleanup_null_only_occurrences management command for per-project cleanup of the field bug. Dry-run by default. Deletes phantom occurrences (no valid detections OR null determination) and dangling null-marker Detection rows on source images that have no real detection. Idempotent.

Test plan

  • test_null_marker_not_persisted_when_broker_dispatch_fails — RED, then GREEN after move-to-end.
  • TestDetectionNullMarker — is_null_marker for None / [] / real bbox, build_null_marker field setup, valid() / null_markers() disjointness.
  • TestOccurrenceValidQuerySet — fixture with real / null-only-detection / null-determination occurrences; asserts valid() returns only the real, fully-determined one.
  • TestCleanupNullOnlyOccurrencesCommand — dry-run reports without deleting; --commit deletes phantoms (both the no-real-detection arm and the null-determination arm) while preserving valid rows and null markers on images that also have a real detection; idempotent on second run.
  • Full ami/main/tests.py + ami/ml/tests.py + ami/jobs/tests/ pass locally.

Manual e2e (dev deployment)

  1. Happy path async_api job — a small collection through an ML pipeline. No new phantom occurrences.
  2. Broker-outage simulation — patched create_detection_images.delay to raise mid-job. 0 null markers persisted, 0 phantoms; image stays in the filter_processed_images yield list.
  3. Calc-field DB error — patched update_calculated_fields_for_events to raise. Same result: 0 null markers persisted.
  4. Cleanup command — dry-run on a dev deployment reported the expected phantom + dangling-null-marker counts; --commit removed exactly those rows and left valid() counts unchanged; a second dry-run reported 0 / 0 (idempotent). Running it against any affected production deployment is a post-merge ops step.

Out of scope — deferred follow-up

transaction.atomic() wrap. The persistence block (real detections → classifications → occurrences → calc-fields → null marker) can still partially commit if a mid-block step raises. This PR closes the ordering window (null marker writing before downstream steps); it does not close the within-block partial-commit window. A narrow transaction.atomic() wrap with transaction.on_commit for celery dispatch is the structural fix, deferred to a separate PR because transaction changes carry concurrency risk (see the select_for_update + ATOMIC_REQUESTS contention introduced by #1261) and need their own multi-worker e2e.

Dual-form bbox=None vs bbox=[]. New writes go through Detection.NULL_BBOX = None; legacy rows still carry bbox=[]. .null_markers() / .is_null_marker / null_detections_q() all recognise both, so no consumer breaks, but the dual form persists until a data migration backfills legacy rows. Worth a follow-up ticket.

Re-classification gap. Adjacent: filter_processed_images currently reprocesses from scratch because there is no mechanism to reclassify existing detections. Worth a separate ticket.

Summary by CodeRabbit

  • New Features

    • Added management command to clean up phantom detection records from existing data.
  • Bug Fixes

    • Improved detection pipeline handling of images with no detections to prevent orphan records.
    • Fixed API detection filtering to consistently exclude empty-detection markers for accurate results.
  • Tests

    • Added regression test coverage for empty detection scenarios and cleanup operations.

Loading
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

PSv2 Async & distributed ML backend (PSv2): job state, NATS dispatch, result handling. Umbrella #515.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Fix captures are marked as processed with zero detections when they shouldn't be

2 participants