Repository navigation
Record which job created each detection and classification, and filter occurrences by job - #1471
Conversation
Detections and classifications gain a nullable `job` foreign key. Pipeline result saving, class masking and the size filter set it on the rows they create; rows that already exist keep the job they had. The detection and classification API responses expose it as a read-only id. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`?job=<id>` on the occurrence list matches occurrences with a detection or a classification written by that job, using EXISTS subqueries so each occurrence appears once. A non-integer id returns 400. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
✅ Deploy Preview for antenna-preview ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
✅ Deploy Preview for antenna-ssec ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (26)
Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThis change records the job that creates detections and classifications, exposes job choices and occurrence filtering by job, and adds UI controls and links for viewing occurrences associated with processing jobs. ChangesJob Provenance and Occurrence Filtering
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant User
participant JobFilter
participant JobViewSet
participant Occurrences
participant OccurrenceJobFilter
participant OccurrenceQuerySet
User->>JobFilter: Open job filter
JobFilter->>JobViewSet: GET jobs/choices for active project
JobViewSet-->>JobFilter: Return job choices
User->>Occurrences: Select job
Occurrences->>OccurrenceJobFilter: Request occurrences with job ID
OccurrenceJobFilter->>OccurrenceQuerySet: created_or_updated_by_job(job_id)
OccurrenceQuerySet-->>Occurrences: Return matching occurrences
Merge Risk: ⚪ Minimal · up to The job-filtering path is ready to merge after normal checks. The frontend has not yet been exercised in a browser. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The change is additive and preserves project visibility checks. No access-control regression was established. Remaining uncertainty concerns provenance consistency across job execution paths and recovery from interrupted jobs or deployment. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (2 warnings)
✅ Passed checks (3 passed)
Full details: Out of Scope Changes checkExplanation The occurrence job filter, choices endpoint, and related job UI use the provenance required by
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…ilter /api/v2/jobs/choices/ returns the pipeline and post-processing jobs of a project, most recently created first, in one response capped at 100, the same shape as the capture set choices. Failed jobs are included because they may have written results before failing. The capped pagination class is now shared under a generic name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The occurrence list gains a job filter fed by the job choices endpoint. A job selected from outside the listed choices, such as one linked from an older job's page, is loaded on its own so the filter stays visible and clearable. The job details page links pipeline and post-processing jobs to the occurrences they wrote results for. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
d9c97c6 to
8f92753
Compare
The job choices endpoint, its serializer, the shared pagination class and the job filter now point at the capture set choices they are modelled on, and the capture set side points back, so the pattern is visible from either end. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An occurrence matches the filter for every job that created one of its detections or classifications, so a later job that only adds a classification also matches it. "Created or updated" says that plainly; the queryset method is renamed to match. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…bles Replace the plain foreign key indexes on Detection.job and Classification.job with partial covering indexes, (job, occurrence) and (job, detection) where the job is set, built concurrently in their own non-atomic migration. The filter is then answered by index-only scans, the indexes leave out the rows written before jobs were recorded, and adding the columns no longer holds a lock on the two largest tables while an index builds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7
Replace the pinned 65-query count, which measured the occurrence list's existing per-row cost rather than the filter, with a check that the filter adds no queries of its own. Add the positive case to the draft project test, so a choices endpoint that hid every draft project would fail. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7
The occurrence filter list grows with the job filter, so the toggle that changes what every other filter returns moves to the top, on the species page as well for consistency. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Job choices retain unnecessary per-row database work, and the job selector needs an accessible label.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
Adds job provenance to detections and classifications, supporting occurrence filtering and future history features.
Changes:
- Records creating jobs without overwriting existing provenance.
- Adds job-filtered occurrence queries and capped job choices.
- Connects job details and occurrence filters in the UI.
| File | Description |
|---|---|
| ui/src/utils/useFilters.ts | Registers job filter metadata. |
| ui/src/utils/language.ts | Adds job-filter and navigation text. |
| ui/src/utils/getAppRoute.ts | Supports job route filters. |
| ui/src/pages/species/species.tsx | Moves default filters first. |
| ui/src/pages/occurrences/occurrences.tsx | Adds job filter controls. |
| ui/src/pages/occurrences/occurrence-filters.ts | Preserves job filters during navigation. |
| ui/src/pages/job-details/job-details.tsx | Links to matching occurrences. |
| ui/src/data-services/models/job.ts | Labels post-processing jobs. |
| ui/src/data-services/hooks/jobs/useJobChoice.ts | Loads selected jobs outside recent choices. |
| ui/src/data-services/constants.ts | Defines job choices endpoint. |
| ui/src/components/filtering/filters/job-filter.tsx | Implements job selector. |
| ui/src/components/filtering/filter-control.tsx | Registers job selector component. |
| ami/ml/tests.py | Tests pipeline and size-filter provenance. |
| ami/ml/post_processing/tests/test_class_masking.py | Tests class-masking provenance. |
| ami/ml/post_processing/small_size_filter.py | Records classification jobs. |
| ami/ml/post_processing/class_masking.py | Passes jobs to new classifications. |
| ami/ml/models/pipeline.py | Records jobs while saving results. |
| ami/main/tests.py | Tests occurrence job filtering. |
| ami/main/models.py | Adds job relationships and filtering. |
| ami/main/migrations/0097_detection_and_classification_job_indexes.py | Builds partial indexes concurrently. |
| ami/main/migrations/0096_detection_and_classification_job.py | Adds nullable job columns. |
| ami/main/api/views.py | Adds job filtering and shared choices pagination. |
| ami/main/api/serializers.py | Exposes read-only job IDs. |
| ami/jobs/views.py | Adds project-scoped job choices. |
| ami/jobs/tests/test_jobs.py | Tests choices visibility, ordering, and limits. |
| ami/jobs/serializers.py | Adds job choices serializer. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…nce migrations Rebased onto #1471 (head 797a649), which takes main/0096 and main/0097 for the job columns and their indexes. The result migrations now follow them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01APWsYj7AAYL5aiq7T3HUwq
The job choices serializer inherited the default per-object permission hook, which looked up each job's project and the user's permissions and loaded the deferred status field for a debug message: six queries per job, so 57 queries for eight jobs where three took 27. A picker needs no per-job permissions, so it now returns an empty list, as other nested serializers do, and a test pins that the query count does not grow with the number of jobs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7
Each filter's heading was a plain span with no link to its control, so a screen reader announced only the placeholder or selected value. FilterControl now gives the heading an id and marks the control and its clear button as a group labelled by it (WCAG technique ARIA17), which covers every filter at once. The default filters switch, which had no accessible name at all, is labelled by its heading. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7
…p ci] DefaultSerializer resolves object permissions for every row, so a new list or dropdown serializer can cost several queries per row without anyone noticing, as the job choices did in #1471. The endpoint checklist now says how to opt out or cache per request, and how to test for it. The systemic fix is #1475. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7
…points for per-row queries (#1474) * docs: attach PR screenshots with gh --attach [skip ci] GitHub CLI 2.99 and later uploads images referenced in a PR body and rewrites the references in place, so screenshots no longer need hosting on a branch, fork or bucket. Includes the check for production data in screenshots, since the repository is public. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7 * docs: add the per-row permission check to the endpoint checklist [skip ci] DefaultSerializer resolves object permissions for every row, so a new list or dropdown serializer can cost several queries per row without anyone noticing, as the job choices did in #1471. The endpoint checklist now says how to opt out or cache per request, and how to test for it. The systemic fix is #1475. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7 --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nce migrations Rebased onto #1471 (head 797a649), which takes main/0096 and main/0097 for the job columns and their indexes. The result migrations now follow them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01APWsYj7AAYL5aiq7T3HUwq


Summary
When a pipeline run or a post-processing task writes detections and classifications, nothing on those rows says which run made them. That makes it hard to answer simple questions after the fact, such as "what did last night's run actually change?", and it is the missing link that the occurrence history work in #1461 needs.
This PR records the job on every detection and classification a job creates, shows it in the API, and adds a job filter to the occurrence list so a user can see the occurrences a given job created or updated. An occurrence matches every job that created one of its detections or classifications, so a later job that only adds a classification also matches it. The job details page links straight to that filtered list. Rows that already exist are left alone: if a later run returns the same box or label, the row keeps the job that first wrote it. Existing data has no job recorded; only rows created after deployment will.
Requests without
?job=are unchanged: the filter returns the queryset as it is, so the occurrence list and stats endpoints run the same SQL as before.This is the first PR in the series that builds on job provenance (#1461, then #1469 and #1462), split out so the schema change can land and be reviewed on its own.
Closes #1156
List of Changes
jobforeign key onDetectionandClassification(on_delete=SET_NULL, so deleting a job keeps the rows). Migrationmain/0096_detection_and_classification_jobadds the columns only.(job, occurrence)on detections and(job, detection)on classifications,WHERE job_id IS NOT NULL. BuiltCONCURRENTLYin the non-atomic migrationmain/0097_detection_and_classification_job_indexes. Measurements below.save_resultspassesjob_idthroughcreate_detections/get_or_create_detectionandcreate_classifications/create_classification; only newly created rows get it.make_classifications_filtered_by_taxa_listtakes the task's job;SmallSizeFilterTasksetsjob=self.job.jobid (PrimaryKeyRelatedFieldwith help text) on the detection and classification serializers. It reads thejob_idalready on the row, so it adds no query.?job=<id>via a newOccurrenceJobFilterbackend andOccurrenceQuerySet.created_or_updated_by_job(), built from twoEXISTSsubqueries (a detection by the job, or a classification by the job). Human identifications are not matched. A non-integer id returns 400. Because the occurrence filter backends are shared, the occurrence stats endpoints accept the same parameter./api/v2/jobs/choices/: pipeline and post-processing jobs only, most recently created first, one response capped at 100, the same approach as the capture set choices. Failed jobs are included because they may have written results before failing. The capped pagination class is now shared asChoicesPagination.JobFilterunder "More filters". When the selected job is not among the listed choices (for example, an older job linked from its page), it is loaded on its own and added to the options.jobis added to the filters carried over into the occurrence list.FilterControlgives each heading an id and labels the control and its clear button as a group (role="group",aria-labelledby, WCAG technique ARIA17), which covers every filter. The switch getsaria-labelledby.post_processing, which previously had no label.Detailed Description
Why EXISTS and not a join
A join through
detectionsordetections__classificationsreturns one row per matching result, so an occurrence with three detections from the same job would be listed three times and the paginator count would be inflated. The.distinct()fix for that sorts whole occurrence rows and is slow on large projects. TwoEXISTSsubqueries return each occurrence once, matching how the existing?algorithm=filter is written.Schema and index choice (measured)
All figures below are single runs on one machine against a local copy of production data (about 640,000 detections and 835,000 classifications), with a warm cache. They are measurements, not a benchmark, and production tables are larger.
Adding the columns. Measured inside a transaction that was rolled back:
ALTER TABLE … ADD COLUMN job_id … REFERENCES jobs_job(job_id)index on the new, all-null column(job_id, occurrence_id / detection_id) WHERE job_id IS NOT NULLAdding a nullable foreign key column does not scan the table, so the lock it takes is brief. A plain index on the new column stores an entry for every existing row, all of them null; the partial index skips them, so it builds in a third of the time and starts empty.
Querying. To measure the filter once jobs are recorded, both tables were copied into temporary tables with every row assigned to a job, then vacuumed and analysed so index-only scans were possible. The query is the occurrence count, and the first page ordered by determination score, with only the job filter applied (not the full occurrence list query). Four cases: a job that wrote every occurrence of the largest project (179,466 detections and 197,222 classifications); a job that wrote one night in another project (1,492 detections, 888 classifications); a job that wrote only classifications, as post-processing does (876); and a job with nothing in the queried project.
EXISTSEXISTS(this PR)IN (… UNION ALL …)EXISTS … OR EXISTS …walks every occurrence in the project and checks it against the job's rows, so its cost follows the size of the project: 110 to 190 ms here even for a small job.IN (… UNION ALL …)starts from the job's rows, so its cost follows the size of the job: a few milliseconds for a night, but 650 ms for the first page of a job that wrote a whole project, because every matching occurrence has to be collected before the page can be sorted.EXISTS, whose worst case here is about 190 ms, against about 650 ms forUNION ALL. Both shapes use the same indexes, so the schema does not depend on this choice. The query lives in one method,created_or_updated_by_job(), and can change shape, or pick a shape based on the job's size, later without a migration.WHERE job_id = $1) also uses these indexes, so deleting a job does not scan either table.An earlier measurement on the real tables, with the plain index and the full occurrence list query for an anonymous request with the project's default score threshold, is kept for reference: a page of 20 took 233 ms without
?job=, 104 ms for a job that wrote one night and 240 ms for a job that wrote nothing; the count took 367 ms, 921 ms and 325 ms. Most of the filtered count was spent outside the job filter, in the list's existingLEFT JOINto detections.Migration and deployment note
0096adds the two columns. It needs a brief exclusive lock on each table, but it waits for running queries on those tables to finish, and new queries queue behind it while it waits, so avoid deploying while a long export is running.0097builds both indexesCONCURRENTLY, which blocks neither reads nor writes; it clears the statement timeout for the build and restores it afterwards, as in0093. A local database that applied an earlier version of this branch has the old plain indexes (main_detection_job_id_d3c071b8andmain_classification_job_id_bae87240); they can be dropped by hand.Tests
?job=abcreturns 400; and the filter adds no queries of its own (one query for the rows, one for the count).makemigrations --checkreports no changes, the affected test modules pass locally (105 tests), and the full backend suite passes in CI at797a6495.UI verification
The frontend passes
tsc, ESLint and Prettier, and was checked in a browser at797a6495against a local stack, on a test project where a pipeline job and a size-filter job had recorded their rows:job=<id>in the request and the page URL; clearing the filter returns all 27. The size-filter job matches all 27, since it classified every one.job, ornullfor rows without one, and?job=abcreturns 400.DialogTitlewarning, which this PR does not touch.One cosmetic issue for later: a long job name is truncated in the dropdown once selected, because the clear button takes part of the width.
Notes for the PRs that build on this
jobcolumns in its own migration. When it rebases onto this PR it should drop those twoAddFieldoperations and renumber its migrations to follow0097.?job=yet. When Record any algorithm's results on occurrences in a standard way, and show them in each occurrence's history #1461 adds algorithm results with a job,created_or_updated_by_job()should gain a third branch for them.🤖 Generated with Claude Code
https://claude.ai/code/session_01HEbXiX2mZk7efVmdgSNcN7