Skip to content

Add feature vectors to existing detections from the admin, without detecting again - #1479

Draft
mihow wants to merge 4 commits into
feat/detection-embeddings-taskfrom
feat/add-feature-vectors-task
Draft

mihow wants to merge 4 commits into
feat/detection-embeddings-taskfrom
feat/add-feature-vectors-task

Conversation

@mihow

@mihow mihow commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Pipelines that return feature vectors only fill them in for detections they create from now on. This PR lets an admin add vectors to detections that already exist, without running the detector or the classifiers again, so existing projects can use tracking's appearance matching, retraining and similarity search. It builds on the vector storage in #1462 and adopts the idea from #1407 of sending stored detections to the processing service and keeping only the vectors that come back.

It is admin only on purpose: an "Add feature vectors" action on capture sets in the Django admin, built like the class masking and size filter actions. There is no new job type, no migration and no interface change; the job form (#1447) can offer it later through the post-processing registry.

List of Changes

# Change (effect) How (implementation)
1 Admins can add feature vectors to the existing detections of a capture set AddFeatureVectorsTask (ami/ml/post_processing/feature_vectors.py), registered post-processing task; capture set admin action built with make_post_processing_action; settings form lists only pipelines with a feature-vector algorithm
2 Only detections still missing a vector are sent, so a second run sends nothing detections_missing_vectors from #1462 selects them per (algorithm, key); captures with nothing missing are skipped
3 Only vectors are written: no detections, classifications or occurrences change save_embedding_results (ami/ml/embeddings/writer.py) matches returned detections to stored ones by capture and box, ignores the echoed detector reference, counts unmatched boxes instead of creating them, and reuses the shared store and length check
4 The request carries exactly the detections to embed process_detections (ami/ml/models/pipeline.py) sends stored detections in one synchronous request; the per-detection request builder is shared with the existing request path, whose output is unchanged
5 A regular ML job cannot be started with a feature-only pipeline by mistake Pipeline.raise_if_embedding_only() in the ML job, with a message pointing to this admin action
6 Progress shows what happened one stage with captures, detections sent, vectors stored, vectors unchanged and boxes unmatched

Related Issues

Part of #1464 (the synchronous path; the queued path needs the processing service's worker to run feature-only pipelines). Stacked on #1462. Processing-service side: RolnickLab/ami-data-companion#175, whose feature-only pipelines skip the detector when detections are sent.

Detailed Description

  • Synchronous only. The processing service's queued worker does not run feature-only pipelines yet, so the task posts to /process directly, one batch of captures per request.
  • Captures without a public image URL are skipped and their detections are not sent; the counts on the stage show how many were sent.
  • Query cache: the "what is still missing" read depends on Store feature vectors from any model for every detection, indexed for the ways they are read #1462's fix that lets new vectors invalidate it (django-cachalot does not track tables inside a negated Exists).

Direction and follow-ups

How to Test

  1. Run a feature-only pipeline from ami-data-companion Bump react-admin from 4.8.4 to 4.11.4 in /frontend #175 (for example the BioCLIP 2.5 embeddings pipeline) and register its processing service in Antenna.
  2. In the Django admin, select a capture set with processed captures, choose "Add feature vectors", pick that pipeline and submit.
  3. The job stores vectors for the set's detections (vector_counts_by_algorithm(<project id>) in a shell); the counts of detections and classifications do not change. Running it again sends nothing.
  4. Tests: python manage.py test ami.ml.post_processing.tests.test_feature_vectors (16 tests). Full backend suite: 809 tests OK, 2 skipped.

🤖 Generated with Claude Code

https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy

@netlify

netlify Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for antenna-preview ready!

Name Link
🔨 Latest commit 680ffd3
🔍 Latest deploy log https://app.netlify.com/projects/antenna-preview/deploys/6ac6cf40ea05d00007235fc7
😎 Deploy Preview https://deploy-preview-1479--antenna-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
Lighthouse
Lighthouse
1 paths audited
Performance: 50 (🔴 down 15 from production)
Accessibility: 81 (🔴 down 8 from production)
Best Practices: 92 (🔴 down 8 from production)
SEO: 92 (no change from production)
PWA: 80 (no change from production)
View the detailed breakdown and full score reports
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@mihow
mihow force-pushed the feat/add-feature-vectors-task branch from a501e8f to f1aea0b Compare October 6, 2026 19:26
@mihow

mihow commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Claude says: Notes on how this task fits the plan for the default pipeline and the other open PRs, agreed in a planning discussion today.

  1. Declare the id settings as references. source_image_collection_id and pipeline_id are plain ints, so the occurrence history (Record any algorithm's results on occurrences in a standard way, and show them in each occurrence's history #1461) will show them as bare numbers, and the jobs dialog (Create any job from one dialog generated from each job type's own settings model, & simplify the job creation API #1447) will neither render pickers for them nor check that they belong to the job's project. Declaring them with reference("capture_set", ..., title=...) and a pipeline reference fixes both. Record any algorithm's results on occurrences in a standard way, and show them in each occurrence's history #1461's reference types do not include pipelines yet, so this needs a "pipeline" entry in REFERENCE_TYPES and its route in ui/src/utils/references.ts. The plan is for Create any job from one dialog generated from each job type's own settings model, & simplify the job creation API #1447 to read the same declaration for its pickers (see the comment there).
  2. Where it sits. This task sends existing detections to a processing service, so in the jobs dialog it belongs with "Process images", not with the methods that refine results inside Antenna. Longer term it becomes a mode of the ML job, "use existing detections, skip the detector" (Add feature vectors to existing detections, and reprocess detections without re-detecting #1464). That same mode covers re-classifying existing detections with a new classifier, and it could keep the classifications the response already carries, which this task currently discards.
  3. Outputs. Post-processing tasks now declare the result models they write (result_models in Record any algorithm's results on occurrences in a standard way, and show them in each occurrence's history #1461). This task writes vectors rather than algorithm results, so it leaves that empty; a later, general declaration of a stage's outputs would list feature vectors.
  4. A backfill, not the main path. For the default pipeline, new captures should get their vectors in the same pass as detection and classification. Workers on the queue-based (async) path emit vectors only when configured to, so that setting should default to on; this task then fills in older detections.
  5. One job and one table for vectors. Other work that needs embeddings, such as the classifier retraining work, can reuse this task and DetectionEmbedding rather than adding its own job type.
  6. When Create any job from one dialog generated from each job type's own settings model, & simplify the job creation API #1447 lands, the task will also need its group (Process images) and a feature flag, or it stays superuser-only as it is now.

mihow and others added 4 commits October 7, 2026 16:00
…pipelines in ML jobs

A feature-only processing service returns the detections it was sent, unchanged and
still naming the detector, with an embedding attached. The new
save_embedding_results() stores only those vectors on the stored detections they
match by capture and box, so it can never create a detection, classification or
occurrence, and it ignores the echoed detector reference. An embedding from an
algorithm outside the pipeline still raises, and the per-(algorithm, key) length
check still applies. Unmatched boxes are counted in the result instead of created.

Pipeline gains embedding_algorithms() and is_embedding_only(). A regular ML job
(and Pipeline.process_images) now refuses an embedding-only pipeline up front with
a message pointing to the "Add feature vectors" admin action, instead of failing
later in save_results with PipelineNotConfigured. process_detections() sends a
chosen set of stored detections in one synchronous request, sharing the request
building with collect_detections().

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…etections

Admins can now select capture sets and run "Add feature vectors" with a pipeline
that has an embedding algorithm. Only valid detections that lack a vector from that
extractor are sent, in batches of captures, to the pipeline's processing service,
and the returned vectors are stored on those detections. Captures with nothing
missing are skipped, so a second run sends nothing. One job tracks the run and
reports captures, detections sent, vectors stored, vectors unchanged and boxes
unmatched on its stage.

The task is registered like the other post-processing tasks and triggered through
make_post_processing_action; no new job type, migration or UI is involved. The
"still missing" reads bypass the query cache, which otherwise served the answer from
before the vectors were stored.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The reader that finds detections without a vector now invalidates correctly when
vectors are stored, so the add-feature-vectors task no longer needs to bypass the
query cache for those reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The three test classes for the add-feature-vectors task rebuilt their
project, pipeline and service in setUp for every test. They now build
the data once in setUpTestData; only the fake processing-service
session and the status-check stub stay per test.

Measured on the 16 tests in test_feature_vectors.py: 13.5 s before,
3.6 s after. The test count and results are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
@mihow
mihow force-pushed the feat/add-feature-vectors-task branch from 93b59f1 to 680ffd3 Compare October 7, 2026 23:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant