Repository navigation
Conversation
✅ Deploy Preview for antenna-preview ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
mihow
force-pushed
the
feat/add-feature-vectors-task
branch
from
October 6, 2026 19:26
a501e8f to
f1aea0b
Compare
Collaborator
Author
|
Claude says: Notes on how this task fits the plan for the default pipeline and the other open PRs, agreed in a planning discussion today.
|
This was referenced Oct 6, 2026
Open
mihow
force-pushed
the
feat/add-feature-vectors-task
branch
from
October 6, 2026 23:02
f1aea0b to
93b59f1
Compare
…pipelines in ML jobs A feature-only processing service returns the detections it was sent, unchanged and still naming the detector, with an embedding attached. The new save_embedding_results() stores only those vectors on the stored detections they match by capture and box, so it can never create a detection, classification or occurrence, and it ignores the echoed detector reference. An embedding from an algorithm outside the pipeline still raises, and the per-(algorithm, key) length check still applies. Unmatched boxes are counted in the result instead of created. Pipeline gains embedding_algorithms() and is_embedding_only(). A regular ML job (and Pipeline.process_images) now refuses an embedding-only pipeline up front with a message pointing to the "Add feature vectors" admin action, instead of failing later in save_results with PipelineNotConfigured. process_detections() sends a chosen set of stored detections in one synchronous request, sharing the request building with collect_detections(). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…etections Admins can now select capture sets and run "Add feature vectors" with a pipeline that has an embedding algorithm. Only valid detections that lack a vector from that extractor are sent, in batches of captures, to the pipeline's processing service, and the returned vectors are stored on those detections. Captures with nothing missing are skipped, so a second run sends nothing. One job tracks the run and reports captures, detections sent, vectors stored, vectors unchanged and boxes unmatched on its stage. The task is registered like the other post-processing tasks and triggered through make_post_processing_action; no new job type, migration or UI is involved. The "still missing" reads bypass the query cache, which otherwise served the answer from before the vectors were stored. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The reader that finds detections without a vector now invalidates correctly when vectors are stored, so the add-feature-vectors task no longer needs to bypass the query cache for those reads. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The three test classes for the add-feature-vectors task rebuilt their project, pipeline and service in setUp for every test. They now build the data once in setUpTestData; only the fake processing-service session and the status-check stub stay per test. Measured on the 16 tests in test_feature_vectors.py: 13.5 s before, 3.6 s after. The test count and results are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
mihow
force-pushed
the
feat/add-feature-vectors-task
branch
from
October 7, 2026 23:01
93b59f1 to
680ffd3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Pipelines that return feature vectors only fill them in for detections they create from now on. This PR lets an admin add vectors to detections that already exist, without running the detector or the classifiers again, so existing projects can use tracking's appearance matching, retraining and similarity search. It builds on the vector storage in #1462 and adopts the idea from #1407 of sending stored detections to the processing service and keeping only the vectors that come back.
It is admin only on purpose: an "Add feature vectors" action on capture sets in the Django admin, built like the class masking and size filter actions. There is no new job type, no migration and no interface change; the job form (#1447) can offer it later through the post-processing registry.
List of Changes
AddFeatureVectorsTask(ami/ml/post_processing/feature_vectors.py), registered post-processing task; capture set admin action built withmake_post_processing_action; settings form lists only pipelines with a feature-vector algorithmdetections_missing_vectorsfrom #1462 selects them per (algorithm, key); captures with nothing missing are skippedsave_embedding_results(ami/ml/embeddings/writer.py) matches returned detections to stored ones by capture and box, ignores the echoed detector reference, counts unmatched boxes instead of creating them, and reuses the shared store and length checkprocess_detections(ami/ml/models/pipeline.py) sends stored detections in one synchronous request; the per-detection request builder is shared with the existing request path, whose output is unchangedPipeline.raise_if_embedding_only()in the ML job, with a message pointing to this admin actionRelated Issues
Part of #1464 (the synchronous path; the queued path needs the processing service's worker to run feature-only pipelines). Stacked on #1462. Processing-service side: RolnickLab/ami-data-companion#175, whose feature-only pipelines skip the detector when detections are sent.
Detailed Description
/processdirectly, one batch of captures per request.Exists).Direction and follow-ups
source_image_collection_idandpipeline_idwithreference()(a capture set and a pipeline), adding a "pipeline" reference type and its UI route, so the history shows names and the jobs dialog renders pickers.result_modelsstays empty: the task writes vectors, not algorithm results.group = "process", and either a feature flag or superuser-only access.DetectionEmbeddingrather than keep its own.How to Test
vector_counts_by_algorithm(<project id>)in a shell); the counts of detections and classifications do not change. Running it again sends nothing.python manage.py test ami.ml.post_processing.tests.test_feature_vectors(16 tests). Full backend suite: 809 tests OK, 2 skipped.🤖 Generated with Claude Code
https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy