Repository navigation
Store a feature vector for every detection and show what happened to each occurrence in one timeline - #1439
Store a feature vector for every detection and show what happened to each occurrence in one timeline#1439mihow wants to merge 49 commits into
Conversation
…ervice sends Pipeline results can now carry an `embeddings` list on each detection: one 2048-float vector per algorithm, including detections the moth/non-moth filter rejected. They are stored in a new DetectionEmbedding table, one row per (detection, algorithm), rather than as an extra classification, because an extra species classification on a rejected crop changes its occurrence's determination. - Re-saving the same results replaces the stored vector, so redelivery and reprocessing are idempotent. - Responses are matched to detections by image and box, because create_detections returns existing detections ahead of new ones. - An algorithm key the pipeline has not registered raises PipelineNotConfigured at the point an unregistered classification algorithm would. - Migration 0101 only creates the table, so it is safe on a populated database. Refs #1417 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…king and merge ranking Tracking, the merge-candidate ranking, the extend-mode preview and the "has a vector" counts read a detection's vector from its DetectionEmbedding when there is one, and otherwise from the most recent classification vector of the same algorithm. Detections the moth/non-moth filter rejected are then compared by appearance too, and older data keeps working. - One reader (models_future/embeddings.py) reads both stores in one UNION ALL query, always keyed by algorithm. - Merge ranking scores every candidate with one algorithm, the one with vectors on the most of the occurrence's scored frames, the lowest id on a tie. - Tracking loads the vectors for a pair of captures in one query instead of one per detection, through the shared greedy matcher. - has_features is true when the classification's own algorithm stored a vector for its detection; frames_with_vectors and detections_with_features count a detection with a vector from any algorithm. - A test pins that a box whose only vector is a stored embedding is scored and linked in extend mode. Refs #1417 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…tracks CSV The has_feature_vector column only looked at classification vectors, so a detection whose only vector is an embedding, such as a crop the moth filter rejected, was reported as having none. It now uses the same rule as Detection.objects.has_vector(). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…that stored it DetectionEmbedding.features_2048 is renamed to vector, so the embedding table no longer shares a column name with Classification.features_2048, which stays as it is. Each embedding now points at the job whose results stored it; saving the same detection again from another job moves the pointer, and deleting the job clears it without removing the vector. The change is folded into the 0102 migration, which has not been applied anywhere shared. See #1431. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The merge appended the test to the vector-reader test case, which has no captures fixture, so it failed with an AttributeError. It belongs with the other feature-presence tests that share its setup. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…s as history records Add OccurrenceHistoryRecord, one dated entry in an occurrence's history: either the result of an algorithm run or a person's review. Each record carries a kind, a subtype, the job, algorithm and user behind it, and a JSON payload that is validated against a pydantic schema chosen by kind and subtype, both when saved and when built for a bulk insert. Identifications and predictions keep their own tables. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…d track reviews in occurrence history Tracking, class masking and the small size filter now leave one algorithm-result history record on each occurrence a run changes, bulk-created at the end of the run. A run that changes nothing about an occurrence writes nothing for it: a chain that is already one occurrence, a classification masking leaves unchanged, or an occurrence with no flagged detection. Tracking records the settings, feature extractor, frames linked, occurrences merged in and the link costs; the post-processing tasks record the determination before and after and the detections they changed. Confirming a track's grouping posts a track_complete review with the detection ids, frame count, time span and the detections added or removed since the previous review, but only when the set differs from the last confirmed one. The confirmation fields on the occurrence stay as a cache of the latest review. Merging occurrences, by hand or by tracking, moves their history to the surviving occurrence. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…/occurrences/{id}/history/
Add a read-only history action on the occurrence viewset that returns the occurrence's
history records, identifications and predictions merged into one list, newest first.
Each entry names its source table in a type field and carries its timestamp, subtype,
algorithm, job, taxon (and the taxon before, for an algorithm result), score and a
payload. People appear by name and picture only. A prediction made by an algorithm that
also left a history record on the occurrence is left out, since the record already
stands for that change.
The endpoint is visible to exactly the people who can open the occurrence itself and
uses the same default filters as the detail view. It costs a fixed number of queries
however many entries there are, which a test pins on a multi-row fixture, alongside a
member, non-member, anonymous and superuser matrix on a public and a draft project.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…l detection is skipped The batch flush sat at the end of the loop body, after the checks that skip a detection with no box, no image dimensions or an invalid box. When the last detection in scope was skipped, the final partial batch was never written: its Not identifiable classifications were dropped and its occurrences kept their old determination. Each full batch is now written before the next begins and the remainder once the loop ends, so the history writer no longer needs to guard against occurrences that were never saved. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…rs hide The session view lists occurrences with the project's default filters off and opens their paths without them, and the history timeline is meant to sit beside it. The history endpoint still applied the score threshold and taxa filters, so an occurrence a reviewer could see and edit there returned 404. History now skips the default filters like the path and track-edit actions do, and visibility still goes through the usual object lookup. Skipping the filters also drops two queries. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…ctions Three cases left the track review history out of step with the confirmation. A re-confirmation of an unchanged set was skipped even after the confirmation had been withdrawn or when a different person made it, so the timeline kept showing the old reviewer. Reviews moved onto an occurrence by a merge were compared with its next confirmation, so it reported the occurrence's own detections as added. A session split copied the confirmation onto each piece without any review, and the earliest piece's last review still listed the other pieces' detections. A review is now skipped only when the same person re-confirms a set that is still confirmed. Each review records the occurrence it was written for, and only those reviews are compared with a later confirmation. A split restates the confirmation as a review of each piece's own detections, with the original reviewer and time. The model docstrings now describe the two confirmation fields as the current confirmation rather than a cache of the latest review. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Adds types for the occurrence history endpoint, a query hook keyed under the occurrences prefix so existing identification and track mutations refresh it, and helpers that map each history entry to the card that shows it and tell whether a track changed since it was last marked complete. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…views in one timeline The identification tab now lists everything that happened to an occurrence newest first, from the history endpoint. Identification and prediction cards keep their existing actions; tracking, class masking and size filter runs get a compact result card, and each "Mark complete" review gets a card with its frames, time span and the change since the previous review. The track panel says when a track was edited after it was marked complete. When the history cannot load, the tab falls back to identifications and predictions. See #1433. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…ermination differs An algorithm result's taxon is the occurrence's determination after the run, not the taxon the algorithm predicted, so matching the folded prediction on both dropped it whenever a person or another algorithm set a different determination. The server folds every prediction of an algorithm that wrote a record, and keeps one prediction per algorithm, so the card now matches on the algorithm alone and shows that prediction's taxon and score. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…nk algorithm page A prediction read from the history with no algorithm was given an empty Algorithm object, which the card treated as present, so clicking its title opened the algorithm page with no id. The algorithm is now left undefined, and the card only links when it has one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
… was confirmed The panel compared the latest review's detections with the occurrence's detections on the client. The review lists only valid detections, while the occurrence detail does not filter them, and the client also accepted a review with no occurrence id as its own where the server does not. Either difference could show 'Edited since marked complete' on a track nobody edited. The occurrence detail now returns grouping_edited_since_verified, computed with the same review lookup and detection queryset the server uses when it records a review, and the panel reads that field. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…ine card Confirming an identification from a prediction, algorithm result or human identification card in the occurrence timeline now calls the same callback as the header Confirm button. When the dialog is opened from the occurrences list, this moves the reviewer on to the next occurrence instead of leaving them on the one they just confirmed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…tory loads The identification and prediction cards the occurrence already carries are shown on their own while the full history loads, and the spinner appears only when there is nothing else to show. Those cards now break timestamp ties on id, newest first, the same way the server orders the history, so they keep their place when the history arrives. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…nymous Timeline cards, the track review card, the "Track marked complete by" line and the verified-by tooltip now name the viewer's own actions "You" when they have not set a display name, and label other users without one "Unnamed user". "Anonymous user" is kept for actions that have no user at all. A shared helper decides the label so these places stay consistent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The before and after determination on class masking and size filter cards is built from a translated string, and the row is left out when there was no determination before or after the run, rather than reading "N/A (unchanged)". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
… confirming person changed Undoing a confirmation and marking the same unchanged track complete again wrote a duplicate track_complete review every time. A new review is now written only when the detection set differs from the latest review or a different person confirms it; the same person confirming the same set only refreshes the current confirmation fields. Reviews carried over a session split still apply per piece. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
… history When an algorithm's classifications tied for the top score across a track's detections, the history timeline showed one identical prediction entry per detection. The timeline now keeps one prediction per algorithm, preferring the highest score, then a terminal prediction, then the most recent. This changes the history response: tied predictions collapse into a single entry. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…firms the current ID A timeline card's button reads "Confirm" when its taxon is the current determination and "Apply ID" when it would change the determination to a different taxon. Only the confirm case now moves the reviewer on to the next occurrence, matching the header Confirm button. Applying a different ID keeps the reviewer on the occurrence so they can see the result of the change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…roject summary The "Track confirmed by" line in the session capture toolbar and the top identifiers list on the project summary now use the shared user label, so a signed-in user without a display name reads as "Unnamed user" there too, matching the occurrence details. "Anonymous user" stays for a confirmation with no user attached. The capture model keeps the server's empty name instead of folding it into null, since the two cases are labelled differently. The toolbar only receives a name, not a user id, so it cannot say "You". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…udget On a branch that keeps occurrence history, merging two tracks also moves the merged track's history records to the target and cascade-deletes the rest when the merged occurrence is removed. Those are two fixed queries per merge, not per frame, so the budget rises from 29 to 31 and the test still guards against a per-frame query. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
… history loads When several frames tie for an algorithm's top score, the occurrence lists each of them, so the timeline drew two identical prediction cards until the history arrived and dropped one. The fallback list now keeps the best prediction of each algorithm by the same rule the server's history uses (highest score, then terminal, then most recent, then id), so the cards stay the same when the history loads. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…one line isort wants the shorter import on one line now that the embeddings reader replaced one of the names it imported. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
pgvector 0.5 returns a vector field as a Python list rather than a numpy array, so the embedding tests' .tolist() calls raised AttributeError. list() works for both. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for antenna-preview canceled.
|
…h repair migration and the feature-only queue rules [skip ci] Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…side its feature extractor A processing service that embeds existing boxes lists the detector whose boxes it echoes back alongside the feature extractor, so requiring every algorithm to be a feature extractor meant such a pipeline was run as a full detection pipeline. A pipeline is now feature-only when it has at least one feature extractor and no classifier, and only the extractors decide which detections still need a vector. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…line also embeds When a pipeline lists a feature extractor beside its detector and classifiers, the check for already processed captures counted the extractor as a classifier. The extractor never writes a classification, so every capture looked unprocessed and was sent again, adding a second set of classifications. Feature extractors are now left out of that check. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The results schema no longer fixes the vector length, since extractors differ and each algorithm's length is checked when vectors are saved. The schema test now pins that any non-empty vector parses and an empty one is refused. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…ithm is a detector A pipeline counted as feature-only unless one of its algorithms was typed as a classifier. A classifier registered with a blank or unknown task type therefore made its pipeline feature-only whenever it also listed a feature extractor, and saving its results would have kept only the vectors and dropped every new detection and classification. The rule now fails closed: besides its feature extractors, a feature-only pipeline may list detectors and nothing else. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The processing service sends each detection's vector under "features" (ami-data-companion#175), so the fallback that also accepted "vector" is removed. One key keeps the contract with the service unambiguous. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
A feature-only pipeline returns the boxes Antenna sent it, each with a
vector. A returned box that matches no stored detection was skipped with a
warning, so if a whole batch matched nothing the job still succeeded, no
vector was stored, and the next run sent the same detections again.
Now the job progress counts the returned boxes that matched no detection
("unmatched" on the results stage), the log names the first few, and a
batch in which no box matches raises FeatureResultsMatchNoDetections. In the
async path that exception acks the message and fails the job instead of
leaving it to be redelivered.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Resolved in tracking_task.py and merge_candidates.py: vectors are read with vectors_for_detections (embeddings first, then classification vectors), each tracking run records its history entry, and the link options from the tracking settings are passed to every scoring call. Same resolution as the integration branch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Brings in the tracking settings branch, which now also carries #1439. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…etections left without one A feature-only batch could store nothing and still end in success in two ways the previous guard missed: the service echoed the boxes without vectors, or it returned fewer boxes than it was sent (including none). Either way the same detections were queued again on every run. After each feature-only save, the detections on the batch's images that still have no vector are counted and reported on the results stage as "without_vector". A batch that stores no vector while such detections remain, or whose returned boxes all match no detection, raises FeatureResultsStoredNothing (renamed from FeatureResultsMatchNoDetections, which no longer described every case). A response that reports an error is left to the existing error handling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…rom #1439) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
| detection = models.ForeignKey(Detection, on_delete=models.CASCADE, related_name="embeddings", db_index=False) | ||
| algorithm = models.ForeignKey("ml.Algorithm", on_delete=models.CASCADE, related_name="detection_embeddings") | ||
| # No fixed length: extractors differ. Each algorithm keeps one (Algorithm.embedding_dimensions). | ||
| vector = pgvector.django.VectorField( |
There was a problem hiding this comment.
Will this support semantic embeddings? Or is the algorithm field enough? Where should we store a description of what the vector represents?
|
Can we split this PR and create one focused only on adding the DetectionEmbedding model? The features introduced in #1407 needs to use it as well |
|
Will this approach be appropriate for storing logits in their own table? Otherwise what's the plan for moving logits out of the Classification table? |
|
|
||
| # No separate index: the unique constraint's index leads with detection_id. | ||
| detection = models.ForeignKey(Detection, on_delete=models.CASCADE, related_name="embeddings", db_index=False) | ||
| algorithm = models.ForeignKey("ml.Algorithm", on_delete=models.CASCADE, related_name="detection_embeddings") |
There was a problem hiding this comment.
How can we specify the type of embedding from the algorithm? what if we store a softmax-ed copy or a reduced version of the vector after running PCA? We could use different Algorithm entries since those include post-processing. Or just notes or a enum field?
…-and-embeddings Keeps both sides: the occurrence history endpoint and timeline from this branch, and the paged detections endpoint and frame strip from the tracking UI branch. The detections and history actions both open a single occurrence and skip the detail prefetch and the identifications annotation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The note body now describes the design as decided on 1 October: two abstract bases and five tables (per-target embeddings on halfvec, AlgorithmResult for what a job decided or measured about an occurrence, capture or session, ValidationReview for what a person verified, PipelineResultsBatch for raw service responses kept in object storage, and Job as the run on every output row), the use cases each phase serves, what the research found and where the notes are, the entity diagram, ORM usage and endpoints, migration steps, six implementation phases, how #1439 splits and converges with #1407, export mappings, the alternatives that were rejected and why, and a decision log. The three appendices are folded in. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
Brings in the session lock, the confirmed-track hold-out, update_fields saves and the evaluation report changes. Tracking history is still recorded for every occurrence a chain changes and moved onto the keeper on merge; a chain that reaches a confirmed track is skipped before any history is written. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MFVhTNyVvH3Pf3oygZzRZ8
|
Claude says: The merge order and plan for tracking, agreed with the owner today, are on #1412: #1412 (comment) This PR's place: its content was split into #1461 (the occurrence history and algorithm results) and #1462 (feature vectors for every detection). |
Summary
Three things tracking needs before it leaves the development deployment, built on the tracking server (#1272) and interface (#1432) branches.
First, every detection can now carry a feature vector, including the crops the moth/non-moth filter rejects. Until now the only vectors came from the species classifier, which never sees those crops, so between a fifth and a third of the detections in the benchmark sessions could not be compared by appearance, and every frame reviewers could not attach to a track was one of them. The vectors are stored in their own table and never become a classification, so storing them on an already-annotated project changes no identification.
Second, an occurrence keeps a history of what happened to it: which tracking run grouped it, what class masking or the size filter changed, and when a person marked its track complete and accurate. The occurrence page shows that history in one timeline, beside human identifications and machine predictions, newest first. A review is dated, says who confirmed which detections, and the page can say whether the track changed since.
Third, detections that already exist can be given vectors afterwards. An ordinary ML job run with a feature-only pipeline sends the boxes Antenna already has to the processing service and stores only the vectors that come back, so a project processed months ago can gain appearance vectors without new detections, classifications or changed identifications. Vectors can be of any length, because different feature extractors produce different sizes, and each extractor keeps one length.
Merge order. This PR is based on #1432 and merges after it, and before the tracking settings PR (#1442), which is now based on this branch. It stays one PR: the vectors, the history and the feature-only jobs share one reader and one set of migrations, and splitting them would mean a second round of the same conflicts.
Checked end to end. On a local copy of a development database, three feature-only jobs run through a local processing service (RolnickLab/ami-data-companion#175) stored one 1024-float vector for each of 12,168 existing detections in three one-hour sessions, and the counts of detections, classifications and occurrences, and every determination, were unchanged.
Why it is still a draft. GitHub runs the backend tests only on PRs based on
main, so this stacked PR has no CI result. On this head (32da6cfb), the full backend suite ran locally in the CI compose stack:makemigrations --checkreports no changes, and 1,065 tests passed (2 skipped). It comes out of draft once it is based onmainand its CI is green.List of Changes
DetectionResponse.embeddings(optional list of{features, algorithm}; the vector is read fromfeaturesonly, which is what the processing service sends), stored inDetectionEmbedding(migration 0102): one row per detection and algorithm, replaced when the same results are saved again, with the job that stored it (SET_NULLwhen that job is deleted). An unregistered algorithm key stops the batch the same way it does for classificationsmodels_future/embeddings.py) reads embeddings and classification vectors in one query, always for one algorithm. Merge ranking scores every candidate with the algorithm that has vectors on most of the occurrence's framesAlgorithm.embedding_dimensions(ml migration 0029) is set from/infoor from the first vector stored; a vector of another length is refused for the whole batchembeddingorfeature_extractionalgorithm and no classifier (a detector may be listed, since the service names the detector of the boxes it echoes back). Its jobs send each capture's existing boxes, skip detections that already have a vector from that extractor, never queue a capture with no boxes, and save onlyDetectionEmbeddingrows matched by capture and box. Classifications in the response are ignoredwithout_vector), and returned boxes that match no stored detection are counted asunmatchedwith the first few named in the job log. A batch that stores no vector while such detections remain (no box matches, the boxes come back without vectors, or no boxes come back at all) raisesFeatureResultsStoredNothingand fails the job; in the async path the message is acknowledged rather than redelivered. A batch that stores some vectors but misses others does not fail; its missing detections are counted and will be sent again by the next run. A response that reports an error is left to the existing error handling. In the synchronous path the batch still fails the job, but the two counts are not added to its progresshas_feature_vectoris true for an embedding or a classification vectorOccurrenceHistoryRecord(migration 0103) with a kind (algorithm result or review) and a typed payload. Records move with the keeper when occurrences merge, and a confirmation carries over to the pieces when a regroup splits a trackGET /occurrences/{id}/history/: algorithm results, reviews, identifications and one prediction per algorithm, newest first. Visible to whoever can open the occurrence, including occurrences the default filters hide; users are shown by name and picture onlyOpen before merge
features_2048) are still written and read beside the new table. Whether new processing services should stop sending them is worth deciding here too.How to test
python manage.py test ami.ml ami.main ami.exports ami.jobs.embeddingson its detections, then check thatDetectionEmbeddinghas one row per detection and algorithm, including detections with no species classification.embeddingalgorithm), run an ML job with it on captures that already have detections, and check thatDetectionEmbeddinggains one row per real detection while the counts of detections, classifications and occurrences, and every determination, stay the same.GET /api/v2/occurrences/{id}/history/?project_id=Preturns the same entries.ui/:yarn tsc --noEmit,yarn lint,yarn test.Part of #1431. Part of #1433. Refs #1417. Based on #1432. The processing-service side of feature-only pipelines is RolnickLab/ami-data-companion#175.
🤖 Generated with Claude Code
https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8