You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Stop saving a detection with another detection's classification #1436
When a batch of results contains both detections Antenna has seen before and new ones, each detection can be saved with a different detection's classification. Measured on a development database against a real processing service: 30 of 31 crops ended up carrying another crop's species label and score. A batch that is entirely new, or entirely already known, is unaffected, which is probably why this has not been noticed.
Found while working on #1407, which adds a second consumer of the same code path (feature vectors). The defect itself is older and independent of that work, so it is filed on its own.
Observation
Three real captures were sent to a processing service, which returned 31 detections with a distinct score each. Alternating detections were removed beforehand so that the batch was genuinely mixed, then create_detections and create_classifications were run exactly as save_results runs them. Everything happened inside a transaction that was rolled back.
Batch shape
Crops carrying another crop's classification
Every detection already known (no mix)
0 of 31
Alternating known and new
30 of 31
A sample of the mismatches, where "its own" is the score the service returned for that exact bounding box:
image 270: stored score 0.946286, its own was 0.950745 -> that score belongs to image 269
image 271: stored score 0.950750, its own was 0.953485 -> that score belongs to image 270
image 271: stored score 0.941504, its own was 0.951437 -> that score belongs to image 270
The pairing itself was checked separately, comparing source image id and bounding box rather than scores, and mispaired every detection in a mixed batch (4 of 4 and 6 of 6 on two runs).
Interpretation
create_detections (ami/ml/models/pipeline.py) walks the responses, sorts the rows into two lists as it goes, and returns them concatenated:
Both lists come from the same in-memory response list, so this does not depend on the processing service returning anything in a particular order: Antenna walks the responses in whatever order they arrive, then reorders its own output before pairing. A well-behaved service is affected equally.
The two orders coincide only when one of the piles is empty, or when the already-known detections happen to be the first ones in the batch. Otherwise every pair is shifted.
A second, smaller cause is in the same loop: a response whose source image cannot be found is skipped with continue and produces no row, which shortens the returned list and shifts everything after it even when nothing was re-sorted.
What produces a mixed batch
filter_processed_images skips images a pipeline has already fully processed, so re-running the same pipeline over the same images sends nothing. These cases still send a mix:
an interrupted run where boxes were saved but classifications were not (case 2 in that function's docstring)
a classifier added to an existing pipeline, so images with existing boxes are sent again alongside new ones (case 3)
a different pipeline run over the same images: the existing-detection lookup matches on (source_image, bbox) and is deliberately algorithm-agnostic, so boxes found by another detector count as already known
Directions to discuss
Carry the response and the row it produced together, so position is never used. create_detections would return pairs and both consumers would iterate them. This removes the class of bug rather than one instance, and needs no change to the API contract: each response already carries source_image_id and bbox, which is what the existing-row lookup matches on.
Alternatively, match a response to its row explicitly by (source_image_id, bbox) at each call site. Same effect, more code repeated in two places.
Whichever is chosen, the skipped-response case above should stop silently shortening the list.
What still needs verifying
How much stored data this has affected. This was measured on a development database; production has not been looked at, and I do not have access to it. A check would look for detections whose classification is inconsistent with the image, or re-run a pipeline over a known sample and compare.
Whether any verified identifications were made against a misattributed prediction. That is the case worth knowing about: agreeing with a prediction promotes it to a human-confirmed label, which is then treated as ground truth, including as training data.
Claude says: A note from #1462, which stores the feature vectors a processing service returns with each detection. It is the second consumer of the same pairing problem.
#1462 avoids positional pairing by matching each returned vector to a stored detection by capture and exact box coordinates, the same identity get_or_create_detection uses to reuse a detection. When more than one stored detection on a capture has exactly that box (for example, two detectors returning the same box in one batch), the vector for that box is skipped and logged rather than guessed. That rule is safe but loses vectors in that rare case.
The fix proposed for this issue, having create_detections return the detections in the same order as the responses, would let both writers pair by index. Once it lands, the vector writer in ami/ml/embeddings/writer.py can switch to index pairing and drop the box lookup, and its duplicate-box test becomes a test that each of the two detections keeps its own vector.
Summary
When a batch of results contains both detections Antenna has seen before and new ones, each detection can be saved with a different detection's classification. Measured on a development database against a real processing service: 30 of 31 crops ended up carrying another crop's species label and score. A batch that is entirely new, or entirely already known, is unaffected, which is probably why this has not been noticed.
Found while working on #1407, which adds a second consumer of the same code path (feature vectors). The defect itself is older and independent of that work, so it is filed on its own.
Observation
Three real captures were sent to a processing service, which returned 31 detections with a distinct score each. Alternating detections were removed beforehand so that the batch was genuinely mixed, then
create_detectionsandcreate_classificationswere run exactly assave_resultsruns them. Everything happened inside a transaction that was rolled back.A sample of the mismatches, where "its own" is the score the service returned for that exact bounding box:
The pairing itself was checked separately, comparing source image id and bounding box rather than scores, and mispaired every detection in a mixed batch (4 of 4 and 6 of 6 on two runs).
Interpretation
create_detections(ami/ml/models/pipeline.py) walks the responses, sorts the rows into two lists as it goes, and returns them concatenated:create_classificationsthen pairs that result against the responses by position:Both lists come from the same in-memory response list, so this does not depend on the processing service returning anything in a particular order: Antenna walks the responses in whatever order they arrive, then reorders its own output before pairing. A well-behaved service is affected equally.
The two orders coincide only when one of the piles is empty, or when the already-known detections happen to be the first ones in the batch. Otherwise every pair is shifted.
A second, smaller cause is in the same loop: a response whose source image cannot be found is skipped with
continueand produces no row, which shortens the returned list and shifts everything after it even when nothing was re-sorted.What produces a mixed batch
filter_processed_imagesskips images a pipeline has already fully processed, so re-running the same pipeline over the same images sends nothing. These cases still send a mix:(source_image, bbox)and is deliberately algorithm-agnostic, so boxes found by another detector count as already knownDirections to discuss
create_detectionswould return pairs and both consumers would iterate them. This removes the class of bug rather than one instance, and needs no change to the API contract: each response already carriessource_image_idandbbox, which is what the existing-row lookup matches on.(source_image_id, bbox)at each call site. Same effect, more code repeated in two places.What still needs verifying
create_classificationsis the consumer that is live today. Retrain a classifier head from verified identifications, and score what it produces #1407 addscreate_detection_embeddings, which pairs the same way; a misplaced feature vector is invisible to a reviewer, unlike a wrong species name, so this matters more once embeddings are stored.