Repository navigation
Conversation
… Antenna instances Four management commands and the modules behind them. export_validated_occurrences writes the occurrences people confirmed or identified in a project to a versioned JSON bundle keyed by natural keys only (capture path, timestamp, station, box, detector; taxon by GBIF key or name and rank; user by email). import_validated_occurrences replays a bundle onto another project: it finds each detection again (exact box, else IoU >= 0.7), regroups confirmed occurrences with the track-edit operations the review interface uses, records each confirmation under its original reviewer and time, and re-creates the identifications with their user and timestamp. Partially found occurrences are reported and never confirmed; a second run changes nothing. export_embeddings and import_embeddings do the same for detection feature vectors (vectors.npy plus an index of detection keys), writing DetectionEmbedding rows where that table exists. The confirmation write is isolated in confirm_grouping_as_of so it can become a ValidationReview row in a few lines once the model-outputs schema lands. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
✅ Deploy Preview for antenna-preview canceled.
|
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…sewhere A replayed confirmation must keep its original reviewer time. Instead of stamping now and overwriting the field afterwards, verify_grouping accepts an optional timestamp, so the import goes through the same function the review interface uses and whatever that function records (today the cached fields, later a review row as well) carries the original time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…ow is missing Two problems found on a trial merge with the model-outputs branch. The vector export read only DetectionEmbedding when that table existed and so wrote nothing for data whose vectors live on the classifications; it now carries both stores, one matrix file per source since their lengths differ, and the import writes each back to its own home (embedding rows, or features_2048 on the matching classification, reported as skipped when the target has no such classification). The "already confirmed" check compared only the cached grouping_verified_at/by, so an occurrence confirmed before reviews existed never got its ValidationReview row; where that model exists the check now also requires a standing grouping review by that person at that time, and re-confirms through verify_grouping when it is missing. On a branch without reviews the check stays cache-only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…oth stores The matcher claims a target detection for one key, which is right between two boxes and wrong between two index rows of the same box. The vector import now matches the distinct keys and maps every row to its key's match. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
… instead of assuming it The outputs stack lands in pieces (results and reviews, then embeddings, then tracking), so a branch may have any subset of the tables and columns this code reads. The vector transfer now checks for the classification feature column as it already did for the embeddings table: an export notes a store it could not read in its manifest, and an import counts rows it has nowhere to put instead of failing. Embedding rows are written through the model's own insert-mostly writer where it has one. The confirmation check requires the grouping aspect to exist on the review model, not just the model, and falls back to the cached fields otherwise. The vector tests move to their own module so they run on a branch without the tracking code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Collaborator
Author
|
Claude says: The merge order and plan for tracking, agreed with the owner today, are on #1412: #1412 (comment) This PR's place: it waits for the review PRs that follow #1469. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
When a project is re-imported or re-processed, every capture, detection and occurrence gets a new id, and the work reviewers did on the old rows is lost: which detections they confirmed as one individual, and which taxon they identified each occurrence as. This PR adds a way to carry that work from one Antenna database to another. A project's confirmed occurrences and identifications are exported to a bundle that names everything by stable facts (capture path, time and station, bounding box, detector name, taxon GBIF key, reviewer email), and that bundle can be replayed onto a fresh project where the ids are different. The same is done for detection feature vectors, which are expensive to recompute. This is the replay route chosen over migrating the draft tables for the model-outputs schema change (see #1412 and #1208).
List of Changes
export_validated_occurrences --project N --output file.json;models_future/validated_occurrences.pybuilds a versioned bundle keyed by natural keys only (DetectionKey,TaxonKey, user email).import_validated_occurrences --project N --input file.json [--execute] [--iou-threshold 0.7] [--create-missing-detections] [--report out.json]. The dry run only reads and prints exact / overlap / missing counts.models_future/detection_matching.py: capture by path (fallback station + timestamp), box exact within 0.5 px, else highest IoU above the threshold; each target detection is claimed once. Two queries per batch.tracks.detach_detectionandtracks.add_detections, so chain links, track statistics and session counts stay right. An occurrence with a missing detection is reported as partial and never confirmed.verify_grouping()gains an optionaltimestampfor a confirmation made elsewhere;confirm_grouping_as_of()is the import's one call to it. "Already confirmed" means the cachedgrouping_verified_at/bymatch and, where the schema hasValidationReview, a standing grouping review by that person at that time exists; a cache without its review is re-confirmed throughverify_groupingso the row gets written.Identification.save()replays the withdraw behaviour;created_atrestored afterwards; "agreed with" links resolved by prediction taxon and by bundle ref. Missing users or taxa are reported and skipped, never created.unchanged; identifications that exist with the same user, taxon and time are skipped.export_embeddings --project N --algorithm KEY --output dir/writes onevectors.<source>.npyper store (DetectionEmbeddingrows and the classifier features onClassification.features_2048, which differ in length), plusindex.csvandmanifest.json;import_embeddingsmatches the keys and writes embeddings asDetectionEmbeddingrows and classifier features back onto the matching classification, reporting the ones whose detection has no such classification.ami/main/test_validated_occurrences.py: round trip, overlap fallback and a stricter threshold, partial not confirmed, hand-drawn box recreated, idempotent re-run, dry run writes nothing, missing user or taxon skipped, cache-only confirmation re-confirmed, vector round trips for both stores. Two tests skip on this base and run on the model-outputs branch (review row written; embedding rows written).Detailed description
What the bundle is not. It does not carry occurrences, captures or detections themselves, only what people decided about them. The target project must already hold the captures and have been processed by a detector; the bundle finds the boxes again. An unconfirmed occurrence that was identified keeps whatever grouping the target has; its identifications land on the occurrence holding most of its detections and the report says over how many occurrences they were spread.
Where this sits in the outputs stack. The model-outputs work lands as a stack:
main→ results and reviews → detection embeddings → the tracking backend (#1272) → the tracking UI (#1432). This PR builds on #1432 and moves onto its rebased head when that is pushed. Because a branch may have any subset of those tables and columns, the code here finds out what the schema has instead of assuming it: the vector export readsDetectionEmbeddingrows where that table exists and classifier feature vectors where classifications have afeatures_2048column, and notes in its manifest any store it could not read; the import writes embedding rows through the model's own insert-mostlystore()writer where it has one (a row holding the same vector is left alone, a different one replaced), setskeyandprojectwhen the model has those fields, leavesjobnull (a replayed vector has no run of its own on the target), and counts rows it has nowhere to put asskipped_no_fieldrather than failing. The "already confirmed" check looks for a grouping review only when the review model exists and has thegroupingaspect, which arrives with the tracking backend, and otherwise uses the cached fields alone.Checked against the stack's branches. On this base: 32 tests pass (2 skip, for tables it lacks). The vector modules alone, copied onto the embeddings branch (
DetectionEmbeddingin its final shape, nofeatures_2048): the embedding round trip passes, writing throughstore()withkeyandprojectset, and the manifest notes the missing feature column.makemigrations --checkreports no changes on both.Measured on a copy of a partner database (same data on both sides, so exact matches are expected): a read-only dry run of a bundle with 183 occurrences, 49 of them confirmed, matched 1,626 of 1,626 detections exactly (all 1,261 in the confirmed occurrences) and reported 182 records as already in place. The one record reported as changed was an identification made after the copy was taken, which is the right answer. The overlap fallback has not yet been exercised against a real detector rerun; the 0.7 default is the plan's starting point, to be tuned from the IoU distribution in the report.
How to test
On a stack with data:
Refs #1412, #1208. Based on #1432; moves onto its rebased head once that is pushed.
🤖 Generated with Claude Code
https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8