Repository navigation
Add tracking settings that avoid merging different insects, calibrate them per image model, and tune them against confirmed tracks - #1442
Draft
mihow wants to merge 22 commits into
Conversation
…tationary-first linking Tracking links detections in consecutive captures by one summed cost. This adds settings that change which pairs link, all off by default so a run with the default configuration makes exactly the links it made before: - A weight on each cost term (appearance, overlap, size, distance). - A species gate that forbids, or adds a penalty to, a link between two detections whose confident top labels name unrelated taxa. Labels are read once per session, not per pair. - Activity scaling that makes the distance term weigh more when a pair of captures holds many detections (a log curve or a step table). - A stationary-first pass that links pairs that barely moved before any other pair, optionally even when one of them has no embedding. The pair terms can also be computed once per session and re-scored under many settings, which the evaluation command uses for parameter sweeps. Every new setting is validated by the config schema and is staff-only through the existing member allow-list. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
evaluate_tracking gains --sweep (a list of settings or a grid on a base), --config for any tracking setting, --vectors-file to compare embeddings that are not stored as classifications, --output-dir for a Markdown and JSON table, and --per-track-csv. Each setting is scored per session and overall. The scores add what precision-first tuning needs: exact recovery and completeness over confirmed tracks of two or more detections (single-detection tracks are recovered by doing nothing and are reported apart), merges across species, and for the whole session the predicted track lengths and the number of occurrences and distinct determinations before and after tracking. The pairs of each session are read and scored once and reused for every setting. The run stays inside a transaction that is always rolled back. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The regression test compares the links a default run proposes on a synthetic session with a frozen copy of the matcher as it was before the new rules. The other tests cover each rule on hand-built pairs, config validation, that the new settings stay staff-only, that precomputed pairs give the same links as a run for non-default settings, and the sweep output (tables, per-track CSV, vectors file, nothing written). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Sweep tables cut scores to three places instead of rounding, so a precision of 0.9996 reads 0.999 and only a perfect score reads 1.000. The session track length median is taken over tracks of two or more detections, because in a busy session most detections stay alone and would pin the median at one, and the table shows how many such tracks tracking made. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…oss the whole session Confirmed tracks only reveal a wrong merge when two of them are joined; a track extended into an insect nobody confirmed goes unseen. As a session-wide proxy, the evaluation now counts proposed links whose two detections carry confident labels (score 0.5 or more) naming unrelated taxa, and shows it in sweep tables. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…s [skip ci] Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for antenna-preview canceled.
|
The species gate and the evaluation's species counts read each detection's top terminal label across every classifier. When a session has more than one classifier, one crop's label could come from one model and the next crop's from another, and their scores were held to one threshold although different models do not score on the same scale. Labels now come from a single classifier per session. A new setting, species_label_algorithm_id, names it; left unset, the only classifier that labelled the session is used. A session labelled by several classifiers is skipped by a tracking run with a stated reason, and evaluate_tracking stops with an error asking for the setting. The evaluation caches labels per session and classifier, so a sweep can vary the setting. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…ip ci] Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The per-track CSV of a sweep is written before the sweep files, so a run whose output folder did not exist yet failed at the very end and lost every scored setting. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…n appearance gate and a move rule Each feature extractor has its own cosine similarity range: the classifier backbone scores nearly every pair of moths above 0.9, while a BioCLIP model spreads them from about 0.2 to 1. Three optional, staff-only settings make the appearance evidence usable with either: - appearance_similarity_floor / _ceiling map similarity linearly onto the 0..1 appearance cost (the defaults 0 and 1 keep the plain 1 - similarity). - appearance_min_similarity forbids a link whose two embeddings are less similar than the given value; pairs without an embedding are not gated. - motion_min_similarity / motion_max_shift let a pair whose embeddings look alike replace the overlap term with its centre shift in box sizes divided by motion_max_shift, so an insect that moved clear of its old box can link below a threshold of 1. Every rule is off by default, and the default cost stays bit-identical to the plain sum; tests cover each rule and the precomputed-pairs path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…le [skip ci] Adds the new settings to the tracking reference and a short guide to choosing them per feature extractor from confirmed tracks, with the similarity ranges measured on a partner's evaluation project. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
… settings The merge picker and the capture preview used the plain cost sum, so appearance calibration, the appearance gate, cost weights and the move rule had no effect on what they ranked or marked as a link. Both now score with weighted_cost under the session's link options, and would_link applies the appearance gate. The response shape is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The box-size shift is computed for every pair, move rule or not, and a box with zero or negative area made it raise. Such a box now has an infinite shift, which leaves the overlap term unchanged, so the cost is what it was before the move rule. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The merge picker and the capture preview score pairs with whatever settings tracking_config_for returns, and today that is always the default settings: no per-session tracking settings are stored yet. The docstrings and test names described the scoring as using the session's settings, which the code does not do. Reword them to name tracking_config_for and say that only the defaults exist, and note that the tests mock it. No behaviour changes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Resolved in tracking_task.py and merge_candidates.py: vectors are read with vectors_for_detections (embeddings first, then classification vectors), each tracking run records its history entry, and the link options from the tracking settings are passed to every scoring call. Same resolution as the integration branch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
mihow
added a commit
that referenced
this pull request
Sep 29, 2026
Brings in the tracking settings branch, which now also carries #1439. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
mihow
changed the base branch from
claude/revive-tracking-feature-OyMO3
to
feat/occurrence-history-and-embeddings
September 29, 2026 06:47
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
mihow
added a commit
that referenced
this pull request
Sep 29, 2026
…rom #1439) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…f proposed tracks [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…main) into feat/tracking-cost-terms Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…n's tracking settings The previews used the default settings, so a session tracked with other settings showed a different threshold and, where frames carry vectors from two models, could pick the model the tracker did not use and report no similarity. They now read the latest successful tracking run over the session and prefer its feature extractor. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
This was referenced Oct 2, 2026
Collaborator
Author
|
Claude says: The merge order and plan for tracking, agreed with the owner today, are on #1412: #1412 (comment) This PR's place: on hold. Its settings work feeds the calibration in #1468, which follows #1469. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
People reviewing moth tracks told us what matters most: a track that lumps two insects together is much worse than a missed link, tracks of moths that sit still for a long time are very valuable, and even a short correct track helps. Tracking today has one knob, the cost threshold, and its default of 0.2 leaves most real tracks in pieces.
This PR gives tracking the settings that those priorities call for, and a way to measure them before anyone changes a default. Operators (staff) can now weigh each part of the matching cost, stop links between detections that a classifier confidently calls different species, make movement tolerance stricter on a crowded sheet, and link still insects first. Every new setting is off by default, and a default run makes exactly the links it made before; a test pins that. The evaluation command can now score a whole grid of settings in one read-only run and write a table that puts wrong merges first.
Stacked on #1439 (itself on #1432 and #1272) and merges after it, because tracking now reads vectors through #1439's embedding reader and records each run in its history. No default changes here.
Tests on this head (
240722cd), run locally in the CI compose stack, since GitHub runs the backend tests only on PRs based onmain:makemigrations --checkreports no changes, and the full backend suite ran 1,095 tests, OK (2 skipped). An earlier merge of #1439 (before0966b634) hit one import error intest_tracking_cost_terms, for a helper #1439 removed; that import is fixed on this branch. Sweep results on a partner's evaluation project are below, including a general-purpose image model and settings to calibrate the appearance term per model; which defaults to change is a decision for after review.List of Changes
appearance_weight,iou_weight,size_weight,distance_weightonTrackingConfig;weighted_cost()equalstotal_cost()bit for bit at 1.0.species_gate(off/penalty/forbid),species_gate_min_score,species_gate_penalty; labels read once per session (top_labels, two queries), not per pair.applied_to).activity_scaling(logorsteps) multiplies the distance term by a function of the larger detection count of the two captures.stationary_firstpass inchoose_links()with its own threshold, max centre shift and min IoU, then the normal pass; still one link per detection per side.evaluate_tracking --sweep(list or grid),--config,--vectors-file(.npz),--output-dir(Markdown + JSON),--per-track-csv. Pair terms are computed once per session and re-scored per setting.registry.py.appearance_similarity_floor/_ceilingmap similarity linearly onto 0..1 (appearance_term()); defaults 0 and 1 keep1 - similarityexactly.appearance_min_similarity, checked inchoose_links(); pairs missing an embedding are not gated.motion_min_similarity,motion_max_shift: the overlap term becomesmin(1 - IoU, shift in box sizes / motion_max_shift)(overlap_term(),shift_in_box_sizes(), new optionalPairTerms.shift)._pair_scores()usesweighted_cost()underconfig.link_options();would_linkappliesappearance_min_similarity; the preview passes the options toselect_links()and applies the crowd multiplier. Response shape unchanged.tracking_config_for()), not the settings of the run that made the track; see "Still open" below.shift_in_box_sizes()returns infinity for a box with zero or negative area, which leaves the overlap term as it was.First sweep results (measured)
Read-only runs against a local copy of a partner's evaluation project: three one-hour windows of 180 captures (a busy night of small moths, a quiet night with a few large moths, a busy night of large moths), 49 tracks confirmed by a person (40 of two or more detections), 1,212 true links. Embeddings are a classifier backbone's (2048-d). The database session was opened read-only on top of the command's rollback.
Observations:
1 - IoUis 1 for boxes that do not overlap, so at 1.0 only overlapping boxes can link.What these numbers cannot show yet (to verify before changing defaults):
Second sweep: a general-purpose image model, calibration, gate and move rule (measured)
Same benchmark. Embeddings from BioCLIP (1024-d, every detection has one, passed with
--vectors-file) against the classifier backbone above (2048-d, 74% of detections have one). 720 and 1,440 settings.How well each model's cosine similarity tells the same insect from a different one nearby (true consecutive pairs against different insects in the adjacent capture within 3 box sizes):
Both rank pairs equally well, but the backbone squeezes everything into 0.87–1, so its
1 - similarityterm moves the cost by about 0.05 and cannot stop a merge.Overall results, crowd-aware movement on (log), move rule off. Columns: link recall / exact tracks of 40 / merges of two confirmed tracks / links between confidently different labels / occurrences after tracking (from 12,168):
Calibration used floor = median of the nearby different insects, ceiling = 25th percentile of true pairs, gate = 1st percentile of true pairs (BioCLIP 0.39 / 0.89 / 0.47; backbone 0.957 / 0.988 / 0.958). "Plain + gate" uses
1 - similaritywith the same gate. Every threshold, percentile and gate value was chosen and scored on the same three one-hour windows, so all of these numbers are in-sample; there is no held-out set yet.Observations:
Visual check of all predicted tracks (measured, by eye)
The merge counts above only see joins between two confirmed tracks, and the confirmed tracks are mostly clear, well-separated moths, so those counts under-detect joins. To measure precision on everything else, every predicted track of 4 or more detections in a sample was looked at by eye: 365 tracks in the same three one-hour windows and 99 on three full nights, 464 in all, labelled as one insect, two or more insects, or unsure. This is the precision measure to trust over the merge columns above. All three settings use the classifier backbone with no calibration and no gate; unsure tracks count as correct, and the worst case counting them as joins is given in brackets.
A direction to discuss, not a decision: threshold 1.0 without requiring embeddings is the only looser setting checked by eye, and it joins two insects in about 3% of long tracks in the hour windows and about 10% on the busiest full night, while cutting the occurrences to review on the full nights by 29%, 50% and 93%. The default joined nothing in the sample but saves little on busy nights (11% on the busiest). The best result in the second sweep, BioCLIP at 1.5 with crowd-aware movement and the gate at 0.47 ("plain + gate" above: link recall 0.965, 24 of 40 exact tracks, no merges between confirmed tracks), has not been checked by eye. Given what the check found for the ungated 1.5 setting, it should not be adopted until it has. It also cannot run on this branch yet, because the task cannot read BioCLIP embeddings until the detection-embedding work lands.
Detailed Description
choose_links()is the single matcher: the task'sselect_links()(used by runs and by the session view's link preview) and the evaluation's precomputed path both call it, and a test checks that both give identical links for non-default settings.test_default_settings_give_the_same_links_as_before_the_new_rulescompares default-config links on a synthetic session against a frozen copy of the previous matcher, at four threshold/feature combinations.event_transition_pairs()converts each embedding to an array once per detection instead of once per pair, keeping its dtype, so costs are unchanged.Still open
How to test
Tested locally:
ami.ml.post_processing,ami.jobsand the tracking test classes inami.main.testspass (422 tests; two NATS result tests failed once on a transient name lookup of the test ML service and passed on rerun);makemigrations --checkreports no changes. Frontend untouched. After the calibration, gate and move rule:ami.ml.post_processingpasses (136 tests) andmakemigrations --checkreports no changes.🤖 Generated with Claude Code
https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8