You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Calibrate tracking's default settings from experiments on real nights #1468
Once tracking can run as a post-processing method with editable settings (the run-only tracking PR), its default settings need to be chosen from evidence rather than guesses. This ticket covers a dedicated PR that brings together the tracking experiments done so far, runs the new ones that are still needed, and sets the defaults and any named presets. The goal is that a project can run tracking with the defaults and get occurrences a reviewer can trust, without tuning anything by hand.
What earlier experiments found
These came from a database-free simulator that reproduces real tracking runs exactly, on three real nights (quiet, medium and busy), at 20-second capture intervals. Each change was also checked by eye on samples, and against occurrences that reviewers had confirmed.
The baseline links only overlapping boxes. That works at 20-second captures, but it means the settings are tuned to that interval without saying so. At longer intervals, one insect leaving and another landing in the same spot get joined.
A setting called "D3" was approved as the first default. D3 is the baseline plus three things: an appearance-similarity gate on image embeddings, an activity (crowding) adjustment, and a species penalty. On the three nights it reduced occurrences to review by about 93%, 46% and 23%, with no wrong merges in the by-eye audit.
A later combination improved on D3. It added a move rule (link an insect that moved, if its appearance is similar enough and its box is not tiny), threaded bridging across one missed capture, skipping the species check when boxes overlap strongly, and treating the same genus as no conflict. That combination raised the review saving to about 94%, 49% and 26%. Link recall against confirmed occurrences rose, precision stayed at 0.98 or higher, and there were still no wrong merges.
The best setting without embeddings, for projects or runs that have no vectors yet: the baseline plus the activity adjustment, the species rules and the interval limits.
Longer capture intervals (1 minute, 5 minutes): which limits must scale with the interval, and whether the defaults should be set per capture interval.
The guards proposed but not yet simulated, such as refusing to thread a middle detection far larger than its neighbours.
A check on nights from other stations and regions, so the defaults are not tuned to one site.
Deliverable
A PR that:
sets the defaults in the tracking settings schema;
adds named presets where a single default does not fit (for example "with embeddings" and "without embeddings", or per capture interval);
includes a short report with each setting's numbers and how to repeat them.
The experiment code can live in the repository as a management command or script if it is useful to rerun. Otherwise it is summarised in the report.
Related
#1412 (tracking v1), #1272 (source of the tracking code), #1442 (cost terms prototyped earlier), #1462 (vector storage, needed for the appearance terms), #1465 (per-model confidence threshold).
Summary
Once tracking can run as a post-processing method with editable settings (the run-only tracking PR), its default settings need to be chosen from evidence rather than guesses. This ticket covers a dedicated PR that brings together the tracking experiments done so far, runs the new ones that are still needed, and sets the defaults and any named presets. The goal is that a project can run tracking with the defaults and get occurrences a reviewer can trust, without tuning anything by hand.
What earlier experiments found
These came from a database-free simulator that reproduces real tracking runs exactly, on three real nights (quiet, medium and busy), at 20-second capture intervals. Each change was also checked by eye on samples, and against occurrences that reviewers had confirmed.
Experiments still needed
Deliverable
A PR that:
The experiment code can live in the repository as a management command or script if it is useful to rerun. Otherwise it is summarised in the report.
Related
#1412 (tracking v1), #1272 (source of the tracking code), #1442 (cost terms prototyped earlier), #1462 (vector storage, needed for the appearance terms), #1465 (per-model confidence threshold).