Skip to content

Store, show and review what post-processing methods decide #1457

Description

@mihow

Summary

Post-processing methods (class masking, size filters, tracking, and future ones such as out-of-distribution scores) change occurrences after the main pipeline has run. Antenna can already trigger and review them from the admin (#954, #1289), but it does not yet keep a durable record of what each run decided, which job produced each output, or what a person confirmed afterwards. This ticket collects that work so the tracking feature and every later post-processing method share one storage model and one review history.

The design is settled in #1431. It is being implemented in two pull requests, #1461 and then #1462, with tracking (#1272, #1432) building on them:

  1. Embeddings: a feature vector for every detection in its own table, recording the job and algorithm that produced it. Also records which job created each detection and classification (PSv2: add Job relationship to a Classification and Detection objects #1156).
  2. Results and reviews: one table for what any post-processing run decided about an occurrence, capture or session, and one for what a person confirmed, rejected, corrected or commented on. Includes the history endpoint and the identification timeline (Show when a track was reviewed and what changed, in one identification timeline #1433).

Done

In progress

Related

Activity

  1. added
    post-processingPost-processing task framework and tasks: size filter, class masking, rank rollup, tracking
    on Oct 2, 2026
  2. mihow commented on Oct 10, 2026

    @mihow
    CollaboratorAuthor

    Claude says: Plan for where algorithm outputs are stored, agreed with the owner on 2026-10-09. It follows a test of every open PR in this stack together (main, #1461, #1469, #1462 and #1447) on one local deployment.

    Where we are heading

    Every algorithm output has one home, chosen by how it is used:

    Output Home Why
    Labels and scores (classifiers) Classification Already searched, filtered and used to name occurrences
    Feature vectors Their own table per target (detection vectors in #1462) Searched by similarity, each needs its own vector index
    Figures that are shown or kept for the record: a filter's decision, the before and after of class masking, the cost of each tracking link, run summaries AlgorithmResult (#1461) Read in the history, not queried across rows

    The rule for new work: anything that is searched, filtered or indexed gets a typed table; everything else is an AlgorithmResult. Vectors should not go into a result's data field, and labels should not either.

    Steps

    1. New home for algorithm results that are not species classifications #1461 merges with results on occurrences only. It is ready, and the plan is stated in its description.
    2. Results on detections and captures: a follow-up PR from the New home for algorithm results that are not species classifications #1461 work, opened after New home for algorithm results that are not species classifications #1461 merges. It adds optional detection and capture links next to the occurrence link, with a database check that exactly one is set. Results already written keep their occurrence link, so nothing is migrated. The design goes in that PR's description.
    3. Per-detection outputs move to detections once step 2 lands:
      • Tracking (Run automated tracking from the admin as a post-processing method, with editable settings #1469) stores the cost of the links each run makes on the occurrence, as a list with one entry per detection in the result's detection order (empty when the link came from an earlier run, was refused, or the detection is the last one). Because each cost is aligned with its detection, a data migration can rewrite the results as one result per detection without rerunning tracking.
      • Class masking currently records one result per occurrence. Masking works on each detection's classifications, so its results belong on detections. Until then, results pile up when tracking merges occurrences: one merged occurrence in a test project showed 51 masking results. The interim fix, in Run automated tracking from the admin as a post-processing method, with editable settings #1469, shows one history entry per run: a run's results on a merged occurrence are grouped, with every classification they created and a count, and the figures of any single result are left off, since they may describe an absorbed occurrence. Results on detections can use the same grouping.
    4. Occurrences built by tracking (see the plan in Tracking usable & available behind feature flag #1412) depends on step 2, since masking runs before any occurrence exists.

    Vectors for captures and taxa

    Each target gets its own vector table, each with its own vector index. #1462 adds a shared abstract model base (BaseEmbedding) with detection vectors (DetectionEmbedding) as its first table, so a capture or taxon vector table is a new model plus a migration. The abstract base adds no table of its own.

    Edited on 2026-10-09: corrected the tracking migration step. The link cost list in #1469 is not aligned with the detection list today, so it needs to be keyed by detection before it merges.

    Edited on 2026-10-10: #1469 now aligns the link costs with the detections and groups a run's results on merged occurrences into one history entry; step 3 describes both.

    Edited on 2026-10-10: the vector section now matches #1462, which adds the shared base.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    post-processingPost-processing task framework and tasks: size filter, class masking, rank rollup, tracking

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions