Skip to content

Fill in the job for detections and classifications written before job tracking #1473

Description

@mihow

Summary

#1471 records which job created each detection and classification, but only for rows written after it is deployed. Every existing row has no job, so the new occurrence job filter finds nothing for past runs. This ticket is about whether, and how, to fill in the job for some existing rows. It is parked to return to later.

What does not work

The async job state cannot be used. Redis keeps only the sets of pending and failed image ids for a running job, with an expiry, and empties them as images finish. NATS streams are torn down when a job ends. Neither holds anything for a completed job.

The candidate approach: infer from job windows

A row could be attributed to a job when it was created in the same project, during the job's run (started_at to finished_at), by an algorithm in the job's pipeline, on a capture inside the job's scope (its capture set, or its single capture). Post-processing jobs would match their own algorithm in the same way.

Rough measurement on a copy of production data (about 640,000 detections, about 500 started pipeline jobs), read-only, not validated against ground truth:

Matching rule for a detection No candidate job Exactly one Two or more
Project, pipeline has the detector, created during the run ~9% ~14% ~77%
...and the capture is in the job's scope ~44% ~27% ~28%

Reasons the windows are unreliable: about 40 jobs have no finished_at, so their window has to be guessed (falling back to updated_at produced a window of more than two years for one job); overlapping runs on the same pipeline and capture set are common; and captures removed from a capture set since the run no longer match its scope.

Open questions

  1. Should inferred values share the job column with recorded ones? If they do, the filter would claim provenance that was only guessed. Options: a separate inferred_job column, a flag, or leaving old rows empty.
  2. Is ~27% of detections worth it, given that only rows with exactly one candidate job would be filled?
  3. How accurate is the rule? Once Record which job created each detection and classification, and filter occurrences by job #1471 has been deployed for a while, run the inference over rows whose job is recorded and compare. That gives a real accuracy figure before anything is written.

What we still need to verify

  • Accuracy of the matching rule against rows with a recorded job (question 3).
  • The same measurement for classifications, where class masking and the size filter add post-processing jobs to the candidates.
  • Behaviour on production, which has more data than the local copy.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions