You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Fill in the job for detections and classifications written before job tracking #1473
#1471 records which job created each detection and classification, but only for rows written after it is deployed. Every existing row has no job, so the new occurrence job filter finds nothing for past runs. This ticket is about whether, and how, to fill in the job for some existing rows. It is parked to return to later.
What does not work
The async job state cannot be used. Redis keeps only the sets of pending and failed image ids for a running job, with an expiry, and empties them as images finish. NATS streams are torn down when a job ends. Neither holds anything for a completed job.
The candidate approach: infer from job windows
A row could be attributed to a job when it was created in the same project, during the job's run (started_at to finished_at), by an algorithm in the job's pipeline, on a capture inside the job's scope (its capture set, or its single capture). Post-processing jobs would match their own algorithm in the same way.
Rough measurement on a copy of production data (about 640,000 detections, about 500 started pipeline jobs), read-only, not validated against ground truth:
Matching rule for a detection
No candidate job
Exactly one
Two or more
Project, pipeline has the detector, created during the run
~9%
~14%
~77%
...and the capture is in the job's scope
~44%
~27%
~28%
Reasons the windows are unreliable: about 40 jobs have no finished_at, so their window has to be guessed (falling back to updated_at produced a window of more than two years for one job); overlapping runs on the same pipeline and capture set are common; and captures removed from a capture set since the run no longer match its scope.
Open questions
Should inferred values share the job column with recorded ones? If they do, the filter would claim provenance that was only guessed. Options: a separate inferred_job column, a flag, or leaving old rows empty.
Is ~27% of detections worth it, given that only rows with exactly one candidate job would be filled?
Summary
#1471 records which job created each detection and classification, but only for rows written after it is deployed. Every existing row has no job, so the new occurrence job filter finds nothing for past runs. This ticket is about whether, and how, to fill in the job for some existing rows. It is parked to return to later.
What does not work
The async job state cannot be used. Redis keeps only the sets of pending and failed image ids for a running job, with an expiry, and empties them as images finish. NATS streams are torn down when a job ends. Neither holds anything for a completed job.
The candidate approach: infer from job windows
A row could be attributed to a job when it was created in the same project, during the job's run (
started_attofinished_at), by an algorithm in the job's pipeline, on a capture inside the job's scope (its capture set, or its single capture). Post-processing jobs would match their own algorithm in the same way.Rough measurement on a copy of production data (about 640,000 detections, about 500 started pipeline jobs), read-only, not validated against ground truth:
Reasons the windows are unreliable: about 40 jobs have no
finished_at, so their window has to be guessed (falling back toupdated_atproduced a window of more than two years for one job); overlapping runs on the same pipeline and capture set are common; and captures removed from a capture set since the run no longer match its scope.Open questions
jobcolumn with recorded ones? If they do, the filter would claim provenance that was only guessed. Options: a separateinferred_jobcolumn, a flag, or leaving old rows empty.What we still need to verify