Skip to content

Let a job type declare what it needs, and an algorithm say whether it can be retrained - #1501

Open
mohamedelabbas1996 wants to merge 24 commits into
feat/detection-embeddings-taskfrom
feat/retrain-job-plumbing
Open

mohamedelabbas1996 wants to merge 24 commits into
feat/detection-embeddings-taskfrom
feat/retrain-job-plumbing

Conversation

@mohamedelabbas1996

Copy link
Copy Markdown
Contributor

Groundwork for retraining a classifier head, with nothing about training in it. Split out of #1494 so it can be read on its own.

A job type can describe itself

Four declarations on JobType, and the two checks that read them:

  • required_fields and required_params are what a job of this type cannot run without. The API refuses to create one that is missing any, so a gap is a 400 when the job is made rather than a failure minutes later in a worker.
  • user_creatable says whether a person starts one by hand. It is not enforced on create: several types the platform makes as a side effect are also posted directly by existing clients. It says which types a job form should offer, which is the UI's question.
  • stalled_after_minutes is how long one may go untouched before the stale check calls it dead.

params also becomes writable when a job is created, so a type's own settings can reach it. One field rather than a column per type, because only the type knows what its params mean.

The stale sweep judges a job by its own type

The sweep failed anything untouched for ten minutes. A job that hands work to an external service and waits to be called back writes nothing to its own row meanwhile, so for those types silence is not evidence that the job died. Candidates are still gathered at the default, then each is judged against its own type's deadline. The type is looked up rather than read off the job, so an unrecognised job_type_key leaves that one job alone instead of stopping the whole sweep.

An algorithm says whether it can be retrained

trainable, training_config and training_info on Algorithm, mirrored from the service's /info at registration the way task_type and uri already are.

This was a real gap: a service could report trainable: true and Antenna stored false, because nothing read the field. Found by registering a service that does report it.

The settings are seeded only when the algorithm is new, so an admin's edits are not overwritten the next time pipelines are re-registered. Algorithms can also be filtered by trainable, which is what a form offering "retrain this one" needs.

Stack

#1462 detection embeddings  ->  this  ->  training set  ->  training job  ->  job form

Based on feat/detection-embeddings-task because the algorithm migration follows that branch's 0030.

mihow and others added 24 commits October 5, 2026 22:46
This is the first install of pgvector in Antenna: nothing on main uses it yet, and no
database has it today. The migration first checks that the server offers pgvector 0.8 or
later, reading default_version from pg_available_extensions, and stops with one clear
message when the package is missing or older. Only then does it create the extension.
The reverse leaves the extension in place, because it may be shared with other databases
on the server and dropping it can be restricted in hosted environments.

The local and CI Postgres image installs the 0.8 series of the postgresql-16-pgvector
package, so every environment starts on the same release. The Python client library
(pgvector 0.5.0) provides the Django field and distance expressions the next commits use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
…ch detection

A processing service can now return a feature vector (embedding) with each detection, as
an item in the detection's new embeddings list that names the algorithm whose backbone
produced it. Antenna stores it in a new DetectionEmbedding table: one row per detection,
algorithm and key, with the project copied from the detection's capture (or its station),
the job whose results stored it, and the vector itself in an unsized half-precision pgvector
column kept uncompressed out of line. Vectors land on their detection by matching the
returned box, not by position, because detection creation returns existing detections ahead
of new ones. Storing a vector never adds a classification, so no determination can change.

Writes are insert-mostly: an identical vector is left alone, a different one replaces the
row in place, and a value half precision cannot hold is skipped with a warning. Each
algorithm records the length of its first stored vector and refuses any other length, since
vectors of different lengths could never be compared. The one reader,
vectors_for_detections(), returns the vectors of one algorithm and key at a time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
…ne occurrence

GET /occurrences/?ordering=visual_similarity sorts a project's occurrences by the cosine
distance between their feature vectors and a seed occurrence's, most similar first;
-visual_similarity reverses it. The seed is similar_to=<occurrence id>, or by default the
most recently updated occurrence with a vector that the default filters show. One
algorithm's vectors are compared at a time (similarity_algorithm=<id>, or the one with the
most vectors in the project), because distances between algorithms are meaningless.
Occurrences without a vector come last in either direction, and the sort goes through the
existing default filters and visibility rules. Bad parameters return 400.

An occurrence's vector is its representative detection's: the earliest detection that has
one, which is the crop the list shows when that crop has a vector. The distance is computed
inside the correlated subquery so the list's aggregate annotations evaluate it once per
occurrence rather than once per joined detection. This is an exact scan; measured at
20,000 occurrences and 40,000 2048-d vectors it takes about 0.45 s with Postgres JIT off
and about 1.1 s with the default JIT settings, and the pagination count is unaffected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
…ccurrences

The snapshots column of the occurrence table can now be sorted, which orders the list by
visual similarity, and an occurrence's details page gets a "Show similar occurrences" link
that opens the list sorted by similarity to that occurrence. The seed is shown as a
read-only "Similar to occurrence" filter so it is visible and clearable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
The feature-vector table now lives in ami/ml/models/embedding.py next to the other
model-output tables, instead of in ami/main/models.py. The field definitions, the
constraints and the out-of-line (STORAGE EXTERNAL) vector column are unchanged; only the
app label of the table and constraint names changes.

No feature vectors exist in any deployment yet, so the migrations are regenerated rather
than chained onto the earlier ones: ml/0029 enables pgvector (with the version check) and
ml/0030 creates the table and Algorithm.embedding_dimensions. The branch's own
main/0096 and main/0097 are removed, because those numbers now belong to the job columns
added to detections and classifications.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The code that stores vectors (create_detection_embeddings, the box matching and the
dimension check) moves from ami/ml/models/pipeline.py to ami/ml/embeddings/writer.py, and
the readers move from ami/main/models_future/embeddings.py to ami/ml/embeddings/reader.py.
Behaviour is unchanged; imports in the API views, the occurrence queryset, the tests and
the canonical-patterns reference are updated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
A model can return several outputs under different keys, and those need not share a
length. The writer now holds each (algorithm, key) to the length of one existing row
of that pair, so Algorithm.embedding_dimensions is no longer needed and is removed
from the model, the serializer and the unpublished ml/0030 migration.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Reading all of one model's vectors in a project, or counting vectors per model, now
has an index on (project, algorithm, key, detection). The trailing detection column
lets a project's vectors be read in detection order without a sort.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Each helper serves one known query and names the index it relies on: one model's vectors
for a set of detections, a project's vectors in bounded id-ordered chunks, vector counts
per (algorithm, key), and the detections still missing a vector. The default-algorithm
lookup now reuses the counts helper.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The writer checks a batch against the length of one stored vector of the same
(algorithm, key). Neither existing index leads with those columns, so the lookup
sequentially scanned the table: until the first matching row for a model that
has rows (the cost grows with where that model's rows sit in the table), and
across the whole table for a model with none yet.

An index on (algorithm, key, detection) plus ordering the lookup by detection
lets the planner read the first entry of the pair directly in both cases.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…y cover

The algorithm and project foreign keys on DetectionEmbedding no longer get their own
index: the (algorithm, key, detection) and (project, algorithm, key, detection) indexes
lead with those columns, so deletes that cascade from an algorithm or a project still use
an index. Measured on a seeded copy of the largest project's layout, the planner picked
the bare algorithm index for a 5,000-detection read and filtered on detection afterwards
(26 ms); the extra indexes also cost every insert.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
detections_missing_vectors used filter(~Exists(...)). django-cachalot 2.6 does not
record the tables inside a negated Exists, so the cached result kept listing detections
after their vectors were stored, and a second feature-vector run would send them again.
This was reproduced against a regular database with the cache enabled, not only inside a
test transaction. exclude(Exists(...)) emits the same NOT EXISTS and is tracked
correctly; a test pins that the query depends on the vector table.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
A job that only stores feature vectors creates no detections or classifications, so the
"View occurrences" link for that job showed an empty list. The job filter now also matches
occurrences that have a feature vector stored by the job.

Vector lookups by job are served by a new partial index on (job, detection). The job foreign
key no longer gets its own single-column index, and the unreleased 0030 migration is edited
in place to match.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The six test classes added for feature vectors rebuilt their project,
captures and occurrences in setUp for every test, and the fixture
registered a processing service over HTTP each time. They now build
the data once in setUpTestData, and a new fixture helper skips the
processing-service calls these tests never use.

Measured on the 39 tests in test_detection_embeddings.py and
test_visual_similarity.py: 38.5 s before, 13.4 s after. The test
count and results are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…ise first writers

Vectors are now matched to stored detections by the exact bounding box coordinates, the same identity that detection reuse in get_or_create_detection relies on, instead of coordinates rounded to three decimals. When two stored detections on one capture share the same box, the vector for that box is skipped and a warning names the capture and box, rather than guessing which detection it belongs to.

The first vectors of an (algorithm, key) pair are now written under a transaction-scoped Postgres advisory lock, with the stored length re-read under the lock, so two workers cannot concurrently store different lengths. Pairs that already have rows take no lock and run the same queries as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…er the sort with a seed

Choosing any sort other than visual similarity now removes the similar_to parameter, so the filter chip and URL no longer claim a seed that the backend ignores. The Snapshots column is sortable only while a similar_to filter is active, because requesting the similarity ordering without a seed returns a 400 in projects without vectors.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Retraining needs two things an algorithm did not carry: whether the processing
service can retrain it at all, and the settings to retrain it with. Asking the
service on every page load would make the UI depend on the service being up, so
both are mirrored onto the algorithm when it registers.

The config is seeded from the service and then owned here, so an admin's edits
survive the next /info read. The settings are split on purpose: Antenna reads the
dataset half when it builds the training set and passes the fitting half through,
since the service owns the fitting. for_run() takes a job's params on top of the
config and checks them, because a test_fraction outside (0, 1) either holds out
everything or scores a head on the rows it just learned, and both finish looking
like a successful run.

training_info is the other direction: every retrain makes a new algorithm
version, so each version records where its weights came from.
The trainable flag and the training settings were on the algorithm but nothing
ever filled them: a service could say it retrains a head and Antenna stored false
anyway, so the job form offered nothing and the job refused every algorithm. Found
by registering a real service that does report both.

They now come from the /info response at registration, like task_type and uri. The
settings are seeded only when the algorithm is new, so an admin's edits are not
overwritten the next time the pipelines are re-registered.

AlgorithmTrainingConfig moves up the module because the config response now refers
to it.
The training job form offers a classifier to retrain, and listing every algorithm
in the project would offer detectors and heads no service can train. The flag is
already mirrored from the service, so the list endpoint now filters on it.
…go quiet

Four things a job type knows about itself that nothing could ask it:

``required_fields`` and ``required_params`` are what a job of this type cannot
run without, so a gap becomes a 400 when the job is created rather than a failure
minutes later in a worker. ``user_creatable`` says whether a person starts one by
hand, since the rest are made by the platform as a side effect of something else
and nothing should offer those as a choice. ``stalled_after_minutes`` is how long
one may go untouched before the stale check calls it dead, which the default of
ten minutes gets wrong for any type that waits on something external.

Declared here on their own; the checks that read them follow.
A job form had no way to send a type's own settings: params were readable on the
model and absent from the serializer, so which algorithm to retrain and what split
to hold out could not reach the API at all. They are now one writable field rather
than a column per type, because only the type knows what its params mean.

Creation also refuses a job its type could not run, naming the setting that is
missing. The alternative is a worker discovering the gap minutes later, which
leaves a failed job in the list and a person guessing what they got wrong.

Not enforced: user_creatable. Several types the platform creates as a side effect
are also posted directly by existing clients, so it says which types a job form
should offer, which is the UI's question rather than the API's.
The stale sweep fails any job untouched for ten minutes. A job that hands work to
an external service and waits to be called back writes nothing to its own row
meanwhile, so for those types silence is not evidence that the job died, and a
training run that takes an hour was being marked lost while it was still running.

Candidates are still gathered at the default, then each is judged against the
deadline its own type declares. The type is looked up rather than read off the
job, so an unrecognised job_type_key leaves that one job alone instead of raising
and stopping the whole sweep.
@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 607c10d2-a2ba-4021-8cdd-3ca6386d09fc

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@netlify

netlify Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for antenna-preview ready!

Name Link
🔨 Latest commit d6a3cc3
🔍 Latest deploy log https://app.netlify.com/projects/antenna-preview/deploys/6ac87548ac2e450008e31b3f
😎 Deploy Preview https://deploy-preview-1501--antenna-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
Lighthouse
Lighthouse
1 paths audited
Performance: 56 (🔴 down 9 from production)
Accessibility: 81 (🔴 down 8 from production)
Best Practices: 92 (🔴 down 8 from production)
SEO: 92 (no change from production)
PWA: 80 (no change from production)
View the detailed breakdown and full score reports
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants