Repository navigation
Let a job type declare what it needs, and an algorithm say whether it can be retrained - #1501
Open
mohamedelabbas1996 wants to merge 24 commits into
Open
mohamedelabbas1996 wants to merge 24 commits into
mohamedelabbas1996 wants to merge 24 commits into
Conversation
This is the first install of pgvector in Antenna: nothing on main uses it yet, and no database has it today. The migration first checks that the server offers pgvector 0.8 or later, reading default_version from pg_available_extensions, and stops with one clear message when the package is missing or older. Only then does it create the extension. The reverse leaves the extension in place, because it may be shared with other databases on the server and dropping it can be restricted in hosted environments. The local and CI Postgres image installs the 0.8 series of the postgresql-16-pgvector package, so every environment starts on the same release. The Python client library (pgvector 0.5.0) provides the Django field and distance expressions the next commits use. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
…ch detection A processing service can now return a feature vector (embedding) with each detection, as an item in the detection's new embeddings list that names the algorithm whose backbone produced it. Antenna stores it in a new DetectionEmbedding table: one row per detection, algorithm and key, with the project copied from the detection's capture (or its station), the job whose results stored it, and the vector itself in an unsized half-precision pgvector column kept uncompressed out of line. Vectors land on their detection by matching the returned box, not by position, because detection creation returns existing detections ahead of new ones. Storing a vector never adds a classification, so no determination can change. Writes are insert-mostly: an identical vector is left alone, a different one replaces the row in place, and a value half precision cannot hold is skipped with a warning. Each algorithm records the length of its first stored vector and refuses any other length, since vectors of different lengths could never be compared. The one reader, vectors_for_detections(), returns the vectors of one algorithm and key at a time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
…ne occurrence GET /occurrences/?ordering=visual_similarity sorts a project's occurrences by the cosine distance between their feature vectors and a seed occurrence's, most similar first; -visual_similarity reverses it. The seed is similar_to=<occurrence id>, or by default the most recently updated occurrence with a vector that the default filters show. One algorithm's vectors are compared at a time (similarity_algorithm=<id>, or the one with the most vectors in the project), because distances between algorithms are meaningless. Occurrences without a vector come last in either direction, and the sort goes through the existing default filters and visibility rules. Bad parameters return 400. An occurrence's vector is its representative detection's: the earliest detection that has one, which is the crop the list shows when that crop has a vector. The distance is computed inside the correlated subquery so the list's aggregate annotations evaluate it once per occurrence rather than once per joined detection. This is an exact scan; measured at 20,000 occurrences and 40,000 2048-d vectors it takes about 0.45 s with Postgres JIT off and about 1.1 s with the default JIT settings, and the pagination count is unaffected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
…ccurrences The snapshots column of the occurrence table can now be sorted, which orders the list by visual similarity, and an occurrence's details page gets a "Show similar occurrences" link that opens the list sorted by similarity to that occurrence. The seed is shown as a read-only "Similar to occurrence" filter so it is visible and clearable. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L52AN9tabp76yjhjyCZkSJ
The feature-vector table now lives in ami/ml/models/embedding.py next to the other model-output tables, instead of in ami/main/models.py. The field definitions, the constraints and the out-of-line (STORAGE EXTERNAL) vector column are unchanged; only the app label of the table and constraint names changes. No feature vectors exist in any deployment yet, so the migrations are regenerated rather than chained onto the earlier ones: ml/0029 enables pgvector (with the version check) and ml/0030 creates the table and Algorithm.embedding_dimensions. The branch's own main/0096 and main/0097 are removed, because those numbers now belong to the job columns added to detections and classifications. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The code that stores vectors (create_detection_embeddings, the box matching and the dimension check) moves from ami/ml/models/pipeline.py to ami/ml/embeddings/writer.py, and the readers move from ami/main/models_future/embeddings.py to ami/ml/embeddings/reader.py. Behaviour is unchanged; imports in the API views, the occurrence queryset, the tests and the canonical-patterns reference are updated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
A model can return several outputs under different keys, and those need not share a length. The writer now holds each (algorithm, key) to the length of one existing row of that pair, so Algorithm.embedding_dimensions is no longer needed and is removed from the model, the serializer and the unpublished ml/0030 migration. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Reading all of one model's vectors in a project, or counting vectors per model, now has an index on (project, algorithm, key, detection). The trailing detection column lets a project's vectors be read in detection order without a sort. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Each helper serves one known query and names the index it relies on: one model's vectors for a set of detections, a project's vectors in bounded id-ordered chunks, vector counts per (algorithm, key), and the detections still missing a vector. The default-algorithm lookup now reuses the counts helper. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The writer checks a batch against the length of one stored vector of the same (algorithm, key). Neither existing index leads with those columns, so the lookup sequentially scanned the table: until the first matching row for a model that has rows (the cost grows with where that model's rows sit in the table), and across the whole table for a model with none yet. An index on (algorithm, key, detection) plus ordering the lookup by detection lets the planner read the first entry of the pair directly in both cases. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…y cover The algorithm and project foreign keys on DetectionEmbedding no longer get their own index: the (algorithm, key, detection) and (project, algorithm, key, detection) indexes lead with those columns, so deletes that cascade from an algorithm or a project still use an index. Measured on a seeded copy of the largest project's layout, the planner picked the bare algorithm index for a 5,000-detection read and filtered on detection afterwards (26 ms); the extra indexes also cost every insert. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
detections_missing_vectors used filter(~Exists(...)). django-cachalot 2.6 does not record the tables inside a negated Exists, so the cached result kept listing detections after their vectors were stored, and a second feature-vector run would send them again. This was reproduced against a regular database with the cache enabled, not only inside a test transaction. exclude(Exists(...)) emits the same NOT EXISTS and is tracked correctly; a test pins that the query depends on the vector table. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
A job that only stores feature vectors creates no detections or classifications, so the "View occurrences" link for that job showed an empty list. The job filter now also matches occurrences that have a feature vector stored by the job. Vector lookups by job are served by a new partial index on (job, detection). The job foreign key no longer gets its own single-column index, and the unreleased 0030 migration is edited in place to match. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
The six test classes added for feature vectors rebuilt their project, captures and occurrences in setUp for every test, and the fixture registered a processing service over HTTP each time. They now build the data once in setUpTestData, and a new fixture helper skips the processing-service calls these tests never use. Measured on the 39 tests in test_detection_embeddings.py and test_visual_similarity.py: 38.5 s before, 13.4 s after. The test count and results are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…ise first writers Vectors are now matched to stored detections by the exact bounding box coordinates, the same identity that detection reuse in get_or_create_detection relies on, instead of coordinates rounded to three decimals. When two stored detections on one capture share the same box, the vector for that box is skipped and a warning names the capture and box, rather than guessing which detection it belongs to. The first vectors of an (algorithm, key) pair are now written under a transaction-scoped Postgres advisory lock, with the stored length re-read under the lock, so two workers cannot concurrently store different lengths. Pairs that already have rows take no lock and run the same queries as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
…er the sort with a seed Choosing any sort other than visual similarity now removes the similar_to parameter, so the filter chip and URL no longer claim a seed that the backend ignores. The Snapshots column is sortable only while a similar_to filter is active, because requesting the similarity ordering without a seed returns a 400 in projects without vectors. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0121zMVjnPsqeDFSBXRCPvMy
Retraining needs two things an algorithm did not carry: whether the processing service can retrain it at all, and the settings to retrain it with. Asking the service on every page load would make the UI depend on the service being up, so both are mirrored onto the algorithm when it registers. The config is seeded from the service and then owned here, so an admin's edits survive the next /info read. The settings are split on purpose: Antenna reads the dataset half when it builds the training set and passes the fitting half through, since the service owns the fitting. for_run() takes a job's params on top of the config and checks them, because a test_fraction outside (0, 1) either holds out everything or scores a head on the rows it just learned, and both finish looking like a successful run. training_info is the other direction: every retrain makes a new algorithm version, so each version records where its weights came from.
The trainable flag and the training settings were on the algorithm but nothing ever filled them: a service could say it retrains a head and Antenna stored false anyway, so the job form offered nothing and the job refused every algorithm. Found by registering a real service that does report both. They now come from the /info response at registration, like task_type and uri. The settings are seeded only when the algorithm is new, so an admin's edits are not overwritten the next time the pipelines are re-registered. AlgorithmTrainingConfig moves up the module because the config response now refers to it.
The training job form offers a classifier to retrain, and listing every algorithm in the project would offer detectors and heads no service can train. The flag is already mirrored from the service, so the list endpoint now filters on it.
…go quiet Four things a job type knows about itself that nothing could ask it: ``required_fields`` and ``required_params`` are what a job of this type cannot run without, so a gap becomes a 400 when the job is created rather than a failure minutes later in a worker. ``user_creatable`` says whether a person starts one by hand, since the rest are made by the platform as a side effect of something else and nothing should offer those as a choice. ``stalled_after_minutes`` is how long one may go untouched before the stale check calls it dead, which the default of ten minutes gets wrong for any type that waits on something external. Declared here on their own; the checks that read them follow.
A job form had no way to send a type's own settings: params were readable on the model and absent from the serializer, so which algorithm to retrain and what split to hold out could not reach the API at all. They are now one writable field rather than a column per type, because only the type knows what its params mean. Creation also refuses a job its type could not run, naming the setting that is missing. The alternative is a worker discovering the gap minutes later, which leaves a failed job in the list and a person guessing what they got wrong. Not enforced: user_creatable. Several types the platform creates as a side effect are also posted directly by existing clients, so it says which types a job form should offer, which is the UI's question rather than the API's.
The stale sweep fails any job untouched for ten minutes. A job that hands work to an external service and waits to be called back writes nothing to its own row meanwhile, so for those types silence is not evidence that the job died, and a training run that takes an hour was being marked lost while it was still running. Candidates are still gathered at the default, then each is judged against the deadline its own type declares. The type is looked up rather than read off the job, so an unrecognised job_type_key leaves that one job alone instead of raising and stopping the whole sweep.
Contributor
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for antenna-preview ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Groundwork for retraining a classifier head, with nothing about training in it. Split out of #1494 so it can be read on its own.
A job type can describe itself
Four declarations on
JobType, and the two checks that read them:required_fieldsandrequired_paramsare what a job of this type cannot run without. The API refuses to create one that is missing any, so a gap is a 400 when the job is made rather than a failure minutes later in a worker.user_creatablesays whether a person starts one by hand. It is not enforced on create: several types the platform makes as a side effect are also posted directly by existing clients. It says which types a job form should offer, which is the UI's question.stalled_after_minutesis how long one may go untouched before the stale check calls it dead.paramsalso becomes writable when a job is created, so a type's own settings can reach it. One field rather than a column per type, because only the type knows what its params mean.The stale sweep judges a job by its own type
The sweep failed anything untouched for ten minutes. A job that hands work to an external service and waits to be called back writes nothing to its own row meanwhile, so for those types silence is not evidence that the job died. Candidates are still gathered at the default, then each is judged against its own type's deadline. The type is looked up rather than read off the job, so an unrecognised
job_type_keyleaves that one job alone instead of stopping the whole sweep.An algorithm says whether it can be retrained
trainable,training_configandtraining_infoonAlgorithm, mirrored from the service's/infoat registration the waytask_typeandurialready are.This was a real gap: a service could report
trainable: trueand Antenna storedfalse, because nothing read the field. Found by registering a service that does report it.The settings are seeded only when the algorithm is new, so an admin's edits are not overwritten the next time pipelines are re-registered. Algorithms can also be filtered by
trainable, which is what a form offering "retrain this one" needs.Stack
Based on
feat/detection-embeddings-taskbecause the algorithm migration follows that branch's0030.