Repository navigation
Build a classifier training set from verified occurrences, and serve it - #1502
Open
mohamedelabbas1996 wants to merge 10 commits into
Open
mohamedelabbas1996 wants to merge 10 commits into
mohamedelabbas1996 wants to merge 10 commits into
Conversation
A project needs the same occurrences again later: comparing two classifiers only means something if both saw the same rows, and a review pass wants the list it started from. A filter answers differently as data arrives, so the membership is stored rather than described. An OccurrenceSet holds its occurrences and the projects it belongs to. A set with no project is global and is offered everywhere, following how TaxaList already treats a list with no project. Membership is decided when the set is created and nothing adds to or removes from one afterwards, because anything recorded against a set was measured on exactly those occurrences. Creating, renaming and deleting are gated on new project permissions, held by the roles that already curate a project's data. A global set has no single project to check against, so it cannot be edited through the API at all. The occurrence list takes an occurrence_set filter, and the sets are offered as choices the same way capture sets are.
Adds the occurrence set to the occurrence filter panel, picked from the set choices endpoint the same way a capture set is. A set is only useful if you can look at what is in it, and this is where someone reviewing one starts. The field is carried over from other views like the existing filters, so arriving with a set already chosen shows it in the panel where it can be cleared.
Someone filtering the occurrence list to the rows they care about had no way to keep that selection. Selecting occurrences now offers saving them as a set, next to the identification actions already there. The action does not change any occurrence, so it is not behind update rights on them; the endpoint gates it on the project's own permission instead.
Registering a filter in the shared list is not enough for it to appear: the occurrences page renders one FilterControl per field it offers, and the set was missing from that list, so the filter existed everywhere except on screen. It sits under More filters beside the capture set, and that section now opens on arrival when a set is already applied, as it does for the other filters there.
The selection bar holds identification actions and is hidden from anyone without update rights on the occurrences, so saving a set — which changes none of them — was unavailable to a reader who could still create one. It also sat there as an unlabelled icon among three others. It now sits beside Export as a labelled button, and appears only while something is selected, since that is the only time it does anything.
…t/retrain-training-set
Training is run per project and rewrites what every later identification is compared against, so it belongs behind its own permission rather than any existing run_*_job one. ML data managers get it with the other job permissions they already hold; the migration grants it to the groups that exist today so current managers do not lose the ability when the job appears.
A head trained only on the species someone happened to verify cannot predict the rest, so a region's expected species have to come from somewhere other than the verified data. The project now names that list, and training reads it as the class list. It is optional: without one, the classes are whatever has been verified.
A head is only worth retraining on the crops a person has already named, so this collects verified detections together with the embeddings stored for them and writes one dataset file per run. Decisions worth knowing about: The train/test split is grouped by occurrence, not by detection. Crops from one occurrence are near-identical frames of the same insect, so splitting per detection puts the same animal on both sides and the held-out score flatters the head. It is also a hash of the occurrence id rather than a random draw, so the held-out set stays the same between runs and two heads can be compared. The vector width is read from the stored vectors rather than assumed, because the embedding column is unsized and two algorithms may differ. Species with too few verified crops are dropped, since one example cannot be both trained and evaluated on. The classes come from the project's taxa list where it has one, so the head covers the region and not only what has been verified, and the classes with no verified data are recorded in the file rather than silently left out. The metadata is declared as a schema instead of assembled as a dict, so the shape is checked once here rather than at every reader, and the processing service echoes it back for the job to store. The rows are also served over the API, paged at a limit set for their size: a 1024-dimension vector is roughly 20 KB of JSON, so the platform default of 10 is useless and no limit at all returns hundreds of megabytes.
The training job form asks for a split ratio and an optional occurrence set. Both questions are unanswerable without knowing what is there, so the summary endpoint now answers for exactly the run being set up. It takes the occurrence set, or reports every verified occurrence when none is chosen, which is what a run without a set does. It defaults the settings to the algorithm's own training config rather than to module defaults, so the first numbers someone sees are the numbers a run would use, and it accepts the settings as parameters so the form can preview a changed ratio without starting a job. The same bounds a job is held to are applied here, since stats computed under settings no run could use are worse than no stats. It also reports the classes that survive min_per_species, and names the species dropped: a raw class count overstates what the head would come out knowing. While adding the set parameter: the endpoint was gated on project visibility, which a non-draft project grants to everyone, so any account could read every verified label in a project together with the vector for each crop. It is now gated on the same permission as the training job.
Contributor
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for antenna-preview ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
mohamedelabbas1996
added this pull request to stack #1505
October 9, 2026 05:05
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Turns the species people have verified into a training set, and serves it. Nothing runs a training job yet: that is the PR above this one.
Split out of #1494.
What a training set is
Verified crops and the embeddings already stored for them. The backbone is frozen, so a head is trained on vectors rather than images and nothing here opens a file.
Decisions worth a look:
The train/test split is grouped by occurrence, and hashed rather than drawn at random. Crops from one occurrence are near-identical frames of the same insect, so splitting per detection puts the same animal on both sides and the held-out score flatters the head. Hashing the occurrence id keeps the held-out set the same between runs, so two heads can be compared.
The vector width is read from the stored vectors, not assumed: the embedding column is unsized and two algorithms may differ.
The class list comes from the project's taxa list where it has one, so a head covers the region rather than only what has been verified. Classes with no verified data are recorded in the file instead of silently dropped.
Project.default_taxa_listis new for this.Species with too few crops are dropped - one example cannot be both trained on and evaluated.
The metadata is a schema, not a dict, so the shape is checked once here rather than at every reader. It is written inside the file, which is what lets a run say later what it learned from.
Serving it
GET /api/v2/ml/training-data/pages the rows;GET .../summary/answers the questions a form has to ask before a run: how many crops, how many species survive the minimum, how the split falls, how many verified crops have no embedding yet.Paging is set for the payload: a 1024-dimension vector is roughly 20 KB of JSON, so the platform default of 10 is useless and no limit at all returns hundreds of megabytes.
Both are gated on a new
run_train_classifier_jobproject permission. Project visibility is not enough: a non-draft project is readable by anyone, and this returns every verified label together with the vector for each crop.Stack
feat/occurrence-sets(#1492) is merged in, since a training set can be narrowed to one. Land that first.