Skip to content

Build a classifier training set from verified occurrences, and serve it - #1502

Open
mohamedelabbas1996 wants to merge 10 commits into
feat/retrain-job-plumbingfrom
feat/retrain-training-set
Open

mohamedelabbas1996 wants to merge 10 commits into
feat/retrain-job-plumbingfrom
feat/retrain-training-set

Conversation

@mohamedelabbas1996

Copy link
Copy Markdown
Contributor

Turns the species people have verified into a training set, and serves it. Nothing runs a training job yet: that is the PR above this one.

Split out of #1494.

What a training set is

Verified crops and the embeddings already stored for them. The backbone is frozen, so a head is trained on vectors rather than images and nothing here opens a file.

Decisions worth a look:

The train/test split is grouped by occurrence, and hashed rather than drawn at random. Crops from one occurrence are near-identical frames of the same insect, so splitting per detection puts the same animal on both sides and the held-out score flatters the head. Hashing the occurrence id keeps the held-out set the same between runs, so two heads can be compared.

The vector width is read from the stored vectors, not assumed: the embedding column is unsized and two algorithms may differ.

The class list comes from the project's taxa list where it has one, so a head covers the region rather than only what has been verified. Classes with no verified data are recorded in the file instead of silently dropped. Project.default_taxa_list is new for this.

Species with too few crops are dropped - one example cannot be both trained on and evaluated.

The metadata is a schema, not a dict, so the shape is checked once here rather than at every reader. It is written inside the file, which is what lets a run say later what it learned from.

Serving it

GET /api/v2/ml/training-data/ pages the rows; GET .../summary/ answers the questions a form has to ask before a run: how many crops, how many species survive the minimum, how the split falls, how many verified crops have no embedding yet.

Paging is set for the payload: a 1024-dimension vector is roughly 20 KB of JSON, so the platform default of 10 is useless and no limit at all returns hundreds of megabytes.

Both are gated on a new run_train_classifier_job project permission. Project visibility is not enough: a non-draft project is readable by anyone, and this returns every verified label together with the vector for each crop.

Stack

#1462 embeddings  ->  #1501 plumbing  ->  this  ->  training job  ->  job form

feat/occurrence-sets (#1492) is merged in, since a training set can be narrowed to one. Land that first.

A project needs the same occurrences again later: comparing two classifiers only
means something if both saw the same rows, and a review pass wants the list it
started from. A filter answers differently as data arrives, so the membership is
stored rather than described.

An OccurrenceSet holds its occurrences and the projects it belongs to. A set with
no project is global and is offered everywhere, following how TaxaList already
treats a list with no project. Membership is decided when the set is created and
nothing adds to or removes from one afterwards, because anything recorded against
a set was measured on exactly those occurrences.

Creating, renaming and deleting are gated on new project permissions, held by the
roles that already curate a project's data. A global set has no single project to
check against, so it cannot be edited through the API at all.

The occurrence list takes an occurrence_set filter, and the sets are offered as
choices the same way capture sets are.
Adds the occurrence set to the occurrence filter panel, picked from the set
choices endpoint the same way a capture set is. A set is only useful if you can
look at what is in it, and this is where someone reviewing one starts.

The field is carried over from other views like the existing filters, so arriving
with a set already chosen shows it in the panel where it can be cleared.
Someone filtering the occurrence list to the rows they care about had no way to
keep that selection. Selecting occurrences now offers saving them as a set, next
to the identification actions already there.

The action does not change any occurrence, so it is not behind update rights on
them; the endpoint gates it on the project's own permission instead.
Registering a filter in the shared list is not enough for it to appear: the
occurrences page renders one FilterControl per field it offers, and the set was
missing from that list, so the filter existed everywhere except on screen.

It sits under More filters beside the capture set, and that section now opens on
arrival when a set is already applied, as it does for the other filters there.
The selection bar holds identification actions and is hidden from anyone without
update rights on the occurrences, so saving a set — which changes none of them —
was unavailable to a reader who could still create one. It also sat there as an
unlabelled icon among three others.

It now sits beside Export as a labelled button, and appears only while something
is selected, since that is the only time it does anything.
Training is run per project and rewrites what every later identification is
compared against, so it belongs behind its own permission rather than any
existing run_*_job one. ML data managers get it with the other job permissions
they already hold; the migration grants it to the groups that exist today so
current managers do not lose the ability when the job appears.
A head trained only on the species someone happened to verify cannot predict the
rest, so a region's expected species have to come from somewhere other than the
verified data. The project now names that list, and training reads it as the
class list.

It is optional: without one, the classes are whatever has been verified.
A head is only worth retraining on the crops a person has already named, so this
collects verified detections together with the embeddings stored for them and
writes one dataset file per run.

Decisions worth knowing about:

The train/test split is grouped by occurrence, not by detection. Crops from one
occurrence are near-identical frames of the same insect, so splitting per
detection puts the same animal on both sides and the held-out score flatters the
head. It is also a hash of the occurrence id rather than a random draw, so the
held-out set stays the same between runs and two heads can be compared.

The vector width is read from the stored vectors rather than assumed, because the
embedding column is unsized and two algorithms may differ.

Species with too few verified crops are dropped, since one example cannot be both
trained and evaluated on. The classes come from the project's taxa list where it
has one, so the head covers the region and not only what has been verified, and
the classes with no verified data are recorded in the file rather than silently
left out.

The metadata is declared as a schema instead of assembled as a dict, so the shape
is checked once here rather than at every reader, and the processing service
echoes it back for the job to store.

The rows are also served over the API, paged at a limit set for their size: a
1024-dimension vector is roughly 20 KB of JSON, so the platform default of 10 is
useless and no limit at all returns hundreds of megabytes.
The training job form asks for a split ratio and an optional occurrence set. Both
questions are unanswerable without knowing what is there, so the summary endpoint
now answers for exactly the run being set up.

It takes the occurrence set, or reports every verified occurrence when none is
chosen, which is what a run without a set does. It defaults the settings to the
algorithm's own training config rather than to module defaults, so the first
numbers someone sees are the numbers a run would use, and it accepts the settings
as parameters so the form can preview a changed ratio without starting a job. The
same bounds a job is held to are applied here, since stats computed under settings
no run could use are worse than no stats.

It also reports the classes that survive min_per_species, and names the species
dropped: a raw class count overstates what the head would come out knowing.

While adding the set parameter: the endpoint was gated on project visibility,
which a non-draft project grants to everyone, so any account could read every
verified label in a project together with the vector for each crop. It is now
gated on the same permission as the training job.
@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 6ee96b54-931a-4fba-a7f7-d8bf854756b8

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@netlify

netlify Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for antenna-preview ready!

Name Link
🔨 Latest commit ae5da96
🔍 Latest deploy log https://app.netlify.com/projects/antenna-preview/deploys/6ac875612a45910008b50f6d
😎 Deploy Preview https://deploy-preview-1502--antenna-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
Lighthouse
Lighthouse
1 paths audited
Performance: 56 (🔴 down 9 from production)
Accessibility: 81 (🔴 down 8 from production)
Best Practices: 92 (🔴 down 8 from production)
SEO: 92 (no change from production)
PWA: 80 (no change from production)
View the detailed breakdown and full score reports
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant