Skip to content

Retrain a classifier head from a training set, and take the result back - #1503

Open
mohamedelabbas1996 wants to merge 8 commits into
feat/retrain-training-setfrom
feat/retrain-training-job
Open

mohamedelabbas1996 wants to merge 8 commits into
feat/retrain-training-setfrom
feat/retrain-training-job

Conversation

@mohamedelabbas1996

Copy link
Copy Markdown
Contributor

The job that retrains a head: it builds the training set, hands a processing service the URL, and records what comes back as a new algorithm version.

Split out of #1494. The service side is ami-data-companion#167.

What a run does

  1. builds the dataset (the PR below this one) and writes it to storage
  2. posts the URL to the service and leaves the job STARTED
  3. the service trains, uploads the head, reports its epochs, then reports the result
  4. the result is registered as a new algorithm version recording what it learned from

The occurrence set is optional. Without one a run learns from every verified occurrence in the project, which is the ordinary case; naming a set makes the run repeatable, and either way the dataset file records which it was.

The interface

A retraining run reports back the way an ML job's results already do: a POST to a route on the job, typed with a schema and a serializer, acknowledged with a typed response. TrainingRequest and TrainingResult mirror the service's own TrainRequest and TrainResponse, as PipelineRequest and PipelineResultsResponse already do for processing. The job is marked ASYNC_API, which is what it has always been.

One way in. /train returns 202; the result arrives only at the callback and is parsed once. Antenna previously also read a result from the /train response body, which gave one outcome two shapes - and both of the bugs this fixes lived in that gap:

  • on the callback path the echoed dataset metadata was looked for under a key it is not at, so it was never parsed and a registered version never recorded which occurrence set it came from
  • on the inline path the same code would have raised AttributeError; it never fired only because the duplicate guard returned first

The result is validated because the new version is built from it. The echoed metadata is read leniently: a service echoing an older shape should still have its result recorded, and the cost is provenance, not a lost head.

Three things here have no precedent in the platform:

  1. The callback token. Antenna hands the service a URL and a token signed for that one job, valid 24 hours. The existing async path does it the other way round: the ADC worker holds an Antenna API token and posts as itself. I kept the signed token because it is the tighter of the two, and because POST /jobs/{id}/result/ resolves to a result_ml_job permission that does not exist, so that endpoint effectively needs a superuser. Happy to switch if you would rather have one mechanism. That permission looks unintended and may deserve its own fix.
  2. The head upload. No endpoint existed for a service to send a file back.
  3. The progress ping. Processing infers progress from results arriving; one long fit has nothing to infer from, so the service reports its epoch and the stage follows it.

Progress while it trains

A training job at 52 per cent, with the training stage started

The training stage showing epoch 4,950 of 5,000

The total is written at dispatch from the run's own settings, so the stage reads "0 of 5000" while the service is still starting. A ping never finishes the stage - the result does, since the head still has to be scored and uploaded. Late or out-of-order pings are dropped. Reporting is optional: a service that posts nothing behaves as before.

A run, end to end

Against the BioCLIP 2.5 service on the GPU VM, over a project's real verified crops:

The training job finished, with its three stages green

The job's logs in full

Training set: 96 verified crops over 17 species (74 train / 22 test)
Species list comes from taxa list 'NF retraining list (demo)'
7 species in the list have no verified crops yet
Registered #67 "BioCLIP 2.5 + LogReg head (Newfoundland, 749 species)" v22 as version 22
New head top-1 0.636 does not beat current 0.875 by more than 0.000.

The last line is the point of keeping the score: a retrain that is worse than what it would replace says so rather than quietly taking over.

Stack

#1462 embeddings  ->  #1501 plumbing  ->  #1502 training set  ->  this  ->  job form

Adds the training job: it builds the dataset, hands a processing service the URL,
and records what comes back as a new algorithm version.

Training outlasts the request that starts it, so the service reports its result to
a callback instead of the connection being held open. The service has no Antenna
account, so the callbacks are authorised by a signed token issued when the job was
dispatched, and the job is looked up directly because project visibility would
refuse an unauthenticated caller. The weights are uploaded the same way: without
that, the head exists only in a cache directory on the service's disk and Antenna
records a version it cannot point at. An upload is refused once the job has
finished, since the token lasts a day and the path is fixed by the job, so
otherwise the weights behind a registered version could be swapped out after the
run ended.

A job type now also declares what it needs before it can start and whether a user
may create it at all, rather than each job discovering a missing param partway
through its own run.
The refactor that gathered the training modules into one package left two imports
behind, so running a training job raised ImportError on ami.ml.training_dataset
instead of building a dataset. The tests did not catch it because neither import
is reached until a job actually runs.
Requiring a set meant that retraining on everything verified so far - the ordinary
case - first made someone save a set of everything, which is a step with no
decision in it.

A job without a set now learns from every verified occurrence in the project. The
dataset file records which set was used, or that there was none, so a run is still
legible afterwards either way. Naming a set is still worth doing when it matters:
a set is fixed once created, so that run can be repeated.
A key is shared by every version of an algorithm, while embeddings are stored
against one version row. Taking whichever row the database returned first could
pick an older version and train on an empty set of vectors, or on vectors from a
different fitting, with nothing in the job to say which had happened.
Training is the long stage of the job and the service says nothing while it runs,
so the bar sat at nought for as long as the fitting took and a person could not
tell a slow run from a dead one.

The service is now told where to report, and posts the epoch it has reached. The
callback is authorised by the same signed token as the result and the head upload,
since a service has no Antenna account. The epoch and the total land as stage
parameters, and the stage progress is the ratio of the two.

The total is written at dispatch from the run's own settings, so the stage reads
"0 of 5000" while the service is still starting rather than an empty pair. A ping
never finishes the stage: the result does, because the service still has to score
the head and upload it after the last epoch. Pings that arrive late or out of order
are dropped rather than dragging a run backwards.

Reporting is optional. A service that does not post anything behaves exactly as
before, and a ping that fails to send must not fail the run.
The training request was a dict literal and the result was read with .get() at a
dozen call sites, so neither side of the exchange was written down anywhere. Every
other service contract in this module is a schema; this one was the exception, and
a service renaming a field would have produced an empty metrics dict and no error.

TrainingRequest and TrainingResult mirror the service's own TrainRequest and
TrainResponse, the way PipelineRequest and PipelineResultsResponse already do for
processing. TrainingResult keeps unknown fields: a newer service may report more
than this one knows about, and the result is the only record of what a run did.

incumbent_metrics is nullable because a service reports null when there was no
current head to score, which is exactly what a first retrain does.
… to read

A result could arrive two ways: in the body of the /train response, or at the
job's callback. The two carried the same outcome in different shapes, and both of
these bugs lived in that gap.

On the callback path the dataset metadata the service echoes was read as
payload["dataset"]["metadata"], which is not where it is, so it was never parsed
and a registered version never recorded which occurrence set it learned from. The
provenance this is all for was quietly dropped on every real run.

On the inline path the same code would have raised AttributeError, because there
"dataset" is the schema object rather than a dict. It never fired only because a
service posts its callback first and the duplicate guard returned early.

/train is now an acknowledgement: the result arrives at the callback, is validated
by a serializer at the view, and record_result takes the parsed result rather than
a raw payload. One shape, one path, parsed once at the boundary.

The result is validated because the new algorithm version is built from it. The
echoed dataset metadata is read leniently: a service echoing an older shape should
still have its result recorded, and the cost is a version that cannot say which
set it came from, not a head that is lost.

Also fixes a flaky test: with a handful of rows, a hash split does not give a
particular ratio a particular count, so the ratio is now compared across two
values rather than asserted exactly.
Every training job read as "internal", the mode for work the platform does
itself, because Job.setup() decides the mode from the pipeline and a training job
has none. It is the platform's own word for what this job has always done: hand
the work to a service and wait to be called back.
@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 04e29405-c18d-48b8-9978-01b2f67ca5a7

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@netlify

netlify Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for antenna-preview ready!

Name Link
🔨 Latest commit 2f55e09
🔍 Latest deploy log https://app.netlify.com/projects/antenna-preview/deploys/6ac8757e5993480008d7261b
😎 Deploy Preview https://deploy-preview-1503--antenna-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
Lighthouse
Lighthouse
1 paths audited
Performance: 48 (🔴 down 17 from production)
Accessibility: 81 (🔴 down 8 from production)
Best Practices: 92 (🔴 down 8 from production)
SEO: 92 (no change from production)
PWA: 80 (no change from production)
View the detailed breakdown and full score reports
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant