Skip to content

[Draft] Show how each model performs, and let people start the jobs behind it - #1423

Closed
mohamedelabbas1996 wants to merge 3 commits into
RolnickLab:mainfrom
mohamedelabbas1996:feat/model-performance-ui
Closed

mohamedelabbas1996 wants to merge 3 commits into
RolnickLab:mainfrom
mohamedelabbas1996:feat/model-performance-ui

Conversation

@mohamedelabbas1996

@mohamedelabbas1996 mohamedelabbas1996 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Antenna can now score a model against a fixed set of verified occurrences, and retrain a classifier head from labels people confirmed. Neither was visible: the numbers existed only in the API, and the jobs behind them could only be started by posting to it by hand. This is the interface for both.

Split out of #1407 so the screens can be reviewed on their own. The API fields they read land in that PR; until it merges the new columns read n/a rather than breaking.

Screenshots

Taxa list — verified crops ready to train a head on, per species Table columns — the column can be switched off
Training images ready column Table columns menu
A species that has been scored — every model, its accuracy here, and the set used A species nothing has scored yet — reads n/a, not 0
Algorithms and performance on a species Species with no performance yet
Taxa lists — the model that handles the list's species best, and the average it won on Algorithm panel — training data ready, and what the model has been scored on
Best model on a taxa list Training data and evaluation sections
Jobs table — the new job types are named instead of blank The same table filtered to Evaluate algorithm
Job types in the jobs table Jobs filtered by type
Create job — defaults to an ML pipeline run, as before Train classifier — asks for a head, not captures
Job form in its default state Train classifier fields
Evaluate algorithm — asks for a head and an evaluation set Generate embeddings — asks for a pipeline
Evaluate algorithm fields Generate embeddings fields
The head picker offers only heads that can be retrained
Trainable heads in the picker

List of Changes

# What it does for a user How
1 A species list shows the verified crops ready to train a head on training_crops_ready, requested only when the column is visible
2 A species page lists every model scored on it, with that species' accuracy and the set used algorithm_performance on the taxon, each row linking to the model
3 A taxa list names the model that handles its species best best_model, with the per-species average it was ranked on
4 An algorithm's panel shows what it has been scored on evaluations table in the details dialog
5 It also shows whether a retrain is worth starting crops ready, species covered, and crops still waiting for an embedding
6 Training, embedding and evaluation jobs can be started from the jobs form job type picker; each type asks only for the fields it needs
7 Those jobs show their type in the table and can be filtered by it they were blank before, because the frontend's job type list had never been extended

Detailed Description

  • Null is not zero. Every new field reads n/a when the API returns null — on a project where nothing has been evaluated, and on any deployment where Retrain a classifier head from verified identifications, and score what it produces #1407 has not landed. The crop count is deliberately null rather than 0 when it was not requested; a zero would read as a species with nothing to train on.
  • The form reads each job type's declared requirements: a pipeline to process or embed, a head to retrain, a head and an evaluation set to score. It no longer asks for captures for a job that never looks at one.
  • The crop count is opt-in. ?with_training_crop_counts=true is sent only when that column is visible, following the pattern Add example occurrence to the taxa list to expedite the verification of species presence #1365 established for the Example column. The count grows with the project rather than the page, and most viewers never show it.

How to Test the Changes

Against a project with verified occurrences and at least one scored algorithm:

  1. Taxa — the "Training images ready" column, and the same count under Table columns.
  2. Taxa › a species — "Algorithms & performance", one line per scored model, each opening that model.
  3. Taxa lists — "Best model", naming the model and the per-species average it won on.
  4. Algorithms › one algorithm — "Training data" and "Evaluation" sections.
  5. Jobs › Create new job — switch Type and watch the fields change; create an Evaluate algorithm job and start it.

Verified in a browser against a local stack: a job created from this form ran and scored 1.000 over 15 occurrences. yarn type-check, yarn lint, yarn format --check and yarn test (52 tests) all pass.

Known gaps

  1. The job form offers types the API rejects until Retrain a classifier head from verified identifications, and score what it produces #1407 lands. Picking Train classifier before then returns a validation error rather than a job.
  2. The taxa list shows the best model's per-species average only; the plain share comes back in the same payload and is not displayed.

🤖 Generated with Claude Code

The platform can now score a model against a fixed set of verified occurrences and
retrain a classifier head, but none of it was visible: the numbers existed only in
the API, and the jobs could only be started by posting to it by hand.

The species table gains the verified crops ready to train on, and a species page
lists every model that has been scored on it, with the accuracy for that species and
the set it was scored on. A taxa list names the model that handles its species best,
ranked on the per-species average, and an algorithm's panel shows what it has scored
and the size of the training set waiting for it. Every one of these reads "n/a"
until something has actually been evaluated, so the columns are honest on a project
where nothing has run yet.

The jobs form asks which kind of job to create and then asks only for what that kind
needs — a pipeline for a processing run, a head to retrain, a head and an evaluation
set to score. Training, embedding and evaluation jobs also show their type in the
jobs table and can be filtered by it, which they could not before.

The API fields these read land separately; until then the new columns read "n/a"
rather than breaking. See RolnickLab#1407.

Co-Authored-By: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

👷 Deploy request for antenna-ssec pending review.

Visit the deploys page to approve it

Name Link
🔨 Latest commit 6ab0ebe

@coderabbitai

coderabbitai Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@netlify

netlify Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for antenna-preview ready!

Name Link
🔨 Latest commit 6ab0ebe
🔍 Latest deploy log https://app.netlify.com/projects/antenna-preview/deploys/6aac541cdad41e0008eaf60f
😎 Deploy Preview https://deploy-preview-1423--antenna-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
Lighthouse
Lighthouse
1 paths audited
Performance: 56 (🔴 down 9 from production)
Accessibility: 81 (🔴 down 8 from production)
Best Practices: 92 (🔴 down 8 from production)
SEO: 92 (no change from production)
PWA: 80 (no change from production)
View the detailed breakdown and full score reports
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

mohamedelabbas1996 and others added 2 commits September 17, 2026 16:22
Occurrence sets moved out of the ml namespace to sit beside the other named
collections of domain objects, so the picker follows.

Co-Authored-By: Claude <noreply@anthropic.com>
Comparing two heads on the same set showed 1.00 over 15 occurrences against 0.89
over 9. Read plainly the second head looks worse, but six of those occurrences are
species it cannot predict at all, so it was asked fewer questions rather than
getting more wrong. The count was already in the response and not shown.

The scored column now carries what was left out beneath it, so a comparison that
is not like for like says so.

Co-Authored-By: Claude <noreply@anthropic.com>
@mihow

mihow commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Claude says: Hi! A heads-up rather than a request. Michael asked for one Create Job dialog that every job type can appear in, and a draft of it is up in #1447. It adds GET /api/v2/jobs/types/?project_id=N, which lists the job types a member may create in a project, what each one runs on, and a JSON Schema of its settings. The dialog builds its form from that response.

It is written to fit alongside #1407 and this PR rather than replace anything in them:

  • required_fields, required_params and user_creatable on JobType use the same names and defaults as Retrain a classifier head from verified identifications, and score what it produces #1407, and the endpoint returns them for each type. So train_classifier, generate_embeddings and evaluate_algorithm would show up in the dialog once both land, with a simple field for each required param.
  • If one of those jobs later gets a pydantic config_schema, it gets a generated form (labels, help text, pickers, numeric bounds) with no frontend change.
  • The job-type picker here could read its list from the endpoint instead of a hardcoded one, if that turns out to be useful. No pressure either way.

Nothing needs to change in this PR for now. Happy to line things up whenever it suits you.

@mihow

mihow commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Claude says: A correction to my earlier note. After review, #1447 no longer uses required_fields / required_params. Each job type the Create Job dialog offers now declares one pydantic model for everything a new job takes, and a new job's inputs all arrive in params.config (with job_type_key required). So for train_classifier, generate_embeddings and evaluate_algorithm to appear in the dialog, each would declare a small model, for example an algorithm_key field, instead of the tuples. The dialog then generates the form from it. The job-type list on GET /jobs/types/ is otherwise as described. Sorry for the churn, and happy to help line it up when you get there.

@mohamedelabbas1996

Copy link
Copy Markdown
Contributor Author

Moved to the lab repo and rebased onto the evaluation API it reads: #1506.

This branch lived on my fork, targeted main, and carried a copy of the training job form from before the retraining work was split. That form was rebuilt in #1504, and the training-data tab it added needs an endpoint from #1502, so neither came across. #1506 is the evaluation UI only, sitting on #1493.

Closing this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants