Skip to content

Score a classifier against a fixed set of occurrences, and show where it does well - #1493

Draft
mohamedelabbas1996 wants to merge 1 commit into
feat/occurrence-setsfrom
feat/evaluate-classifiers
Draft

mohamedelabbas1996 wants to merge 1 commit into
feat/occurrence-setsfrom
feat/evaluate-classifiers

Conversation

@mohamedelabbas1996

@mohamedelabbas1996 mohamedelabbas1996 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A project runs more than one classifier over time and has no way to say whether a newer one
is better. Accuracy quoted from different data is not a comparison, so scoring here is
always against an occurrence set: the same occurrences, every time, for every model.

A run compares an algorithm's stored predictions with what people identified. No images are
opened and no model runs, so it is a pass over rows and returns at once.

Based on #1492, which this scores against, so the diff here is the scoring work alone.

List of Changes

# Change (effect) How (implementation)
1 A model can be scored on a fixed set evaluate_algorithm job type; ami/ml/evaluation.py compares stored classifications with identifications
2 Only what people identified counts Filtered on a non-withdrawn Identification, not on determination — that falls back to the model's own prediction, which would have a model agreeing with itself
3 A species is matched by taxon, not by label text A category map label is the name the model was trained under; comparing it as text silently drops species whose stored name differs
4 The score says what it was measured over species_in_set and species_predictable in the result and the job log: 1.00 over six of sixty species reads the same as 1.00 over all of them
5 Scores are stored, not recomputed AlgorithmEvaluation (one per algorithm and set) and TaxonEvaluation (one per species); a re-run replaces rather than appends
6 An algorithm page lists how it has scored evaluations on AlgorithmSerializer
7 A species page shows how each model did on it algorithm_performance on TaxonSerializer
8 A taxa list names its best model best_model on TaxaListSerializer, annotated in the viewset so a page costs one query
9 Only ML data managers can start a scoring run run_evaluate_algorithm_job, with an object-level backfill for existing projects

Related Issues

Based on #1492 (occurrence sets). Split out of #1407, where this began.

Detailed Description

  • Ranking by the per-species average, not the plain share. Trap data is long-tailed, so a
    model that only handles the common species would otherwise win every list.
  • Scoping. Algorithms and taxa are shared across the platform; evaluation sets are not.
    Every read is scoped to the project, or to what the user can see when the route carries no
    project id — the algorithm detail route does not.
  • ami/ml/migrations/0029 collides by number with Store feature vectors from any model for every detection, indexed for the ways they are read #1462. Whichever lands second gets
    renumbered; nothing else about them overlaps.
  • No interface yet. The numbers are on the API; the screens that read them are not in
    this PR.

Testing

ami.ml.test_evaluation — 30 tests: scoring, the self-scoring and label-matching guards,
coverage reporting, cross-project exposure, the per-species and best-model reads, and a
pinned query count for a page of taxa lists.

@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@netlify

netlify Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for antenna-preview canceled.

Name Link
🔨 Latest commit 8b36ef4
🔍 Latest deploy log https://app.netlify.com/projects/antenna-preview/deploys/6ac808fae4ace3000859ee44

A project runs more than one classifier over time and has no way to say whether a
newer one is better. Accuracy quoted from different data is not a comparison, so
scoring here is always against an occurrence set: the same occurrences, every
time, for every model.

A run compares an algorithm's stored predictions with what people identified.
No images are opened and no model runs, so it is a pass over rows and returns at
once. Only occurrences a person identified are counted: determination falls back
to the model's own prediction when nobody has identified one, and scoring those
would have the model agreeing with itself.

The species a model can be asked about are resolved to taxa rather than compared
as label text, because a category map label is the name the model was trained
under and need not be the name the taxon is stored under here. The result also
carries what it was measured over, since an accuracy over six of a model's sixty
species reads the same as one over all of them.

The scores surface in three places: an algorithm's own list of evaluations, a
species page showing how each scored model did on it, and the best model for a
taxa list, ranked on the per-species average rather than the plain share because
trap data is long-tailed.

Stacked on the occurrence sets branch, which the scoring runs against.
@mohamedelabbas1996
mohamedelabbas1996 force-pushed the feat/evaluate-classifiers branch from 28c0994 to 8b36ef4 Compare October 8, 2026 21:19
@mihow
mihow added this pull request to stack #1498 October 8, 2026 21:52
@netlify

netlify Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for antenna-ssec ready!

Name Link
🔨 Latest commit 8b36ef4
🔍 Latest deploy log https://app.netlify.com/projects/antenna-ssec/deploys/6ac81096298c6d0009ea9953
😎 Deploy Preview https://deploy-preview-1493--antenna-ssec.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant