Repository navigation
Score a classifier against a fixed set of occurrences, and show where it does well - #1493
Draft
mohamedelabbas1996 wants to merge 1 commit into
Draft
mohamedelabbas1996 wants to merge 1 commit into
mohamedelabbas1996 wants to merge 1 commit into
Conversation
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for antenna-preview canceled.
|
A project runs more than one classifier over time and has no way to say whether a newer one is better. Accuracy quoted from different data is not a comparison, so scoring here is always against an occurrence set: the same occurrences, every time, for every model. A run compares an algorithm's stored predictions with what people identified. No images are opened and no model runs, so it is a pass over rows and returns at once. Only occurrences a person identified are counted: determination falls back to the model's own prediction when nobody has identified one, and scoring those would have the model agreeing with itself. The species a model can be asked about are resolved to taxa rather than compared as label text, because a category map label is the name the model was trained under and need not be the name the taxon is stored under here. The result also carries what it was measured over, since an accuracy over six of a model's sixty species reads the same as one over all of them. The scores surface in three places: an algorithm's own list of evaluations, a species page showing how each scored model did on it, and the best model for a taxa list, ranked on the per-species average rather than the plain share because trap data is long-tailed. Stacked on the occurrence sets branch, which the scoring runs against.
mohamedelabbas1996
force-pushed
the
feat/evaluate-classifiers
branch
from
October 8, 2026 21:19
28c0994 to
8b36ef4
Compare
mihow
added this pull request to stack #1498
October 8, 2026 21:52
✅ Deploy Preview for antenna-ssec ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A project runs more than one classifier over time and has no way to say whether a newer one
is better. Accuracy quoted from different data is not a comparison, so scoring here is
always against an occurrence set: the same occurrences, every time, for every model.
A run compares an algorithm's stored predictions with what people identified. No images are
opened and no model runs, so it is a pass over rows and returns at once.
Based on #1492, which this scores against, so the diff here is the scoring work alone.
List of Changes
evaluate_algorithmjob type;ami/ml/evaluation.pycompares stored classifications with identificationsIdentification, not ondetermination— that falls back to the model's own prediction, which would have a model agreeing with itselfspecies_in_setandspecies_predictablein the result and the job log: 1.00 over six of sixty species reads the same as 1.00 over all of themAlgorithmEvaluation(one per algorithm and set) andTaxonEvaluation(one per species); a re-run replaces rather than appendsevaluationsonAlgorithmSerializeralgorithm_performanceonTaxonSerializerbest_modelonTaxaListSerializer, annotated in the viewset so a page costs one queryrun_evaluate_algorithm_job, with an object-level backfill for existing projectsRelated Issues
Based on #1492 (occurrence sets). Split out of #1407, where this began.
Detailed Description
model that only handles the common species would otherwise win every list.
Every read is scoped to the project, or to what the user can see when the route carries no
project id — the algorithm detail route does not.
ami/ml/migrations/0029collides by number with Store feature vectors from any model for every detection, indexed for the ways they are read #1462. Whichever lands second getsrenumbered; nothing else about them overlaps.
this PR.
Testing
ami.ml.test_evaluation— 30 tests: scoring, the self-scoring and label-matching guards,coverage reporting, cross-project exposure, the per-species and best-model reads, and a
pinned query count for a page of taxa lists.