You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Store a suggested confidence threshold for each model #1465
Each classification model has a confidence score below which its predictions should not be trusted, and that score differs from model to model. Today Antenna has one score threshold per project (the project's default filters), so a project that uses two models, or switches models, either hides good predictions from one or shows weak predictions from the other. This ticket stores a suggested confidence threshold on each model, so default filters, class masking, exports and the interface can use the right threshold for whichever model produced a prediction.
What is needed
Storage. A suggested score threshold on each classification algorithm (model), with where it came from: declared by the processing service, set by an administrator, or produced by a calibration run.
Source. The processing service can declare it in its algorithm information (/info), the same way it declares the category map. Administrators can override it. A calibration or evaluation run can propose a value.
Use. Default filters fall back to the model's suggested threshold when a project has not set its own. The interface shows the threshold that applies next to a prediction's score. Exports record it.
Provenance. When the suggested threshold changes, earlier runs still record the value they used, since a job's settings are stored with the job.
Open questions
Whether the project threshold overrides the model's, or the model's overrides the project's, or the stricter of the two applies.
Summary
Each classification model has a confidence score below which its predictions should not be trusted, and that score differs from model to model. Today Antenna has one score threshold per project (the project's default filters), so a project that uses two models, or switches models, either hides good predictions from one or shows weak predictions from the other. This ticket stores a suggested confidence threshold on each model, so default filters, class masking, exports and the interface can use the right threshold for whichever model produced a prediction.
What is needed
/info), the same way it declares the category map. Administrators can override it. A calibration or evaluation run can propose a value.Open questions
Related
#1360 (offline evaluation of post-processing filters), #1316 (which taxa need verification), #857 (rank roll-ups), #1457 (post-processing outputs), #1461.