You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Classifier species lists as taxa lists and support for public lists #1424
The larger goal is that a user is only shown results that are accurate and useful for the place
they monitor (#890). Most classifiers in Antenna cover far more species than occur at any one site,
so the platform needs a way to say "these are the species that matter here" and to apply that to
predictions. Taxa lists are that mechanism, and the work has been arriving in stages: taxa lists
and their editor (#1094, #1104), filtering taxa and occurrences by a list (#745, #1008, #1347),
class masking against a list (#999) on the post-processing framework (#1289), and verifying
presence species by species (#1320, #1365). Still ahead are building a list from a region (#1364,
draft #1367), a user-facing way to turn masking on (#1406), and the earlier geo-fencing idea
(#797). Further out, members add their own taxa for identification, every taxon shows whether a
model can predict it or only a person can identify it, and taxa carry regional information.
This ticket supplies the piece those all lean on: an authoritative, browsable answer to "what can
each model predict", shared across projects, that a project can copy and make its own. It also
settles how shared and project-owned objects coexist, which taxa lists, processing services and
future curator roles all need. A full list of related issues and pull requests is at the end.
Summary
People who use Antenna regularly ask "can this model predict species X?" and today the platform
cannot answer, because the full set of classes a classifier knows (its category map) is not visible
anywhere a user can browse it. The closest thing, the "Taxa returned by " lists, only
holds species the model has already predicted somewhere, so it cannot show a "no".
This ticket proposes using the existing Taxa Lists feature to close that gap. Every classifier gets
one taxa list holding everything in its category map, public when the classifier itself is
available to every project. A project member can copy that list
into their own project, edit the copy (remove species that do not occur in their region, add
missing ones), and later choose it as the list for class masking. The first users are partners who
identify with a regional classifier and want a list for a neighbouring region derived from it.
Doing this well depends on a second piece of work: Antenna has several kinds of objects that can be
either shared across all projects or private to some projects (taxa lists, taxa, tags, pipelines,
processing services), and each one handles that split differently today. Public taxa lists already
exist in the database but appear to be unreachable through the API. The ticket therefore includes a
consistent rule for "public versus project-scoped" as related work.
Terminology
Use public for an object available to every project and project-scoped (or private) for an
object tied to specific projects. Avoid "global": in this domain it already means geographic scope
(the global moth classifier, global species lists), and the two meanings collide in exactly the
places this feature touches. Existing code and comments that say "global list" should be renamed as
they are touched.
Current state
Based on code reading unless marked as measured.
Entity
Relation to project
Public rows exist?
Project-scoped list returns public rows?
Filter parameter
TaxaList
M2M projects
Yes. get_or_create_for_project(project=None) creates them; used by taxa import, taxa update, and "Taxa returned by" lists
No. TaxaListViewSet requires a project and filters projects=project
None
Tag
nullable FK project
Yes
Yes. project = X OR project IS NULL
None
Taxon
M2M projects
Nearly all taxa have no project
Partly. The project view is driven by occurrences and the default taxa filters, not the M2M; the M2M only feeds visible_for_user
include_unobserved (different concept)
Pipeline
M2M projects plus ProjectPipelineConfig
None observed
No. Requires the M2M link and an enabled config
None
ProcessingService
M2M projects (blank=True)
One observed, likely an orphan
No. Requires a project and filters projects=project
projects filterset field
Algorithm
none (through pipelines)
n/a
List scoped through enabled pipelines; detail unscoped
None
Measured in one development database: 31 of 39 taxa lists have no project (mostly "Taxa returned
by" lists), 13,291 of 13,298 taxa have no project, 0 of 15 pipelines, 1 of 15 processing services,
and 4 of 4 tags.
Two behaviours that do not show up when reading the viewsets alone:
The shared visible_for_user queryset method filters M2M models on projects__draft = false OR owner OR member. That is a join, so a row with no projects never matches. For any non-superuser
a public TaxaList, Pipeline or ProcessingService is filtered out before the viewset's own
project filter runs. Only Tag (nullable FK handled explicitly) and Taxon (own override) escape
this. Measured in the same development database by calling the queryset method directly: of the
31 taxa lists with no project, 0 are visible to an anonymous user and 0 to a non-superuser
member; the one processing service with no project is likewise hidden. A regression test
should pin this once the rule changes.
The object-permission helper for M2M models returns no update or delete permission unless the
object belongs to the active project. That is the right outcome for public objects, but it is
incidental rather than designed, and the write path (IsProjectMemberOrReadOnly) checks project
membership only, not whether the target object is public.
Earlier precedent: the integration branch in #820 made taxa and taxa lists without a project
visible inside every project, by filtering on projects = X OR projects IS NULL, the same shape Tag
uses on main. It also made both M2M fields blank=True and cleared the historical taxon-to-project
assignments, which it describes as mostly arbitrary. It left a TODO for the permission check on
project-scoped rows. On main those M2M fields are still required in forms, while ProcessingService.projects is already blank=True. None of that branch's scoping has landed.
How taxa come into being from model output today, from reading the result-saving code. When a
processing service registers an algorithm, Antenna stores the category map and creates no taxa.
Taxa are created later, one at a time, the first time a label is the top prediction of a saved
classification: the name the service returned is looked up by name or search name, and if nothing
matches a new taxon is created with that name and the rank of the category at the top score's
index, with no parent. Labels that never come out on top never get a taxon. Nothing checks that the
returned name matches the category map's label at that index, or that the scores have the same
length as the map, so a name outside the category map is accepted silently and given the rank of
an unrelated category. Two workers saving the same new label at the same moment would both try to
create it, and the second would fail on the unique name; this has not been reproduced.
Work that exists already: a local branch adds AlgorithmCategoryMap.resolve_taxa(), Algorithm.get_or_create_taxa_list() and a management command create_taxa_lists_from_category_maps (with --dry-run), with tests. Labels resolve to taxa by
name or search_names, the same rule used when saving classifications. Not yet pushed.
Core work
One taxa list per classifier, holding its full category map, public when the classifier
is. Land the existing branch. Create or refresh the list when a processing service registers
a category map, so new models need no operator step.
Taxa at registration. Resolve every label to a taxon at registration rather than as predictions arrive, so
the list is complete from the first day and saving results becomes a lookup that never writes
to the shared taxonomy. Taxa created this way are public, since a model and its labels are
shared by every project that runs it, and they are marked as created by a pipeline and not yet
placed in the taxonomy, so a Taxonomy Curator has a queue to work through. A large classifier
has tens of thousands of labels, so this runs as a bulk background step, not inside the
registration request. Creation at save time stays as a fallback for services that never
advertised a map.
Predictions outside the category map. Treat the category map as the model's
declared output and the score index as the source of truth. If the returned name disagrees
with the label at that index, keep the classification, resolve the taxon from the index, and
record a warning on the job. If the scores and the map differ in length, fail loudly, because
that almost always means the service is running a newer model than the one registered and the
stored map is stale. Neither case adds a member to the managed list.
Make the taxa list the primary record of what a model can predict, with fixed membership.
The relationship the rest of the platform uses is taxon → taxa list → algorithm: an algorithm
points at its species list, and "which models know this taxon" is answered through list
membership. The category map steps back to being the verbatim copy of what the processing
service published: unchanging, kept for the order of classes that score vectors depend on, and
rarely read elsewhere. Algorithms that share a category map share a list. A list that an
algorithm points at is managed: its name and description can be edited, but taxa cannot be
added or removed by anyone, including superusers and curators. The only writer of its
membership is the sync from the category map, which is safe to re-run and picks up labels that
resolve later (a new taxon, a synonym fix as in Resolve classifier labels and taxa-list matches through synonym_of, not name-only #1405). Anyone who wants a different set of
species copies the list. The serializer says which algorithms a list describes, so the UI can
hide the add and remove controls. Removing a managed list never deletes taxa, including taxa the
sync created, and project copies are unaffected. Taxa lists already have a description field,
shown in the list table; the sync should fill it for managed lists (which algorithms, how many
labels, how many resolved) and the copy should record its source there as well as in a link. A managed list is only as public as its algorithm. Processing services and pipelines can
be added by a user to just their own project, by design. The list for an algorithm that only
such a service offers is scoped to that project and is not public. This ties the list's
visibility to the processing service's, so the is_public field on processing services is
needed for this step rather than later.
Make public lists visible inside a project, and show the relationship. The project's Taxa
Lists page shows the project's own lists plus public ones, clearly labelled and read-only. The
algorithm page links to the classifier's species list in the Taxa views, next to the existing
link to the category map API endpoint, and the list page names the algorithms it describes.
Show which models can recognise a taxon. Start with filters in both directions: "show me
all taxa this algorithm can predict" on the taxa views, and "show me all algorithms that can
predict this taxon" on the algorithm list, with the taxon detail view listing them too. The
answer includes public algorithms and the active project's own; an algorithm private to
another project never appears and is never counted. With taxon → taxa list → algorithm as the
primary relationship these are indexed lookups through list membership, the same cost as a
direct taxon-to-algorithm table, rather than a search through every category map's JSON.
Showing the algorithms as small labels on every row of the taxa list views is wanted but
postponed: it needs one aggregate or prefetch for the whole page, a query-count test on a
multi-row fixture and a query plan checked on the largest classifier, and the filters deliver
most of the value first. This replaces the direct taxon-to-algorithm relationship and coverage
flag in draft Generate a project taxa list from a region, and track which species the models can predict #1367, whose regional lists should read coverage through the managed lists
instead.
Copy a list into a project. A "Copy to project" action that creates a project-scoped list
with the same members and records where it came from. Requires the project's
create-taxa-list permission.
Curate the copy. The existing add and remove endpoints cover single edits. Check whether
bulk removal and search-to-add are good enough for a list of a few thousand species.
Related work: one rule for public versus project-scoped objects
Proposed rule, to apply to TaxaList and ProcessingService first (managed lists depend on both), then
Pipeline, Tag and Taxon:
An explicit is_public boolean field on the model. A public row is returned for every project,
whatever its project links. A row that is not public and has no project is an orphan, visible to
superusers only, never silently shared. The migration backfills is_public = true for existing
rows that have no project, after checking them (see the verification list).
The projects M2M is blank=True so a public row needs no project link.
A shared queryset method, for example for_project(project, include_public=True), returns rows
linked to the project plus public rows. Viewsets call it instead of hand-written filter(projects=project).
One query parameter with the same name and default everywhere: include_public, default true,
parsed with the existing boolean parameter helper so bad input returns 400.
visible_for_user treats public rows as visible to everyone.
Each serializer exposes is_public read-only so the UI can label and lock public rows.
Creating, editing or deleting a public object requires a platform-level permission for that
model (for example "manage public taxa lists"), checked without reference to any project. To
begin with only superusers hold it. The check should be a named permission rather than a
hard-coded superuser test, so that platform roles can be granted it later without touching the
endpoints (see follow-up work). Project members can edit only objects scoped to a project they
belong to. This overlaps with Review M2M permissions for things that belong to multiple projects #1120 and should be designed with it.
Permission matrix tests (member, non-member, anonymous, superuser) cover both a public and a
project-scoped object for each entity.
Pipelines and processing services are the least symmetric today because availability also depends
on ProjectPipelineConfig.enabled. A public processing service probably means "offered to every
project, enabled per project", which is close to what the default processing service setting does
by adding each new project to one service's M2M. That may be better modelled as a public service
than as an ever-growing project list. This is a direction to discuss, not a requirement of this
ticket.
The public versus project-scoped rule on taxa lists and processing services: the is_public field, the shared queryset method, the visibility fix, the include_public parameter, the platform-level permissions, and permission tests for both
none
B
Algorithm points at its managed list; labels resolved and the list synced at registration as a background step; fixed membership; list visibility follows the algorithm; the score index is trusted when saving results
A
C
Copy a list into a project
A, alongside B
D
Which models recognise a taxon: filters in both directions and the taxon page; labels in list rows later
B
E
Interface: public label, locked managed lists, copy button, link from algorithm to list, algorithm labels on taxa
the API shapes from A, B and D
Two registration defects found along the way are tracked separately: empty category maps piling
up (#1426), and pipeline and algorithm key collisions between public and project-scoped services
(#1427).
Follow-up work
Platform-level roles for people who are not superusers. Every role in Antenna today is scoped
to one project; nothing sits between a project role and superuser. Two roles are anticipated: an
Antenna admin or manager role with cross-project access (public processing services, platform
defaults), and a Taxonomy Curator role that can curate public taxa lists and, in time, taxa
themselves. The platform-level permissions introduced here are what those roles would be granted.
Related: Refactor role-project relationship to use foreign key instead of string parsing #1100 (how roles relate to projects).
Per-species accuracy and training data, per model. The aim is to answer "which algorithm
predicts species X most accurately?", and the same for a group such as the micro-moths or the
Crambidae, and to show how many training images stand behind each class (Show the species-wise accuracy in the machine prediction component #788). The published
figures belong in the category map, since they are facts the model's authors state about each
class and the map is the service's verbatim copy; category entries already accept extra keys.
The through-model on list membership then carries them next to the taxon, together with the
original label and class index, which is what makes them comparable across models and
summable up the taxonomy. List membership alone loses the label and index, so a label that
resolved to a differently named taxon through a synonym cannot be explained to the user, and
class masking has to re-derive the index. No model ships these figures today; agree the keys
with the processing-service schema first, and keep the through-model in mind when writing the
sync so it can be added without reworking it.
Let project members add their own taxa, for identification. A project often needs names the
shared taxonomy lacks: an undescribed or provisional species, a morphospecies, a regional name
not yet imported. Members should be able to create such a taxon inside their project, add it to
their lists, and pick it when identifying. This is the project-scoped taxon, and it is the case
that gives is_public a purpose on taxa: shared reference taxa are public, a member's additions
are scoped to their project until a Taxonomy Curator promotes or merges them. Two things stand in
the way today. Taxon names and display names are unique across the whole platform, so two
projects could not both have a "Morphospecies 1"; uniqueness would need to hold among public
taxa and within each project instead. And nothing outside the admin creates a taxon ([integration] OOD Features #820 had a
first version of a create-species form). The identification search must return public taxa plus
the project's own, and nothing from other projects.
Make "available" and "predictable" visibly different. A taxon can be available to identify
with (it exists in the taxonomy or the project's lists), predictable by a model (it is in the
category map of a classifier the project runs), both, or, for a member's own addition, available
only. Wherever taxa are listed or chosen (list pages, the taxon page, the identification picker)
the distinction should be shown rather than implied, for example "predicted by 2 models" against
"identification by people only". The recognised-by lookup in the core work is what feeds this. It
also sets expectations: a species missing from the project's results may be absent from the
site, or simply not something the model can output.
Regional information for taxa. Record where each taxon is known to occur, so lists, masking
and the identification picker can all be narrowed to a project's region without each feature
fetching it again. The shape to mirror is Darwin Core's species distribution record (a taxon, a
location identifier, occurrence status, establishment means, source) with GADM identifiers as
the location vocabulary, since GADM's nested country, state and district codes match how
deployments are already placed and how Generate a project taxa list from a region, and track which species the models can predict #1367 asks GBIF for a region's species. A regional list
then becomes a saved query over this data rather than a one-off fetch. Design only for now;
the exact Darwin Core terms and GADM version need checking against the published standards.
Carry the public and managed flags across the other shared models. This ticket gives taxa
lists two independent answers: is_public says who may see a row, and the managed flag says who
owns its contents (a list an algorithm points at mirrors a category map and refuses hand edits).
Pipelines and algorithms arrive from what a processing service advertises and are then edited in
the admin with nothing recording that the next registration may overwrite the edit; the default
processing service comes from settings rather than from a person; taxa created from classifier
labels are awaiting a curator and nothing says so. Declaring both flags on the shared base, with
the same read-only serializer fields, would let a client label and lock such rows the same way
everywhere. Belongs with Review M2M permissions for things that belong to multiple projects #1120 and the platform roles above.
Import taxa and taxa lists from the UI. Today both arrive only through management commands
run by an operator (import_taxa, update_taxa, from Update command for importing taxa from external lists #939). The aim is a general import
framework that mirrors the export framework: a registry of import formats, a job that runs the
import in the background and reports what it created, matched and rejected, and a page to start
one. Taxa and taxa lists would be the first formats, built on the existing commands' logic; feat: dataset import framework #1254 proposes the framework for datasets. The taxa list CSV export (feat(exports): add taxa_list_csv export format #1293) defines the natural
round-trip format. Imports into public taxa or public lists need the curator permission
described above.
Decide the fate of the "Taxa returned by " lists: keep, rename, or hide.
Roll recognisability up the taxonomy: on a genus or family page, show how many of its species a
given model can recognise, and later how accurately.
Rename remaining uses of "global" for cross-project objects in code and documentation.
Decisions
Settled:
"Public" is an explicit is_public field, not inferred from a row having no projects. Inferring
it carries a silent risk: deleting a project removes its M2M rows, so a private list whose only
project is deleted would become public to everyone. Processing services take the same approach,
in the same first step, because a managed list's visibility follows its algorithm's service.
include_public defaults to true, which matches how tags behave today and makes classifier
lists discoverable.
The managed taxa list is the one store of which models can predict a taxon. The primary
relationship is taxon → taxa list → algorithm; the category map is the service's verbatim,
unchanging copy. Draft Generate a project taxa list from a region, and track which species the models can predict #1367's direct relationship and coverage flag are dropped in favour of
it. Membership is fixed, and removing a managed list never deletes taxa.
A managed list is only as public as its algorithm. "Known by N algorithms" counts public
algorithms and the active project's own, and never one private to another project.
Recognised-by ships as filters in both directions first; per-row labels in list views come later.
Still open:
Where public lists appear: in every project's Taxa Lists page, only via the algorithm page,
or both.
User-facing name of the per-classifier list: "Category map of " (current) or
something friendlier such as " species list".
When Taxon gets is_public. Nearly every taxon has no project today and taxa are treated as
shared reference data, so the field changes nothing until members can add their own taxa (see
follow-up work). It can wait for that work, but the taxa-list rule should be written so taxa
can adopt it unchanged.
What we still need to verify
What makes an algorithm private. Algorithms have no link to projects and no visibility of
their own; they are reached through pipelines and processing services. Algorithm keys and
pipeline slugs are unique across the whole platform, and registration matches on them, so a
project's own service that advertises an existing key attaches to the same shared algorithm
row, and whichever service registers a key first decides its category map. Proposed working
definition: an algorithm is public if at least one public processing service offers it, and
otherwise visible only to the projects of the services that do. Check whether a project-scoped
service can claim or alter a key that a public service uses, and whether keys need a namespace.
Confirm that a category map never changes after it is created. Registration appears to create a
map once per algorithm and never update it, which is what makes fixed list membership sound. The
admin can still edit a map's labels, and the stored labels hash is only computed when empty, so
an edit would leave both the hash and the list stale. Decide whether to make the labels read-only
in the admin.
Empty category maps pile up at registration. Measured in one development database: 1,250
category maps, of which 1,232 have no algorithm, and every one of those 1,232 is empty (no
labels, no description, not referenced by any classification). The 18 maps in use all have
distinct label sets, so identical label sets being duplicated was not observed. The apparent
cause, from reading the registration code: an algorithm whose processing service reports a
category map with no entries (a detector, in the case observed) fails the "has a valid category
map" check every time, so each registration creates a fresh empty map, points the algorithm at
it and orphans the previous one. This deserves its own small fix (treat an empty map as no map,
and clean up the orphans) and needs a count on a production database. For this ticket it means
lists are created only for maps that have labels and an algorithm. Separately, registration
never reuses an existing map with the same labels, so two algorithms with identical classes
would get two lists; the stored labels hash would allow reuse.
Check what creating missing taxa does for labels that are not taxa (for example the two classes
of a binary moth / non-moth classifier), so the sync does not add junk rows to the shared
taxonomy.
The visible_for_user behaviour above is measured at the queryset level only. Add an API-level
test (with and without a project id) so the endpoint behaviour is pinned as well.
Check whether the one processing service with no projects is an orphan or intentional, and review
the taxa lists with no project before the backfill marks them public.
Check what happens to a project-scoped taxa list when its only project is deleted.
Measure the list endpoint with public lists included on the largest project, with EXPLAIN (ANALYZE): the taxa-count annotation over lists of a few thousand taxa, and the OR projects IS NULL condition, which can defeat an index the way other OR filters in this
codebase have.
Confirm how many category-map labels fail to resolve to a taxon per classifier (dry run of the
management command), since unresolved labels make the public list an incomplete answer.
Related issues and pull requests
Area
Reference
Goal
#890 results that are accurate and useful; #797 geo-fencing for regional species lists
Taxa lists, delivered
#1094 configurable taxa lists; #1104 list endpoints follow the membership API pattern; #1119 review follow-ups; #1001 list name as page title; #745 and #1008 filter by list, and inverted; #927 taxa filters in the backend; #1347 keep list filters when opening a taxon's occurrences; #1320 and #1365 presence verification with an example occurrence
Taxa lists, open
#1081 backend extensions, including "only admins can make a list shared by all projects"; #1293 taxa list CSV export; #1364, #1366 and draft #1367 lists from a region and which species models can predict
Algorithms and category maps
#785 section for algorithms and category maps; #1368 filter occurrences by the algorithms that ran; #1405 resolve labels through synonyms; #788 species-wise accuracy
Class masking
#999 masking to a species list; #1289 post-processing framework; #997 (closed) and #1406 user-facing masking; #1360 measuring the effect of post-processing
Public versus project-scoped, roles
#820 integration branch with public taxa and lists; #1120 permissions for objects in several projects; #1100 how roles relate to projects
Import
#939 import and update commands for taxa; #1254 dataset import framework; #1080 (closed) taxa management workflow; #871 (closed) reference images for taxa
Where this fits
The larger goal is that a user is only shown results that are accurate and useful for the place
they monitor (#890). Most classifiers in Antenna cover far more species than occur at any one site,
so the platform needs a way to say "these are the species that matter here" and to apply that to
predictions. Taxa lists are that mechanism, and the work has been arriving in stages: taxa lists
and their editor (#1094, #1104), filtering taxa and occurrences by a list (#745, #1008, #1347),
class masking against a list (#999) on the post-processing framework (#1289), and verifying
presence species by species (#1320, #1365). Still ahead are building a list from a region (#1364,
draft #1367), a user-facing way to turn masking on (#1406), and the earlier geo-fencing idea
(#797). Further out, members add their own taxa for identification, every taxon shows whether a
model can predict it or only a person can identify it, and taxa carry regional information.
This ticket supplies the piece those all lean on: an authoritative, browsable answer to "what can
each model predict", shared across projects, that a project can copy and make its own. It also
settles how shared and project-owned objects coexist, which taxa lists, processing services and
future curator roles all need. A full list of related issues and pull requests is at the end.
Summary
People who use Antenna regularly ask "can this model predict species X?" and today the platform
cannot answer, because the full set of classes a classifier knows (its category map) is not visible
anywhere a user can browse it. The closest thing, the "Taxa returned by " lists, only
holds species the model has already predicted somewhere, so it cannot show a "no".
This ticket proposes using the existing Taxa Lists feature to close that gap. Every classifier gets
one taxa list holding everything in its category map, public when the classifier itself is
available to every project. A project member can copy that list
into their own project, edit the copy (remove species that do not occur in their region, add
missing ones), and later choose it as the list for class masking. The first users are partners who
identify with a regional classifier and want a list for a neighbouring region derived from it.
Doing this well depends on a second piece of work: Antenna has several kinds of objects that can be
either shared across all projects or private to some projects (taxa lists, taxa, tags, pipelines,
processing services), and each one handles that split differently today. Public taxa lists already
exist in the database but appear to be unreachable through the API. The ticket therefore includes a
consistent rule for "public versus project-scoped" as related work.
Terminology
Use public for an object available to every project and project-scoped (or private) for an
object tied to specific projects. Avoid "global": in this domain it already means geographic scope
(the global moth classifier, global species lists), and the two meanings collide in exactly the
places this feature touches. Existing code and comments that say "global list" should be renamed as
they are touched.
Current state
Based on code reading unless marked as measured.
projectsget_or_create_for_project(project=None)creates them; used by taxa import, taxa update, and "Taxa returned by" listsTaxaListViewSetrequires a project and filtersprojects=projectprojectproject = X OR project IS NULLprojectsvisible_for_userinclude_unobserved(different concept)projectsplusProjectPipelineConfigprojects(blank=True)projects=projectprojectsfilterset fieldMeasured in one development database: 31 of 39 taxa lists have no project (mostly "Taxa returned
by" lists), 13,291 of 13,298 taxa have no project, 0 of 15 pipelines, 1 of 15 processing services,
and 4 of 4 tags.
Two behaviours that do not show up when reading the viewsets alone:
visible_for_userqueryset method filters M2M models onprojects__draft = false OR owner OR member. That is a join, so a row with no projects never matches. For any non-superusera public TaxaList, Pipeline or ProcessingService is filtered out before the viewset's own
project filter runs. Only Tag (nullable FK handled explicitly) and Taxon (own override) escape
this. Measured in the same development database by calling the queryset method directly: of the
31 taxa lists with no project, 0 are visible to an anonymous user and 0 to a non-superuser
member; the one processing service with no project is likewise hidden. A regression test
should pin this once the rule changes.
object belongs to the active project. That is the right outcome for public objects, but it is
incidental rather than designed, and the write path (
IsProjectMemberOrReadOnly) checks projectmembership only, not whether the target object is public.
Earlier precedent: the integration branch in #820 made taxa and taxa lists without a project
visible inside every project, by filtering on
projects = X OR projects IS NULL, the same shape Taguses on main. It also made both M2M fields
blank=Trueand cleared the historical taxon-to-projectassignments, which it describes as mostly arbitrary. It left a TODO for the permission check on
project-scoped rows. On main those M2M fields are still required in forms, while
ProcessingService.projectsis alreadyblank=True. None of that branch's scoping has landed.How taxa come into being from model output today, from reading the result-saving code. When a
processing service registers an algorithm, Antenna stores the category map and creates no taxa.
Taxa are created later, one at a time, the first time a label is the top prediction of a saved
classification: the name the service returned is looked up by name or search name, and if nothing
matches a new taxon is created with that name and the rank of the category at the top score's
index, with no parent. Labels that never come out on top never get a taxon. Nothing checks that the
returned name matches the category map's label at that index, or that the scores have the same
length as the map, so a name outside the category map is accepted silently and given the rank of
an unrelated category. Two workers saving the same new label at the same moment would both try to
create it, and the second would fail on the unique name; this has not been reproduced.
Work that exists already: a local branch adds
AlgorithmCategoryMap.resolve_taxa(),Algorithm.get_or_create_taxa_list()and a management commandcreate_taxa_lists_from_category_maps(with--dry-run), with tests. Labels resolve to taxa byname or
search_names, the same rule used when saving classifications. Not yet pushed.Core work
One taxa list per classifier, holding its full category map, public when the classifier
is. Land the existing branch. Create or refresh the list when a processing service registers
a category map, so new models need no operator step.
Taxa at registration. Resolve every label to a taxon at registration rather than as predictions arrive, so
the list is complete from the first day and saving results becomes a lookup that never writes
to the shared taxonomy. Taxa created this way are public, since a model and its labels are
shared by every project that runs it, and they are marked as created by a pipeline and not yet
placed in the taxonomy, so a Taxonomy Curator has a queue to work through. A large classifier
has tens of thousands of labels, so this runs as a bulk background step, not inside the
registration request. Creation at save time stays as a fallback for services that never
advertised a map.
Predictions outside the category map. Treat the category map as the model's
declared output and the score index as the source of truth. If the returned name disagrees
with the label at that index, keep the classification, resolve the taxon from the index, and
record a warning on the job. If the scores and the map differ in length, fail loudly, because
that almost always means the service is running a newer model than the one registered and the
stored map is stale. Neither case adds a member to the managed list.
Make the taxa list the primary record of what a model can predict, with fixed membership.
The relationship the rest of the platform uses is taxon → taxa list → algorithm: an algorithm
points at its species list, and "which models know this taxon" is answered through list
membership. The category map steps back to being the verbatim copy of what the processing
service published: unchanging, kept for the order of classes that score vectors depend on, and
rarely read elsewhere. Algorithms that share a category map share a list. A list that an
algorithm points at is managed: its name and description can be edited, but taxa cannot be
added or removed by anyone, including superusers and curators. The only writer of its
membership is the sync from the category map, which is safe to re-run and picks up labels that
resolve later (a new taxon, a synonym fix as in Resolve classifier labels and taxa-list matches through synonym_of, not name-only #1405). Anyone who wants a different set of
species copies the list. The serializer says which algorithms a list describes, so the UI can
hide the add and remove controls. Removing a managed list never deletes taxa, including taxa the
sync created, and project copies are unaffected. Taxa lists already have a description field,
shown in the list table; the sync should fill it for managed lists (which algorithms, how many
labels, how many resolved) and the copy should record its source there as well as in a link.
A managed list is only as public as its algorithm. Processing services and pipelines can
be added by a user to just their own project, by design. The list for an algorithm that only
such a service offers is scoped to that project and is not public. This ties the list's
visibility to the processing service's, so the
is_publicfield on processing services isneeded for this step rather than later.
Make public lists visible inside a project, and show the relationship. The project's Taxa
Lists page shows the project's own lists plus public ones, clearly labelled and read-only. The
algorithm page links to the classifier's species list in the Taxa views, next to the existing
link to the category map API endpoint, and the list page names the algorithms it describes.
Show which models can recognise a taxon. Start with filters in both directions: "show me
all taxa this algorithm can predict" on the taxa views, and "show me all algorithms that can
predict this taxon" on the algorithm list, with the taxon detail view listing them too. The
answer includes public algorithms and the active project's own; an algorithm private to
another project never appears and is never counted. With taxon → taxa list → algorithm as the
primary relationship these are indexed lookups through list membership, the same cost as a
direct taxon-to-algorithm table, rather than a search through every category map's JSON.
Showing the algorithms as small labels on every row of the taxa list views is wanted but
postponed: it needs one aggregate or prefetch for the whole page, a query-count test on a
multi-row fixture and a query plan checked on the largest classifier, and the filters deliver
most of the value first. This replaces the direct taxon-to-algorithm relationship and coverage
flag in draft Generate a project taxa list from a region, and track which species the models can predict #1367, whose regional lists should read coverage through the managed lists
instead.
Copy a list into a project. A "Copy to project" action that creates a project-scoped list
with the same members and records where it came from. Requires the project's
create-taxa-list permission.
Curate the copy. The existing add and remove endpoints cover single edits. Check whether
bulk removal and search-to-add are good enough for a list of a few thousand species.
Related work: one rule for public versus project-scoped objects
Proposed rule, to apply to TaxaList and ProcessingService first (managed lists depend on both), then
Pipeline, Tag and Taxon:
is_publicboolean field on the model. A public row is returned for every project,whatever its project links. A row that is not public and has no project is an orphan, visible to
superusers only, never silently shared. The migration backfills
is_public = truefor existingrows that have no project, after checking them (see the verification list).
projectsM2M isblank=Trueso a public row needs no project link.for_project(project, include_public=True), returns rowslinked to the project plus public rows. Viewsets call it instead of hand-written
filter(projects=project).include_public, default true,parsed with the existing boolean parameter helper so bad input returns 400.
visible_for_usertreats public rows as visible to everyone.is_publicread-only so the UI can label and lock public rows.model (for example "manage public taxa lists"), checked without reference to any project. To
begin with only superusers hold it. The check should be a named permission rather than a
hard-coded superuser test, so that platform roles can be granted it later without touching the
endpoints (see follow-up work). Project members can edit only objects scoped to a project they
belong to. This overlaps with Review M2M permissions for things that belong to multiple projects #1120 and should be designed with it.
project-scoped object for each entity.
Pipelines and processing services are the least symmetric today because availability also depends
on
ProjectPipelineConfig.enabled. A public processing service probably means "offered to everyproject, enabled per project", which is close to what the default processing service setting does
by adding each new project to one service's M2M. That may be better modelled as a public service
than as an ever-growing project list. This is a direction to discuss, not a requirement of this
ticket.
Order of work
is_publicfield, the shared queryset method, the visibility fix, theinclude_publicparameter, the platform-level permissions, and permission tests for bothTwo registration defects found along the way are tracked separately: empty category maps piling
up (#1426), and pipeline and algorithm key collisions between public and project-scoped services
(#1427).
Follow-up work
to one project; nothing sits between a project role and superuser. Two roles are anticipated: an
Antenna admin or manager role with cross-project access (public processing services, platform
defaults), and a Taxonomy Curator role that can curate public taxa lists and, in time, taxa
themselves. The platform-level permissions introduced here are what those roles would be granted.
Related: Refactor role-project relationship to use foreign key instead of string parsing #1100 (how roles relate to projects).
predicts species X most accurately?", and the same for a group such as the micro-moths or the
Crambidae, and to show how many training images stand behind each class (Show the species-wise accuracy in the machine prediction component #788). The published
figures belong in the category map, since they are facts the model's authors state about each
class and the map is the service's verbatim copy; category entries already accept extra keys.
The through-model on list membership then carries them next to the taxon, together with the
original label and class index, which is what makes them comparable across models and
summable up the taxonomy. List membership alone loses the label and index, so a label that
resolved to a differently named taxon through a synonym cannot be explained to the user, and
class masking has to re-derive the index. No model ships these figures today; agree the keys
with the processing-service schema first, and keep the through-model in mind when writing the
sync so it can be added without reworking it.
shared taxonomy lacks: an undescribed or provisional species, a morphospecies, a regional name
not yet imported. Members should be able to create such a taxon inside their project, add it to
their lists, and pick it when identifying. This is the project-scoped taxon, and it is the case
that gives
is_publica purpose on taxa: shared reference taxa are public, a member's additionsare scoped to their project until a Taxonomy Curator promotes or merges them. Two things stand in
the way today. Taxon names and display names are unique across the whole platform, so two
projects could not both have a "Morphospecies 1"; uniqueness would need to hold among public
taxa and within each project instead. And nothing outside the admin creates a taxon ([integration] OOD Features #820 had a
first version of a create-species form). The identification search must return public taxa plus
the project's own, and nothing from other projects.
with (it exists in the taxonomy or the project's lists), predictable by a model (it is in the
category map of a classifier the project runs), both, or, for a member's own addition, available
only. Wherever taxa are listed or chosen (list pages, the taxon page, the identification picker)
the distinction should be shown rather than implied, for example "predicted by 2 models" against
"identification by people only". The recognised-by lookup in the core work is what feeds this. It
also sets expectations: a species missing from the project's results may be absent from the
site, or simply not something the model can output.
and the identification picker can all be narrowed to a project's region without each feature
fetching it again. The shape to mirror is Darwin Core's species distribution record (a taxon, a
location identifier, occurrence status, establishment means, source) with GADM identifiers as
the location vocabulary, since GADM's nested country, state and district codes match how
deployments are already placed and how Generate a project taxa list from a region, and track which species the models can predict #1367 asks GBIF for a region's species. A regional list
then becomes a saved query over this data rather than a one-off fetch. Design only for now;
the exact Darwin Core terms and GADM version need checking against the published standards.
lists two independent answers:
is_publicsays who may see a row, and the managed flag says whoowns its contents (a list an algorithm points at mirrors a category map and refuses hand edits).
Pipelines and algorithms arrive from what a processing service advertises and are then edited in
the admin with nothing recording that the next registration may overwrite the edit; the default
processing service comes from settings rather than from a person; taxa created from classifier
labels are awaiting a curator and nothing says so. Declaring both flags on the shared base, with
the same read-only serializer fields, would let a client label and lock such rows the same way
everywhere. Belongs with Review M2M permissions for things that belong to multiple projects #1120 and the platform roles above.
run by an operator (
import_taxa,update_taxa, from Update command for importing taxa from external lists #939). The aim is a general importframework that mirrors the export framework: a registry of import formats, a job that runs the
import in the background and reports what it created, matched and rejected, and a page to start
one. Taxa and taxa lists would be the first formats, built on the existing commands' logic;
feat: dataset import framework #1254 proposes the framework for datasets. The taxa list CSV export (feat(exports): add taxa_list_csv export format #1293) defines the natural
round-trip format. Imports into public taxa or public lists need the curator permission
described above.
given model can recognise, and later how accurately.
Decisions
Settled:
is_publicfield, not inferred from a row having no projects. Inferringit carries a silent risk: deleting a project removes its M2M rows, so a private list whose only
project is deleted would become public to everyone. Processing services take the same approach,
in the same first step, because a managed list's visibility follows its algorithm's service.
include_publicdefaults to true, which matches how tags behave today and makes classifierlists discoverable.
relationship is taxon → taxa list → algorithm; the category map is the service's verbatim,
unchanging copy. Draft Generate a project taxa list from a region, and track which species the models can predict #1367's direct relationship and coverage flag are dropped in favour of
it. Membership is fixed, and removing a managed list never deletes taxa.
algorithms and the active project's own, and never one private to another project.
Still open:
or both.
something friendlier such as " species list".
is_public. Nearly every taxon has no project today and taxa are treated asshared reference data, so the field changes nothing until members can add their own taxa (see
follow-up work). It can wait for that work, but the taxa-list rule should be written so taxa
can adopt it unchanged.
What we still need to verify
What makes an algorithm private. Algorithms have no link to projects and no visibility of
their own; they are reached through pipelines and processing services. Algorithm keys and
pipeline slugs are unique across the whole platform, and registration matches on them, so a
project's own service that advertises an existing key attaches to the same shared algorithm
row, and whichever service registers a key first decides its category map. Proposed working
definition: an algorithm is public if at least one public processing service offers it, and
otherwise visible only to the projects of the services that do. Check whether a project-scoped
service can claim or alter a key that a public service uses, and whether keys need a namespace.
Confirm that a category map never changes after it is created. Registration appears to create a
map once per algorithm and never update it, which is what makes fixed list membership sound. The
admin can still edit a map's labels, and the stored labels hash is only computed when empty, so
an edit would leave both the hash and the list stale. Decide whether to make the labels read-only
in the admin.
Empty category maps pile up at registration. Measured in one development database: 1,250
category maps, of which 1,232 have no algorithm, and every one of those 1,232 is empty (no
labels, no description, not referenced by any classification). The 18 maps in use all have
distinct label sets, so identical label sets being duplicated was not observed. The apparent
cause, from reading the registration code: an algorithm whose processing service reports a
category map with no entries (a detector, in the case observed) fails the "has a valid category
map" check every time, so each registration creates a fresh empty map, points the algorithm at
it and orphans the previous one. This deserves its own small fix (treat an empty map as no map,
and clean up the orphans) and needs a count on a production database. For this ticket it means
lists are created only for maps that have labels and an algorithm. Separately, registration
never reuses an existing map with the same labels, so two algorithms with identical classes
would get two lists; the stored labels hash would allow reuse.
Check what creating missing taxa does for labels that are not taxa (for example the two classes
of a binary moth / non-moth classifier), so the sync does not add junk rows to the shared
taxonomy.
The
visible_for_userbehaviour above is measured at the queryset level only. Add an API-leveltest (with and without a project id) so the endpoint behaviour is pinned as well.
Check whether the one processing service with no projects is an orphan or intentional, and review
the taxa lists with no project before the backfill marks them public.
Check what happens to a project-scoped taxa list when its only project is deleted.
Measure the list endpoint with public lists included on the largest project, with
EXPLAIN (ANALYZE): the taxa-count annotation over lists of a few thousand taxa, and theOR projects IS NULLcondition, which can defeat an index the way other OR filters in thiscodebase have.
Confirm how many category-map labels fail to resolve to a taxon per classifier (dry run of the
management command), since unresolved labels make the public list an incomplete answer.
Related issues and pull requests