Skip to content

Classifier species lists as taxa lists and support for public lists #1424

Description

@mihow

Where this fits

The larger goal is that a user is only shown results that are accurate and useful for the place
they monitor (#890). Most classifiers in Antenna cover far more species than occur at any one site,
so the platform needs a way to say "these are the species that matter here" and to apply that to
predictions. Taxa lists are that mechanism, and the work has been arriving in stages: taxa lists
and their editor (#1094, #1104), filtering taxa and occurrences by a list (#745, #1008, #1347),
class masking against a list (#999) on the post-processing framework (#1289), and verifying
presence species by species (#1320, #1365). Still ahead are building a list from a region (#1364,
draft #1367), a user-facing way to turn masking on (#1406), and the earlier geo-fencing idea
(#797). Further out, members add their own taxa for identification, every taxon shows whether a
model can predict it or only a person can identify it, and taxa carry regional information.

This ticket supplies the piece those all lean on: an authoritative, browsable answer to "what can
each model predict", shared across projects, that a project can copy and make its own. It also
settles how shared and project-owned objects coexist, which taxa lists, processing services and
future curator roles all need. A full list of related issues and pull requests is at the end.

Summary

People who use Antenna regularly ask "can this model predict species X?" and today the platform
cannot answer, because the full set of classes a classifier knows (its category map) is not visible
anywhere a user can browse it. The closest thing, the "Taxa returned by " lists, only
holds species the model has already predicted somewhere, so it cannot show a "no".

This ticket proposes using the existing Taxa Lists feature to close that gap. Every classifier gets
one taxa list holding everything in its category map, public when the classifier itself is
available to every project. A project member can copy that list
into their own project, edit the copy (remove species that do not occur in their region, add
missing ones), and later choose it as the list for class masking. The first users are partners who
identify with a regional classifier and want a list for a neighbouring region derived from it.

Doing this well depends on a second piece of work: Antenna has several kinds of objects that can be
either shared across all projects or private to some projects (taxa lists, taxa, tags, pipelines,
processing services), and each one handles that split differently today. Public taxa lists already
exist in the database but appear to be unreachable through the API. The ticket therefore includes a
consistent rule for "public versus project-scoped" as related work.

Terminology

Use public for an object available to every project and project-scoped (or private) for an
object tied to specific projects. Avoid "global": in this domain it already means geographic scope
(the global moth classifier, global species lists), and the two meanings collide in exactly the
places this feature touches. Existing code and comments that say "global list" should be renamed as
they are touched.

Current state

Based on code reading unless marked as measured.

Entity Relation to project Public rows exist? Project-scoped list returns public rows? Filter parameter
TaxaList M2M projects Yes. get_or_create_for_project(project=None) creates them; used by taxa import, taxa update, and "Taxa returned by" lists No. TaxaListViewSet requires a project and filters projects=project None
Tag nullable FK project Yes Yes. project = X OR project IS NULL None
Taxon M2M projects Nearly all taxa have no project Partly. The project view is driven by occurrences and the default taxa filters, not the M2M; the M2M only feeds visible_for_user include_unobserved (different concept)
Pipeline M2M projects plus ProjectPipelineConfig None observed No. Requires the M2M link and an enabled config None
ProcessingService M2M projects (blank=True) One observed, likely an orphan No. Requires a project and filters projects=project projects filterset field
Algorithm none (through pipelines) n/a List scoped through enabled pipelines; detail unscoped None

Measured in one development database: 31 of 39 taxa lists have no project (mostly "Taxa returned
by" lists), 13,291 of 13,298 taxa have no project, 0 of 15 pipelines, 1 of 15 processing services,
and 4 of 4 tags.

Two behaviours that do not show up when reading the viewsets alone:

  1. The shared visible_for_user queryset method filters M2M models on projects__draft = false OR owner OR member. That is a join, so a row with no projects never matches. For any non-superuser
    a public TaxaList, Pipeline or ProcessingService is filtered out before the viewset's own
    project filter runs. Only Tag (nullable FK handled explicitly) and Taxon (own override) escape
    this. Measured in the same development database by calling the queryset method directly: of the
    31 taxa lists with no project, 0 are visible to an anonymous user and 0 to a non-superuser
    member; the one processing service with no project is likewise hidden. A regression test
    should pin this once the rule changes.
  2. The object-permission helper for M2M models returns no update or delete permission unless the
    object belongs to the active project. That is the right outcome for public objects, but it is
    incidental rather than designed, and the write path (IsProjectMemberOrReadOnly) checks project
    membership only, not whether the target object is public.

Earlier precedent: the integration branch in #820 made taxa and taxa lists without a project
visible inside every project, by filtering on projects = X OR projects IS NULL, the same shape Tag
uses on main. It also made both M2M fields blank=True and cleared the historical taxon-to-project
assignments, which it describes as mostly arbitrary. It left a TODO for the permission check on
project-scoped rows. On main those M2M fields are still required in forms, while
ProcessingService.projects is already blank=True. None of that branch's scoping has landed.

How taxa come into being from model output today, from reading the result-saving code. When a
processing service registers an algorithm, Antenna stores the category map and creates no taxa.
Taxa are created later, one at a time, the first time a label is the top prediction of a saved
classification: the name the service returned is looked up by name or search name, and if nothing
matches a new taxon is created with that name and the rank of the category at the top score's
index, with no parent. Labels that never come out on top never get a taxon. Nothing checks that the
returned name matches the category map's label at that index, or that the scores have the same
length as the map, so a name outside the category map is accepted silently and given the rank of
an unrelated category. Two workers saving the same new label at the same moment would both try to
create it, and the second would fail on the unique name; this has not been reproduced.

Work that exists already: a local branch adds AlgorithmCategoryMap.resolve_taxa(),
Algorithm.get_or_create_taxa_list() and a management command
create_taxa_lists_from_category_maps (with --dry-run), with tests. Labels resolve to taxa by
name or search_names, the same rule used when saving classifications. Not yet pushed.

Core work

  1. One taxa list per classifier, holding its full category map, public when the classifier
    is.
    Land the existing branch. Create or refresh the list when a processing service registers
    a category map, so new models need no operator step.

    Taxa at registration. Resolve every label to a taxon at registration rather than as predictions arrive, so
    the list is complete from the first day and saving results becomes a lookup that never writes
    to the shared taxonomy. Taxa created this way are public, since a model and its labels are
    shared by every project that runs it, and they are marked as created by a pipeline and not yet
    placed in the taxonomy, so a Taxonomy Curator has a queue to work through. A large classifier
    has tens of thousands of labels, so this runs as a bulk background step, not inside the
    registration request. Creation at save time stays as a fallback for services that never
    advertised a map.

    Predictions outside the category map. Treat the category map as the model's
    declared output and the score index as the source of truth. If the returned name disagrees
    with the label at that index, keep the classification, resolve the taxon from the index, and
    record a warning on the job. If the scores and the map differ in length, fail loudly, because
    that almost always means the service is running a newer model than the one registered and the
    stored map is stale. Neither case adds a member to the managed list.

  2. Make the taxa list the primary record of what a model can predict, with fixed membership.
    The relationship the rest of the platform uses is taxon → taxa list → algorithm: an algorithm
    points at its species list, and "which models know this taxon" is answered through list
    membership. The category map steps back to being the verbatim copy of what the processing
    service published: unchanging, kept for the order of classes that score vectors depend on, and
    rarely read elsewhere. Algorithms that share a category map share a list. A list that an
    algorithm points at is managed: its name and description can be edited, but taxa cannot be
    added or removed by anyone, including superusers and curators. The only writer of its
    membership is the sync from the category map, which is safe to re-run and picks up labels that
    resolve later (a new taxon, a synonym fix as in Resolve classifier labels and taxa-list matches through synonym_of, not name-only #1405). Anyone who wants a different set of
    species copies the list. The serializer says which algorithms a list describes, so the UI can
    hide the add and remove controls. Removing a managed list never deletes taxa, including taxa the
    sync created, and project copies are unaffected. Taxa lists already have a description field,
    shown in the list table; the sync should fill it for managed lists (which algorithms, how many
    labels, how many resolved) and the copy should record its source there as well as in a link.
    A managed list is only as public as its algorithm. Processing services and pipelines can
    be added by a user to just their own project, by design. The list for an algorithm that only
    such a service offers is scoped to that project and is not public. This ties the list's
    visibility to the processing service's, so the is_public field on processing services is
    needed for this step rather than later.

  3. Make public lists visible inside a project, and show the relationship. The project's Taxa
    Lists page shows the project's own lists plus public ones, clearly labelled and read-only. The
    algorithm page links to the classifier's species list in the Taxa views, next to the existing
    link to the category map API endpoint, and the list page names the algorithms it describes.

  4. Show which models can recognise a taxon. Start with filters in both directions: "show me
    all taxa this algorithm can predict" on the taxa views, and "show me all algorithms that can
    predict this taxon" on the algorithm list, with the taxon detail view listing them too. The
    answer includes public algorithms and the active project's own; an algorithm private to
    another project never appears and is never counted. With taxon → taxa list → algorithm as the
    primary relationship these are indexed lookups through list membership, the same cost as a
    direct taxon-to-algorithm table, rather than a search through every category map's JSON.
    Showing the algorithms as small labels on every row of the taxa list views is wanted but
    postponed: it needs one aggregate or prefetch for the whole page, a query-count test on a
    multi-row fixture and a query plan checked on the largest classifier, and the filters deliver
    most of the value first. This replaces the direct taxon-to-algorithm relationship and coverage
    flag in draft Generate a project taxa list from a region, and track which species the models can predict #1367, whose regional lists should read coverage through the managed lists
    instead.

  5. Copy a list into a project. A "Copy to project" action that creates a project-scoped list
    with the same members and records where it came from. Requires the project's
    create-taxa-list permission.

  6. Curate the copy. The existing add and remove endpoints cover single edits. Check whether
    bulk removal and search-to-add are good enough for a list of a few thousand species.

Related work: one rule for public versus project-scoped objects

Proposed rule, to apply to TaxaList and ProcessingService first (managed lists depend on both), then
Pipeline, Tag and Taxon:

  • An explicit is_public boolean field on the model. A public row is returned for every project,
    whatever its project links. A row that is not public and has no project is an orphan, visible to
    superusers only, never silently shared. The migration backfills is_public = true for existing
    rows that have no project, after checking them (see the verification list).
  • The projects M2M is blank=True so a public row needs no project link.
  • A shared queryset method, for example for_project(project, include_public=True), returns rows
    linked to the project plus public rows. Viewsets call it instead of hand-written
    filter(projects=project).
  • One query parameter with the same name and default everywhere: include_public, default true,
    parsed with the existing boolean parameter helper so bad input returns 400.
  • visible_for_user treats public rows as visible to everyone.
  • Each serializer exposes is_public read-only so the UI can label and lock public rows.
  • Creating, editing or deleting a public object requires a platform-level permission for that
    model (for example "manage public taxa lists"), checked without reference to any project. To
    begin with only superusers hold it. The check should be a named permission rather than a
    hard-coded superuser test, so that platform roles can be granted it later without touching the
    endpoints (see follow-up work). Project members can edit only objects scoped to a project they
    belong to. This overlaps with Review M2M permissions for things that belong to multiple projects #1120 and should be designed with it.
  • Permission matrix tests (member, non-member, anonymous, superuser) cover both a public and a
    project-scoped object for each entity.

Pipelines and processing services are the least symmetric today because availability also depends
on ProjectPipelineConfig.enabled. A public processing service probably means "offered to every
project, enabled per project", which is close to what the default processing service setting does
by adding each new project to one service's M2M. That may be better modelled as a public service
than as an ever-growing project list. This is a direction to discuss, not a requirement of this
ticket.

Order of work

Step Work Depends on
A (#1425) The public versus project-scoped rule on taxa lists and processing services: the is_public field, the shared queryset method, the visibility fix, the include_public parameter, the platform-level permissions, and permission tests for both none
B Algorithm points at its managed list; labels resolved and the list synced at registration as a background step; fixed membership; list visibility follows the algorithm; the score index is trusted when saving results A
C Copy a list into a project A, alongside B
D Which models recognise a taxon: filters in both directions and the taxon page; labels in list rows later B
E Interface: public label, locked managed lists, copy button, link from algorithm to list, algorithm labels on taxa the API shapes from A, B and D

Two registration defects found along the way are tracked separately: empty category maps piling
up (#1426), and pipeline and algorithm key collisions between public and project-scoped services
(#1427).

Follow-up work

  • Platform-level roles for people who are not superusers. Every role in Antenna today is scoped
    to one project; nothing sits between a project role and superuser. Two roles are anticipated: an
    Antenna admin or manager role with cross-project access (public processing services, platform
    defaults), and a Taxonomy Curator role that can curate public taxa lists and, in time, taxa
    themselves. The platform-level permissions introduced here are what those roles would be granted.
    Related: Refactor role-project relationship to use foreign key instead of string parsing #1100 (how roles relate to projects).
  • Choose a taxa list for class masking from the list page (ties into Give project owners a user-facing interface for class masking #1406).
  • Per-species accuracy and training data, per model. The aim is to answer "which algorithm
    predicts species X most accurately?", and the same for a group such as the micro-moths or the
    Crambidae, and to show how many training images stand behind each class (Show the species-wise accuracy in the machine prediction component #788). The published
    figures belong in the category map, since they are facts the model's authors state about each
    class and the map is the service's verbatim copy; category entries already accept extra keys.
    The through-model on list membership then carries them next to the taxon, together with the
    original label and class index, which is what makes them comparable across models and
    summable up the taxonomy. List membership alone loses the label and index, so a label that
    resolved to a differently named taxon through a synonym cannot be explained to the user, and
    class masking has to re-derive the index. No model ships these figures today; agree the keys
    with the processing-service schema first, and keep the through-model in mind when writing the
    sync so it can be added without reworking it.
  • Let project members add their own taxa, for identification. A project often needs names the
    shared taxonomy lacks: an undescribed or provisional species, a morphospecies, a regional name
    not yet imported. Members should be able to create such a taxon inside their project, add it to
    their lists, and pick it when identifying. This is the project-scoped taxon, and it is the case
    that gives is_public a purpose on taxa: shared reference taxa are public, a member's additions
    are scoped to their project until a Taxonomy Curator promotes or merges them. Two things stand in
    the way today. Taxon names and display names are unique across the whole platform, so two
    projects could not both have a "Morphospecies 1"; uniqueness would need to hold among public
    taxa and within each project instead. And nothing outside the admin creates a taxon ([integration] OOD Features #820 had a
    first version of a create-species form). The identification search must return public taxa plus
    the project's own, and nothing from other projects.
  • Make "available" and "predictable" visibly different. A taxon can be available to identify
    with (it exists in the taxonomy or the project's lists), predictable by a model (it is in the
    category map of a classifier the project runs), both, or, for a member's own addition, available
    only. Wherever taxa are listed or chosen (list pages, the taxon page, the identification picker)
    the distinction should be shown rather than implied, for example "predicted by 2 models" against
    "identification by people only". The recognised-by lookup in the core work is what feeds this. It
    also sets expectations: a species missing from the project's results may be absent from the
    site, or simply not something the model can output.
  • Regional information for taxa. Record where each taxon is known to occur, so lists, masking
    and the identification picker can all be narrowed to a project's region without each feature
    fetching it again. The shape to mirror is Darwin Core's species distribution record (a taxon, a
    location identifier, occurrence status, establishment means, source) with GADM identifiers as
    the location vocabulary, since GADM's nested country, state and district codes match how
    deployments are already placed and how Generate a project taxa list from a region, and track which species the models can predict #1367 asks GBIF for a region's species. A regional list
    then becomes a saved query over this data rather than a one-off fetch. Design only for now;
    the exact Darwin Core terms and GADM version need checking against the published standards.
  • Carry the public and managed flags across the other shared models. This ticket gives taxa
    lists two independent answers: is_public says who may see a row, and the managed flag says who
    owns its contents (a list an algorithm points at mirrors a category map and refuses hand edits).
    Pipelines and algorithms arrive from what a processing service advertises and are then edited in
    the admin with nothing recording that the next registration may overwrite the edit; the default
    processing service comes from settings rather than from a person; taxa created from classifier
    labels are awaiting a curator and nothing says so. Declaring both flags on the shared base, with
    the same read-only serializer fields, would let a client label and lock such rows the same way
    everywhere. Belongs with Review M2M permissions for things that belong to multiple projects #1120 and the platform roles above.
  • Import taxa and taxa lists from the UI. Today both arrive only through management commands
    run by an operator (import_taxa, update_taxa, from Update command for importing taxa from external lists #939). The aim is a general import
    framework that mirrors the export framework: a registry of import formats, a job that runs the
    import in the background and reports what it created, matched and rejected, and a page to start
    one. Taxa and taxa lists would be the first formats, built on the existing commands' logic;
    feat: dataset import framework #1254 proposes the framework for datasets. The taxa list CSV export (feat(exports): add taxa_list_csv export format #1293) defines the natural
    round-trip format. Imports into public taxa or public lists need the curator permission
    described above.
  • Decide the fate of the "Taxa returned by " lists: keep, rename, or hide.
  • Roll recognisability up the taxonomy: on a genus or family page, show how many of its species a
    given model can recognise, and later how accurately.
  • Resolve labels through synonyms (Resolve classifier labels and taxa-list matches through synonym_of, not name-only #1405) so copied lists do not carry unresolved names.
  • Seed a project list from a region and intersect it with a classifier's list (Let users build a project taxa list from a region, so class masking works out of the box #1364).
  • Rename remaining uses of "global" for cross-project objects in code and documentation.

Decisions

Settled:

  • "Public" is an explicit is_public field, not inferred from a row having no projects. Inferring
    it carries a silent risk: deleting a project removes its M2M rows, so a private list whose only
    project is deleted would become public to everyone. Processing services take the same approach,
    in the same first step, because a managed list's visibility follows its algorithm's service.
  • include_public defaults to true, which matches how tags behave today and makes classifier
    lists discoverable.
  • The managed taxa list is the one store of which models can predict a taxon. The primary
    relationship is taxon → taxa list → algorithm; the category map is the service's verbatim,
    unchanging copy. Draft Generate a project taxa list from a region, and track which species the models can predict #1367's direct relationship and coverage flag are dropped in favour of
    it. Membership is fixed, and removing a managed list never deletes taxa.
  • A managed list is only as public as its algorithm. "Known by N algorithms" counts public
    algorithms and the active project's own, and never one private to another project.
  • Recognised-by ships as filters in both directions first; per-row labels in list views come later.

Still open:

  1. Where public lists appear: in every project's Taxa Lists page, only via the algorithm page,
    or both.
  2. User-facing name of the per-classifier list: "Category map of " (current) or
    something friendlier such as " species list".
  3. When Taxon gets is_public. Nearly every taxon has no project today and taxa are treated as
    shared reference data, so the field changes nothing until members can add their own taxa (see
    follow-up work). It can wait for that work, but the taxa-list rule should be written so taxa
    can adopt it unchanged.

What we still need to verify

  • What makes an algorithm private. Algorithms have no link to projects and no visibility of
    their own; they are reached through pipelines and processing services. Algorithm keys and
    pipeline slugs are unique across the whole platform, and registration matches on them, so a
    project's own service that advertises an existing key attaches to the same shared algorithm
    row, and whichever service registers a key first decides its category map. Proposed working
    definition: an algorithm is public if at least one public processing service offers it, and
    otherwise visible only to the projects of the services that do. Check whether a project-scoped
    service can claim or alter a key that a public service uses, and whether keys need a namespace.

  • Confirm that a category map never changes after it is created. Registration appears to create a
    map once per algorithm and never update it, which is what makes fixed list membership sound. The
    admin can still edit a map's labels, and the stored labels hash is only computed when empty, so
    an edit would leave both the hash and the list stale. Decide whether to make the labels read-only
    in the admin.

  • Empty category maps pile up at registration. Measured in one development database: 1,250
    category maps, of which 1,232 have no algorithm, and every one of those 1,232 is empty (no
    labels, no description, not referenced by any classification). The 18 maps in use all have
    distinct label sets, so identical label sets being duplicated was not observed. The apparent
    cause, from reading the registration code: an algorithm whose processing service reports a
    category map with no entries (a detector, in the case observed) fails the "has a valid category
    map" check every time, so each registration creates a fresh empty map, points the algorithm at
    it and orphans the previous one. This deserves its own small fix (treat an empty map as no map,
    and clean up the orphans) and needs a count on a production database. For this ticket it means
    lists are created only for maps that have labels and an algorithm. Separately, registration
    never reuses an existing map with the same labels, so two algorithms with identical classes
    would get two lists; the stored labels hash would allow reuse.

  • Check what creating missing taxa does for labels that are not taxa (for example the two classes
    of a binary moth / non-moth classifier), so the sync does not add junk rows to the shared
    taxonomy.

  • The visible_for_user behaviour above is measured at the queryset level only. Add an API-level
    test (with and without a project id) so the endpoint behaviour is pinned as well.

  • Check whether the one processing service with no projects is an orphan or intentional, and review
    the taxa lists with no project before the backfill marks them public.

  • Check what happens to a project-scoped taxa list when its only project is deleted.

  • Measure the list endpoint with public lists included on the largest project, with
    EXPLAIN (ANALYZE): the taxa-count annotation over lists of a few thousand taxa, and the
    OR projects IS NULL condition, which can defeat an index the way other OR filters in this
    codebase have.

  • Confirm how many category-map labels fail to resolve to a taxon per classifier (dry run of the
    management command), since unresolved labels make the public list an incomplete answer.

Related issues and pull requests

Area Reference
Goal #890 results that are accurate and useful; #797 geo-fencing for regional species lists
Taxa lists, delivered #1094 configurable taxa lists; #1104 list endpoints follow the membership API pattern; #1119 review follow-ups; #1001 list name as page title; #745 and #1008 filter by list, and inverted; #927 taxa filters in the backend; #1347 keep list filters when opening a taxon's occurrences; #1320 and #1365 presence verification with an example occurrence
Taxa lists, open #1081 backend extensions, including "only admins can make a list shared by all projects"; #1293 taxa list CSV export; #1364, #1366 and draft #1367 lists from a region and which species models can predict
Algorithms and category maps #785 section for algorithms and category maps; #1368 filter occurrences by the algorithms that ran; #1405 resolve labels through synonyms; #788 species-wise accuracy
Class masking #999 masking to a species list; #1289 post-processing framework; #997 (closed) and #1406 user-facing masking; #1360 measuring the effect of post-processing
Public versus project-scoped, roles #820 integration branch with public taxa and lists; #1120 permissions for objects in several projects; #1100 how roles relate to projects
Import #939 import and update commands for taxa; #1254 dataset import framework; #1080 (closed) taxa management workflow; #871 (closed) reference images for taxa

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions