Skip to content

Transition API urls to require a project prefix #978

Description

@mihow

Summary

We are relying more on projects to check object permissions, add nested data based on project settings (even to single-object detail views), and increase performance of large requests (partitioning based on project). However each view currently retrieves the "current" project in a few different ways and we are requiring the project_id url param on more and more views.

One standard practice is to add a parent object to the URL structure. So instead of

/api/occurrences/?project_id=5

We use

/api/project/6/occurrences/

Related to #734 & #928

Another option is to continue to use the project_id parameter, but make it required for all endpoints. But it's not as elegant!

Activity

  1. mihow commented on Sep 5, 2026

    @mihow
    CollaboratorAuthor

    Claude says: Adding a related finding from #1390 here, since this ticket is where the project-scoping approach across the API is being decided.

    Detection has no project column; it reaches a project only through source_image__project. When #1390 scoped the detections list to the requested project, the pagination COUNT got 10 to 40 times faster on a local copy of production (22 s to 1.8 s on the largest project, 0.5 to 0.6 s on smaller ones), but EXPLAIN (ANALYZE) shows the plan is a Parallel Seq Scan over the whole detection table followed by index probes into source images, on every project size tried. The cost is therefore bounded by the total size of the detection table rather than by the project, and it will grow platform-wide as data accumulates. Classification (scoped via detection__source_image__project) most likely has the same shape, though that one has not been measured.

    Directions to discuss, roughly by effort:

    1. Denormalise project onto Detection (and possibly Classification) with a composite index leading with it, matching Occurrence and SourceImage. The hard part is not the column but keeping it in sync: it has to be set on every write path (pipeline results, imports, crop regeneration) and updated whenever a capture or occurrence moves between projects, otherwise detections end up orphaned or attributed to the wrong project. A data-integrity check in the framework from Framework to check and report data integrity issues #1188 could pin the invariant (detection.project_id == detection.source_image.project_id).
    2. Estimated counts for large filtered lists (Estimated-count paginator for fast pagination on large filtered lists #1328), which sidesteps the COUNT without a schema change but does not help the list query itself.
    3. Leave it and accept the table-bound cost until the tables are partitioned by project, which the URL structure proposed in this ticket would make natural.

    Measurements are from a local copy of the production database with the query cache disabled; the plans and the table are in #1390's description.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions