Skip to content

Stop every list endpoint from looking up permissions once per row #1430

Description

@mihow

Summary

Every list response in Antenna asks the database what the current user may do with each row, one row at a time, although the answer is the same for every row in the response. DefaultSerializer.get_permissions calls add_object_level_permissions, which calls BaseModel._get_object_perms, which resolves the row's project and then asks django-guardian for that user's permissions on it. Six serializers override that method; everything else inherits it.

#1428 removed this from the taxa lists endpoint, where it measured six queries per row: a ten-row page went from 37 queries to 8, and the page now costs the same whatever it holds. This ticket is about doing the same for the endpoints that still pay it, and about choosing a way to do it that does not need repeating by hand in every serializer.

Where it still happens

Measured since this ticket was opened, with a throwaway test as an ordinary project member (a superuser can short-circuit the permission path and hide the cost), and with the permission call patched out to separate its share from everything else on the page:

Endpoint Queries per row Of which the permission pattern Saving per page
Occurrence list 14 2 about 40 queries, and a reproducible 200–230 ms on a project with roughly 180,000 occurrences
Capture set list 3 3, the whole of it (patched, the page is flat at 9 queries however many rows) about 60 queries; no wall-clock win could be shown, see below
Occurrence identifications, nested in the occurrence list 1.5 per identification all of it another 60–90 queries per page at two or three identifications a row

The effective default page size for both endpoints is 20, not the nominal 10: SourceImageUploadViewSet sets default_limit = 20 on the shared paginator class at import time, so every viewset that does not set its own inherits it. That is worth a separate look.

On the capture set list the query win is the cleanest of the three and the wall-clock win could not be demonstrated: that endpoint takes about seven seconds on a large project, dominated by count annotations over very large tables, and sixty indexed permission lookups disappear into that. Fix it for the query count, and treat the seven seconds as a separate, larger problem.

The rest of this section was counted from reading the code rather than measured:

Endpoint What runs per row
Occurrence list (OccurrenceListSerializer.get_permissions) The row's project is not select_related, so reading it is one query, and guardian is then asked the same question again. An earlier investigation put this endpoint at roughly 5 + 12 queries per row, which makes it the most valuable of the three.
Occurrence identifications (OccurrenceIdentificationSerializer.get_permissions) Nested inside the occurrence list. Each identification walks back to its occurrence's project and runs a role check that is itself a query, so it repeats per row inside a repeat per row.
Capture set list (SourceImageCollectionSerializer.get_permissions) The direct-foreign-key version of what #1428 fixed: resolve the project, ask guardian, once per row.
The nested "minimal" serializer helper (MinimalNestedModelSerializer) It routes through DefaultSerializer.to_representation(), so every nested object appends user_permissions and pays the same lookup. Which parents embed it in a list response has not been traced.

Two serializers were checked and are not affected: the nested capture-taxon and classification serializers return an empty permission list without touching the database.

The part worth deciding first

#1428 fixed one endpoint by threading an already-resolved permission set into the many-to-many helper as an optional argument, and caching it on the serializer instance. That is correct there, but repeating it in four or five more serializers means the same six lines copied each time, and each copy is a chance to cache the wrong thing.

A direction worth considering instead: cache inside add_object_level_permissions / _get_object_perms themselves, keyed on the user and the project, on an object whose lifetime is one request. request.user is such an object — DRF builds it per request and the helper already receives it — so no call site would have to change at all, and every list endpoint would benefit at once. Points to settle before doing it that way:

  • Whether anything reuses a User instance across what are logically separate requests (a management command or a Celery task looping over users would keep a stale cache), and whether to key the cache on the request instead when one is available.
  • Whether permissions can change within a single request in a way that must be seen immediately, for example a view that grants a role and then serializes the affected rows.
  • Whether the optional argument added in Speed up the Taxa Lists page & add a reusable way to look up permissions once per page #1428 should then be removed again, to leave one mechanism rather than two.

A trap to guard while doing this

#1428 also made the membership check in add_m2m_object_permissions read projects from the prefetch cache when the caller populated one. That is correct only if the prefetch is the unfiltered relation. A Prefetch("projects", queryset=Project.objects.filter(...)) anywhere would silently narrow the check and deny permissions the user really has. The constraint is currently implicit; it should be stated where the helper reads the cache, and any new prefetch of that relation checked against it.

What to verify

  • Measure each endpoint before and after, as a project member who is not a superuser: superusers may short-circuit the permission path and hide the cost.
  • None of these endpoints has a query-count test today, so a regression would not be noticed. Each fix should land with one, using a multi-row fixture and a strict count, in the shape of the taxa-list test added in Speed up the Taxa Lists page & add a reusable way to look up permissions once per page #1428. A fixture whose row count does not actually change between the two requests cannot see per-row growth, which is how an existing test on the taxa endpoint has been missing it.
  • Check that permissions reported to the client are unchanged for: a project member, a member of another project, an anonymous user, and a superuser.

Depends on #1428.

Activity

  1. mihow commented on Oct 8, 2026

    @mihow
    CollaboratorAuthor

    Claude says: New evidence for one endpoint this ticket does not cover yet, the project list, plus the findings from #1475, which I'm closing as a duplicate of this one.

    The project list spans many projects, so a per-project cache would not help it

    The direction above caches permissions per request, keyed on the user and the project. That works for the occurrence, capture and capture set lists, where every row belongs to the same project. On /api/v2/projects/, every row is a different project, so that cache would miss on every row and the cost would stay the same.

    Measured in a throwaway test with cachalot disabled, 20 projects, ProjectListSerializer (plain list, no extra flags):

    Caller 5 rows 20 rows Per row Of which guardian queries
    Project member 25 queries 85 queries 4 all 4
    Signed out 35 queries 125 queries 6 4

    The two extra queries per row for signed-out callers appear to come from resolving guardian's anonymous user each time. That is a guess from the counts, not yet traced.

    Wall clock on a local copy of production data (34 to 36 visible projects, warm): about 0.2 s for a superuser, who short-circuits the permission path, and about 1.1 s for a signed-out visitor or an ordinary member. That 1.1 s is what the gallery on main already pays. It is more visible now because the projects map (#1485) requests up to 300 projects in one call. At today's project count the map costs the same as the gallery, and it would cost more once there are more than 40 projects.

    Recommendation for lists that span projects

    django-guardian 2.4.0, the version installed, provides ObjectPermissionChecker(user).prefetch_perms(objects). It loads the user's permissions for a whole set of objects in a couple of queries, and later checker.get_perms(obj) calls are answered from memory. Nothing in the codebase uses it yet.

    That fits alongside the direction above as a single mechanism: a per-request cache keyed on the project, which list views spanning several projects fill in one step with prefetch_perms for the page, and which single-project views fill on first use. The anonymous-user lookup is worth checking at the same time.

    Carried over from #1475

    Earlier attempt

    The branch fix/permission-n1-queries (November 2025, never opened as a PR) prototyped the single-project version: the capture list view fetched the project's permissions once and passed them through the serializer context into BaseModel._get_object_perms. That code did not reach main; #1428 later solved the taxa lists case another way. The branch is kept for reference.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions