Skip to content

Declare each model's vector length and add nearest-neighbour indexes when exact search gets slow #1500

Description

@mihow

Summary

#1462 stores feature vectors for detections and compares them with exact scans: every candidate vector is read and its distance computed. That is fast enough for a filtered list today. When it is not, approximate nearest-neighbour search (an HNSW index from pgvector) is the next step, and it needs one thing the table does not have yet: a declared vector length for each model output. This ticket outlines both, so they are designed together.

Why they belong together

The vector column has no declared length, so models with different lengths share one table (a 2,048-value classifier backbone and a 1,024-value BioCLIP model, for example). The writer keeps each (algorithm, key) at one length by reading the length of an existing row before each batch. That is enough for exact scans, but:

  • An HNSW index has to be built on a fixed-length expression, (vector::halfvec(N)), one partial index per (algorithm, key). Building it safely needs the length recorded somewhere, not inferred from whichever row comes first.
  • The retraining work in Retrain a classifier head from verified identifications, and score what it produces #1407 wants to validate incoming vectors against a known length.
  • The writer's first-write check relies on an advisory lock. A declared length under a unique constraint would make the rule explicit and enforced by the database.

Proposed direction

  1. A small record per (algorithm, key), for example EmbeddingSpace(algorithm, key, dimensions, created_at), unique on (algorithm, key). The writer creates it with get_or_create on the first batch, so a second concurrent writer collides and re-reads instead of needing a lock, and every later batch is checked against dimensions. The length belongs here rather than on Algorithm, because one algorithm can return several outputs.
  2. Optional HNSW index per record, created by an admin action or a management command rather than a migration, because building one is slow and large. The record can note whether an index exists and with which parameters.
  3. Query path: when an index exists for the pair, the similarity query orders by vector::halfvec(N) <=> seed inside one (algorithm, key), with hnsw.ef_search set for the query. Otherwise it falls back to the exact scan used today.

Measurements so far

All from an experiment on a throwaway database holding the real layout of two projects from a production copy, with synthetic clustered vectors, about 350,000 to 450,000 rows:

  • Exact scans: about 16 to 19 ms per 1,000 candidate rows with Postgres JIT off.
  • HNSW: 0.9 to 2.7 GB of index and 3.5 to 15 minutes to build per model, with top-20 recall between 0.68 and 0.98 depending on the model and parameters.

At today's sizes the index does not pay for itself, so nothing here is urgent.

What we still need to verify

  • Recall when the query also filters (by station, taxon, or score), which approximate indexes handle poorly. pgvector 0.8 has iterative index scans for this; they need measuring on real filtered queries.
  • Build time and memory on the largest production project, with maintenance_work_mem set the way production sets it.
  • Whether halfvec precision affects recall. On 600 real 2,048-value vectors, half precision changed cosine similarity by at most 1.5e-4.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions