You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Declare each model's vector length and add nearest-neighbour indexes when exact search gets slow #1500
#1462 stores feature vectors for detections and compares them with exact scans: every candidate vector is read and its distance computed. That is fast enough for a filtered list today. When it is not, approximate nearest-neighbour search (an HNSW index from pgvector) is the next step, and it needs one thing the table does not have yet: a declared vector length for each model output. This ticket outlines both, so they are designed together.
Why they belong together
The vector column has no declared length, so models with different lengths share one table (a 2,048-value classifier backbone and a 1,024-value BioCLIP model, for example). The writer keeps each (algorithm, key) at one length by reading the length of an existing row before each batch. That is enough for exact scans, but:
An HNSW index has to be built on a fixed-length expression, (vector::halfvec(N)), one partial index per (algorithm, key). Building it safely needs the length recorded somewhere, not inferred from whichever row comes first.
The writer's first-write check relies on an advisory lock. A declared length under a unique constraint would make the rule explicit and enforced by the database.
Proposed direction
A small record per (algorithm, key), for example EmbeddingSpace(algorithm, key, dimensions, created_at), unique on (algorithm, key). The writer creates it with get_or_create on the first batch, so a second concurrent writer collides and re-reads instead of needing a lock, and every later batch is checked against dimensions. The length belongs here rather than on Algorithm, because one algorithm can return several outputs.
Optional HNSW index per record, created by an admin action or a management command rather than a migration, because building one is slow and large. The record can note whether an index exists and with which parameters.
Query path: when an index exists for the pair, the similarity query orders by vector::halfvec(N) <=> seed inside one (algorithm, key), with hnsw.ef_search set for the query. Otherwise it falls back to the exact scan used today.
Measurements so far
All from an experiment on a throwaway database holding the real layout of two projects from a production copy, with synthetic clustered vectors, about 350,000 to 450,000 rows:
Exact scans: about 16 to 19 ms per 1,000 candidate rows with Postgres JIT off.
HNSW: 0.9 to 2.7 GB of index and 3.5 to 15 minutes to build per model, with top-20 recall between 0.68 and 0.98 depending on the model and parameters.
At today's sizes the index does not pay for itself, so nothing here is urgent.
What we still need to verify
Recall when the query also filters (by station, taxon, or score), which approximate indexes handle poorly. pgvector 0.8 has iterative index scans for this; they need measuring on real filtered queries.
Build time and memory on the largest production project, with maintenance_work_mem set the way production sets it.
Whether halfvec precision affects recall. On 600 real 2,048-value vectors, half precision changed cosine similarity by at most 1.5e-4.
Summary
#1462 stores feature vectors for detections and compares them with exact scans: every candidate vector is read and its distance computed. That is fast enough for a filtered list today. When it is not, approximate nearest-neighbour search (an HNSW index from pgvector) is the next step, and it needs one thing the table does not have yet: a declared vector length for each model output. This ticket outlines both, so they are designed together.
Why they belong together
The vector column has no declared length, so models with different lengths share one table (a 2,048-value classifier backbone and a 1,024-value BioCLIP model, for example). The writer keeps each (algorithm, key) at one length by reading the length of an existing row before each batch. That is enough for exact scans, but:
(vector::halfvec(N)), one partial index per (algorithm, key). Building it safely needs the length recorded somewhere, not inferred from whichever row comes first.Proposed direction
EmbeddingSpace(algorithm, key, dimensions, created_at), unique on (algorithm, key). The writer creates it withget_or_createon the first batch, so a second concurrent writer collides and re-reads instead of needing a lock, and every later batch is checked againstdimensions. The length belongs here rather than onAlgorithm, because one algorithm can return several outputs.vector::halfvec(N) <=> seedinside one (algorithm, key), withhnsw.ef_searchset for the query. Otherwise it falls back to the exact scan used today.Measurements so far
All from an experiment on a throwaway database holding the real layout of two projects from a production copy, with synthetic clustered vectors, about 350,000 to 450,000 rows:
At today's sizes the index does not pay for itself, so nothing here is urgent.
What we still need to verify
maintenance_work_memset the way production sets it.Related
docs/claude/reference/feature-vectors.md.