Skip to content

Give every detection a BioCLIP feature vector, and embed existing detections without re-detecting - #175

Draft
mihow wants to merge 45 commits into
mainfrom
feat/bioclip-features-all-detections
Draft

mihow wants to merge 45 commits into
mainfrom
feat/bioclip-features-all-detections

Conversation

@mihow

@mihow mihow commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Tracking an insect from one capture to the next works best when every detection carries a feature vector from the same model, whether or not the moth/non-moth filter kept it. This PR lets the processing service give every detection a BioCLIP 2.5 embedding, and adds a feature-only pipeline that embeds detections a platform already has, without running the detector again or adding any classification. That second part matters for backfilling: existing detections, including ones people have already reviewed, keep their boxes and gain a vector.

It builds on the per-classification feature vectors in #77 and the "vector for every detection" work tracked in #170, which are merged into this branch, and it reuses the BioCLIP 2.5 ViT-H/14 backbone loading from #174 without its classification heads. The diff against main therefore includes #77. This is a draft until those land and the platform side can store 1024-float vectors.

List of Changes

Change (effect) How (implementation) Notes
1. Any pipeline can attach a BioCLIP embedding to every detection, including those the moth filter rejected, with no extra classification. New request config embedding_extractor (and service setting AMI_EMBEDDING_EXTRACTOR) naming bioclip_2_5_embeddings; vectors go in detections[].embeddings. A key the pipeline does not list in /info (unknown, or not the one in the setting) returns 422, so a response never names an unregistered algorithm; an invalid setting stops the service at startup.
2. Existing detections can be embedded without re-detecting or re-classifying. New pipeline bioclip_2_5_features. When the request carries detections with boxes, the detector is skipped and the same boxes and detector references come back with an embedding each. Without detections, it detects first. Detections without a box, or with a box of no area, are skipped.
3. A caller can register the new algorithm and pipeline from /info. The extractor is described from its class (task type embedding, no category map) without loading the backbone; it is listed in every pipeline when the setting is on, and always in the feature-only pipeline.
4. The backbone loads once, not per request. The open_clip model is cached per process and device. About 4.7 GB of GPU memory with the detector also loaded.
5. The async worker is unaffected. The worker and ami worker register skip feature-only pipelines, which only the API serves. ami worker register never advertises the embedding extractor, even with AMI_EMBEDDING_EXTRACTOR set, because the worker does not run it.
6. Tests never download the backbone. The encoder is replaced by a small deterministic module; the detector and classifiers are real. trapdata/api/tests/test_embedding_extractor.py

Detailed Description

The contract the platform relies on

Antenna's side of this (RolnickLab/antenna#1439) reads the response under these rules, so they should not change without a matching change there:

  1. Task type. The extractor is described in /info with task_type: "embedding". Antenna treats a pipeline whose algorithms are only embedding extractors and detectors as feature-only, and sends it existing boxes instead of images to detect on.
  2. The features key. Each vector is sent as detections[].embeddings[].features, next to the extractor's algorithm reference. Antenna reads only features; no other key is accepted.
  3. Echoed boxes. A feature-only response returns exactly the boxes it was sent, with the same coordinates and the same detector algorithm reference. Antenna matches each returned box to a stored detection by image and box (rounded to three decimals); a box that matches nothing is counted and logged, and a batch in which no box matches fails the job.
  4. /info lists both algorithms. The feature-only pipeline lists the detector whose boxes it echoes and the embedding extractor. Antenna stores a vector only under an algorithm key the pipeline declared in /info, and refuses a key it did not.
  5. One length per extractor. Every vector from one extractor has the same length (1024 for BioCLIP 2.5). Antenna records the length from the first batch and refuses vectors of another length from the same extractor.

Request (feature-only pipeline, embedding existing detections):

{
  "pipeline": "bioclip_2_5_features",
  "source_images": [{"id": "123", "url": "https://example.org/capture.jpg"}],
  "detections": [
    {
      "source_image": {"id": "123", "url": "https://example.org/capture.jpg"},
      "bbox": {"x1": 10, "y1": 20, "x2": 110, "y2": 140},
      "algorithm": {"name": "FasterRCNN for AMI Moth Traps 2023", "key": "fasterrcnn_for_ami_moth_traps_2023"}
    }
  ]
}

Each response detection keeps its box and algorithm, has classifications: [], and one embedding:

"embeddings": [
  {"features": [0.0123, -0.0456, "... 1024 floats, L2-normalised"],
   "algorithm": {"name": "BioCLIP 2.5 ViT-H/14 embeddings", "key": "bioclip_2_5_embeddings"}}
]

For any other pipeline, add "config": {"embedding_extractor": "bioclip_2_5_embeddings"}; this combines with features_for_all_detections, so a detection can carry both the species classifier's vector and the BioCLIP one, each under its own algorithm key.

Measurements

One RTX 3090, 32 synthetic 6080x3420 frames served over local HTTP, 8 images per request, global_moths_2024, 562 detections. Measured once, after a warm-up request.

Run Wall time Images/s Detections/s Peak GPU memory
Species vectors on every detection 110 s 0.29 5.1 11.3 GB
The same plus BioCLIP on every detection 125 s 0.26 4.5 11.1 GB
Feature-only, given the 562 detections 17 s 1.84 32.3 4.7 GB
Feature-only, detector runs 26 s 1.24 21.8 11.1 GB

The feature-only run returned the same 562 boxes, no classifications, and vectors identical (cosine 1.0) to the ones from the full pipeline. Most of the full pipeline's time is the large-vocabulary species classifier's post-processing, not BioCLIP.

End to end with the platform

Run against a local copy of a development database with the Antenna side of this contract (RolnickLab/antenna#1439): three feature-only jobs over three one-hour sessions stored one 1024-float vector for each of 12,168 existing detections, and the counts of detections, classifications and occurrences, and every determination, were unchanged.

What is not verified yet

  • The async worker path, which ignores embedding_extractor.
  • Throughput with num_workers > 0; crops are read in the main process here, as in the other API pipelines.

How to run locally

uv sync
AMI_EMBEDDING_EXTRACTOR=bioclip_2_5_embeddings uv run ami api --port 2010
uv run pytest trapdata/api/tests/test_embedding_extractor.py

The first request downloads the backbone (about 3.6 GB) from the Hugging Face Hub.

🤖 Generated with Claude Code

https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8

mohamedelabbas1996 and others added 30 commits April 13, 2025 21:00
…tion foundation

Merge main into feat/add-classification-features-to-response. Conflicts in
pyproject.toml, poetry.lock, and api/models/classification.py resolved by
taking main's version. Mohamed's get_features() and features schema field
came through auto-merge and will be refined in subsequent commits.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds include_features and include_logits flags to PipelineConfigRequest (API)
and Settings (worker). Adds features field to ClassificationResponse. Makes
logits field conditional (default None). Both default to off for backward
compatibility and reduced response size.
APIMothClassifier now accepts include_features and include_logits flags.
When enabled, predict_batch() extracts features via get_features() and
post_process_batch() conditionally includes logits. Both flow through
ClassifierResult → update_detection_classification() → ClassificationResponse.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
API endpoint passes both flags from PipelineConfigRequest to classifier.
Worker passes both from Settings (AMI_INCLUDE_FEATURES, AMI_INCLUDE_LOGITS env
vars) to classifier constructor. No changes needed to _process_batch() since
the predict_batch()/post_process_batch() overrides handle the flow.
Tests that features are 2048-dim when enabled, logits present when enabled,
both absent when disabled (default), and both present when both flags set.
Replaces Mohamed's original tests with opt-in config pattern.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Resolve conflicts in worker.py by taking origin/main's version (from PR #122)
and re-applying our include_features/include_logits classifier constructor change.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- predict_batch() now stores features in self._last_features instead of
  returning a tuple, preserving compatibility with base class run() which
  uses len(batch_output) for timing calculation
- Existing tests that assert logits are present now pass include_logits=True
  since logits are opt-in (default off)
- Use self.assertEqual for status code assertion in test helper

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Worker path test verifies features flow through predict_batch/post_process_batch
- Validity test checks features are non-zero, have variance, and differ between
  detections (not just checking existence and dimension)
- Clear self._last_features after post_process_batch reads it to free
  GPU memory between batches
- Add include_features and include_logits to Kivy settings fields dict
  for discoverability in the desktop app UI

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Feature extraction ran the backbone twice: once through the model for the
logits, then again through forward_features for the embedding. Measured with
call counters on timm's resnet50, the old path invoked forward_features twice
per batch where one pass is enough, doubling classifier cost in exactly the
configuration that wants embeddings.

Replaces the get_features(input) hook with forward_with_features(input), which
returns the logits and the features from a single set of feature maps. The
split is exact: logits match a plain forward pass with a maximum absolute
difference of 0.0, and timm's own pre_logits pooling matches the manual
adaptive average pool the hook used before.

Also notes, on each of the two feature extractors in the codebase, that the
other exists and how they differ, as asked for in review.

Co-Authored-By: Claude <noreply@anthropic.com>
Classification responses have always carried logits. Putting them behind a new
include_logits flag that defaults to false would have silently stopped that for
every consumer that did not know to ask, including Antenna's class masking,
which re-scores classifications from the stored logits and skips any row where
they are null. Nothing about adding feature vectors requires taking logits
away, so the flag now defaults to on and only turns them off when a caller
asks.

The flag also now reaches the binary moth/non-moth filter, in both the HTTP API
and the worker. It was only ever passed to the terminal classifier, so
include_logits=false still returned logits on non-moth detections. Feature
vectors are deliberately not passed to the binary filter: its model has no
backbone hook and could only return nothing.

Co-Authored-By: Claude <noreply@anthropic.com>
Four small defects in how the classifier passes features from predict_batch to
post_process_batch:

- _last_features was only ever created inside predict_batch, so calling
  post_process_batch first raised AttributeError. It is now initialised in the
  constructor.
- predict_batch ran under no_grad only on the branch that extracted features,
  because the decorator sat on the extraction hook. The worker calls
  predict_batch directly, so the decorator now sits on the method itself and
  covers both branches.
- Asking for features from a model that has no backbone hook returned nothing
  and said nothing. Six of the ten pipelines are in that position. A new
  supports_features() classmethod reports it, and the constructor warns.
- Logits were copied to the CPU on every batch even when they were about to be
  discarded.

Adds the type hints both overrides were missing, and makes include_features and
include_logits keyword-only so they cannot be bound by a stray positional
argument.

Co-Authored-By: Claude <noreply@anthropic.com>
The pipeline tests exercise the real thing but need model weights, so nothing
pinned the mechanics for a developer working offline or for a reviewer reading
the diff. Adds a set of tests that build a random-weight resnet50 and check the
parts that can silently break: that the backbone runs exactly once whether or
not features are requested, that splitting the forward pass leaves the logits
unchanged, that features survive the handoff between predict_batch and
post_process_batch and are released afterwards, that the defaults are logits on
and features off, and which pipelines report that they can extract features.

Removes a test whose name and docstring claimed it drove predict_batch and
post_process_batch directly when it only re-ran the HTTP pipeline, and drops
one that asserted the old logits default. Both are covered above.

The pipeline the API tests run against can now be set with AMI_TEST_PIPELINE,
so they can be pointed at whichever model is already cached locally. CI keeps
the default.

Co-Authored-By: Claude <noreply@anthropic.com>
The plan describes get_features() and a logits flag that defaults to false,
neither of which is how the code works now. Left in place for the history of
how the branch was brought up to date, with a note at the top saying so and
pointing at what changed, so it is not read as current behaviour.

Co-Authored-By: Claude <noreply@anthropic.com>
Model weights and label maps are fetched from a public object store that moved
to a new cluster. Buckets there are namespaced by the tenant that owns them, so
every URL needs both a new host and a "<tenant>:" prefix on the bucket name.
Until now the old host was written out in full 34 times across three modules,
so every download failed and no single place existed to fix it.

The base URL is now defined once in trapdata/common/constants.py as
OBJECT_STORE_BASE_URL, with MODEL_BASE_URL and IMAGE_BASE_URL derived from it.
The model modules interpolate MODEL_BASE_URL instead of repeating the host.

Verified by resolving all 32 distinct model URLs and requesting each one
anonymously: 30 return content. The two that do not are the UK Turing model
and its category map, which are absent from the bucket rather than moved --
the old and new buckets hold identical bytes, so those keys were already
missing before the migration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018PcbaarhFPYxhoW3o2vvqX
mihow and others added 7 commits September 14, 2026 10:14
This reverts commit 6ae6758.

The stopgap that pointed model downloads at the new object store through a
constant in trapdata/common/constants.py has been superseded on main by
#168 and #169, where the download location is a setting and model classes
name their files relative to it. Reverting it here lets main merge in
cleanly without reintroducing the constant.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012iwV2sD1vVtL5xEUN8Rgf3
…on-features-to-response

# Conflicts:
#	trapdata/api/tests/test_api.py
#	trapdata/settings.py
Brings in the open feature-vector branch so every detection's embedding can reuse its schema and settings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Tracking and similarity search in Antenna compare feature vectors, but crops
the moth/non-moth filter rejects never reach the species classifier, so they
carry none. With AMI_FEATURES_FOR_ALL_DETECTIONS on (or the request config's
features_for_all_detections), the species classifier also runs over the
rejected crops, and every detection gets one vector in a new per-detection
`embeddings` field, attributed to the species classifier's algorithm key. That
is the same model and key as the vectors on moth crops, so all of them are
comparable.

The vector travels on the detection rather than as an extra classification.
An extra species classification on a rejected crop would compete with the
filter's label for the occurrence's determination in Antenna, even when marked
non-terminal. Rejected crops keep exactly the classifications they had before.

The setting is off by default, so existing deployments get the same
classifications as before; the only difference in the JSON is
`"embeddings": null` on each detection. When it is on, each rejected crop costs
one more backbone pass, batched with the other rejected crops. The per-class
score post-processing is skipped for those crops, since on large-vocabulary
models it costs far more than the backbone pass itself.

Refs #170 and RolnickLab/antenna#1417

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
BioCLIP 2.5 is published as an open_clip model, so the feature extractor loads it through that library.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
Tracking needs one comparable vector per detection. This extractor wraps the frozen BioCLIP 2.5 ViT-H/14 image encoder from open_clip, without any classification head, and attaches a 1024-float L2-normalised embedding to each detection it is given. The backbone is loaded once per process and device and shared across requests, because it is about 2.5 GB. The species classifier reuses the same helper to attach its own vectors.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…tions without re-detecting

Any pipeline can now attach a BioCLIP embedding to every detection, including those the moth/non-moth filter rejected, through the embedding_extractor request config or the AMI_EMBEDDING_EXTRACTOR setting. The embedding adds no classification. When the setting is on, /info lists the extractor in every pipeline with task type "embedding" so a caller can register it.

A new feature-only pipeline, bioclip_2_5_features, embeds the detections sent with the request and returns the same boxes and detector references with an embedding and no classifications. Without detections it runs the detector first. The Antenna worker skips feature-only pipelines, since only the API serves them.

Tests replace the backbone with a small deterministic encoder, so they never download it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

mihow and others added 6 commits September 28, 2026 12:42
…orker

The AMI_EMBEDDING_EXTRACTOR setting added the extractor to every classifier
pipeline's description, and worker registration builds its pipelines from the
same description. A worker host with the setting on therefore registered an
embedding algorithm that the worker never runs, so Antenna would record an
algorithm that never produces output and could treat every image as
unprocessed. Registration now describes its pipelines without the extractor,
and a test checks that no embedding algorithm is registered while the setting
is on.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
…vertise

A request could ask for an embedding extractor that the service's /info did
not list for that pipeline, because the AMI_EMBEDDING_EXTRACTOR setting was
off. The response then named an algorithm that Antenna had never registered,
and Antenna rejects such a result partway through saving it. A classifier
pipeline now returns HTTP 422 for any extractor other than the one it
advertises, and the request field's description states the rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
A detection sent to a feature-only pipeline with a zero-width, zero-height or
inverted box could not be cropped, and the crop error failed the whole
request with HTTP 500. Such detections are now left out, as detections
without a box already are, and a test covers both shapes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The embeddings field description said vectors appear only when
features_for_all_detections is on. The embedding_extractor option and the
feature-only pipelines fill it too, and vectors from different algorithms
cannot be compared, so the published description now says both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The fake pipeline descriptions in the pipeline selection tests took only the
classifier and slug, so they failed once service info started passing
include_embedding_extractor through.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
mihow and others added 2 commits September 29, 2026 07:54
The worker's pipeline filter kept a slug only when it was in
CLASSIFIER_CHOICES. A pipeline that is neither there nor feature-only,
such as one built on a different detector, would have been silently left
out of the worker's subscription and its registration with Antenna. The
filter now removes exactly the feature-only pipelines and keeps the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
The map of every pipeline slug, classifier and feature-only alike, was
called PIPELINE_CHOICES. The staged pipeline registry proposed for the
general insect detector uses that same name for a map of stage
definitions, so the two branches would have defined one name with two
meanings. The combined map now has its own name, and the tests that
describe a pipeline look it up through select_pipelines rather than
assuming the registry's value type.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C7Xf6VPbwWtTumhjjF15g8
@mihow

mihow commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Claude says: I checked this PR against #160, the general insect detector PR, so that the two can land in either order. In short, they are compatible. Two small changes were needed here, and one integration step will be needed whichever PR merges second.

Pushed to this branch

The full suite passes: 228 passed, 1 skipped.

What a merge involves

I merged #160 into this branch locally in a throwaway worktree and resolved it. Most of the conflicts come from #160 predating #171 and #172 on main; five of the eight conflicting files also conflict between #160 and main alone. The part that needs thought is that #160 replaces the inline filter-then-classify code in both /process and the worker with one shared stage engine, and this PR's "embed the detections the filter rejected" logic lives in that inline code. It moves into the engine cleanly:

  1. The stage runner passes include_features, include_logits and include_embeddings to the terminal classifier, and only include_logits to the gates.
  2. After the terminal stage runs, the engine embeds whatever the gates dropped using the terminal classifier. /process calls classifier.embed(). The worker uses a tensor-based twin of _classify_detections_from_tensors, which also picks up Let a pipeline pair any detector with any classifier, and add a general insect-detection pipeline #160's clamp_to_bounds in place of this branch's unclamped _transform_crops.
  3. make_pipeline_config_response takes a PipelineDefinition (or a feature extractor for the feature-only pipeline) and appends the embedding extractor after the stage loop.
  4. Two of Let a pipeline pair any detector with any classifier, and add a general insect-detection pipeline #160's tests need adjusting. The advertised/subscribed/dispatchable test has to exclude feature-only pipelines, and the score_scale test's mock needs _embedding_only=False.

On the merged tree, with real models, the anybug pipeline lists exactly one detector plus the BioCLIP extractor in /info. Worker registration still leaves the extractor out. On a test frame it returned 43 detections, 8 of which the Lepidoptera gate rejected, and all 43 carried both vectors. The rejected ones carried only their order label. The merged suite passes apart from three tests that fail on #171's model URL rules, which is #160's rebase and not this PR.

Detection and classification schemas do not collide. #160 adds no response fields, only a clamp helper on BoundingBox. The feature-only pipeline already echoes a detection's detector reference unchanged whatever its key; test_embeds_the_given_boxes_without_detecting_or_classifying uses a non-default key.

Suggestions for #160, not changed there

  • Rebase onto main first, which covers the model URL rules and select_pipelines.
  • When the registry moves out of api.py, as requested in review, give the stage engine per-stage options and a hook for "detections a gate dropped". Then this PR plugs in without adding inline code back. Feature-only pipelines could also get a place in the registry type, so there is a single typed map.
  • One thing worth checking on the platform side before running the feature-only pipeline on detections from another detector: its /info advertises the default moth detector, because that is what runs when no detections are sent.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants