Skip to content

feat: expose the encoder architecture, name and family string through the C API - #3

Merged
mudler merged 1 commit into
masterfrom
feat/encoder-family-accessor
Oct 5, 2026
Merged

mudler merged 1 commit into
masterfrom
feat/encoder-family-accessor

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown

What and why

parakeet.cpp records an encoder fingerprint in its speaker registry, so a registry made with one speaker encoder is refused when another encoder is used. Equal embedding sizes do not mean the same embedding space: ECAPA-TDNN and CAM++ both give 192 values. Its family string is built from the encoder GGUF header.

A host that binds libvoicedetect through the C API (for example the LocalAI voice-detect backend) cannot build that string today, because the context does not report the architecture or the model name. Voices that such a host enrolls from audio stay unfingerprinted. This PR exposes what the loader already reads from the header.

It is additive. Existing behavior and the ABI version (1) do not change.

API

const char* voicedetect_capi_encoder_arch(const voicedetect_ctx* ctx);    // voicedetect.arch, e.g. "ecapa_tdnn"
const char* voicedetect_capi_encoder_name(const voicedetect_ctx* ctx);    // general.name
const char* voicedetect_capi_encoder_family(const voicedetect_ctx* ctx);  // see below

C++: vd::Model::config().arch, .name and .family (VoiceDetectConfig gains name and family).

  • The pointers belong to the context, never change, and stay valid until voicedetect_capi_free. They are NULL for a NULL context. They can be read from several threads at once.
  • They work for a context from every load function (path, memory, prefixed memory). The values are read once at load time.
  • A prefixed load reads the prefixed keys (<prefix>general.name, <prefix>voicedetect.arch, ...), so a bundle component gives the same string as the standalone file it was made from. The keys of the bundle header do not leak into a component.

Family format

voicedetect:<voicedetect.arch>:<general.name>:<voicedetect.embedding_dim>

For example voicedetect:ecapa_tdnn:speechbrain/spkrec-ecapa-voxceleb:192.

  • Values are copied as they are, with no escaping.
  • A missing key gives an empty field and the colons stay.
  • The size is in decimal and is empty when it is 0 (an analyze model).
  • The string is empty when general.architecture is not voicedetect.
  • A general.name of the wrong type reads as empty. It does not fail a load that worked before.
  • Another quantization of the same encoder gives the same string.

This is the same formula as speaker_encoder_family in parakeet.cpp. The two must stay identical: the header and the README say so. Compared with that function (compiled from parakeet.cpp master) on the ECAPA-TDNN, CAM++, WeSpeaker ResNet34 and ERes2Net f32 files, the strings are identical:

voicedetect:ecapa_tdnn:speechbrain/spkrec-ecapa-voxceleb:192
voicedetect:campplus:models/3dspeaker_campplus_zh-cn_16k.onnx:192
voicedetect:wespeaker_resnet34:models/wespeaker_voxceleb_resnet34.onnx:256
voicedetect:eres2net:models/3dspeaker_eres2net_base_200k.onnx:512

Tests

  • test_encoder_family (no model files, runs in CI): exact format, equal strings from the path, memory and prefixed loaders, missing general.name, other or absent general.architecture, a wrongly typed name, colons and UTF-8 in the name, two components in one bundle that each report their own identity, no leak from the bundle header, NULL context, stable pointers, 8 threads reading at once.
  • test_encoder_family_models (needs VOICEDETECT_TEST_GGUF_LIST, a :-separated list of GGUF files; skips otherwise): for each file the three loaders give the same strings, equal to the formula applied to the header read straight from the file, and the last field equals voicedetect_capi_embedding_dim. Run with the four encoders above.

Verification

Linux x86_64, GCC, CPU:

  • CI command ctest -LE model: 5 passed, 1 skipped (test_capi_dim, needs a model), 0 failed.
  • With the four models: test_encoder_family_models passes. test_capi_dim and test_load_memory_embed pass too.
  • AddressSanitizer + UndefinedBehaviorSanitizer build: test_encoder_family, test_encoder_family_models (leak detection on), test_load_memory and test_load_memory_embed are clean.
  • check_baseline (Python reference check) does not run in my environment because the reference Python packages are missing. It does not touch this code.

Not run: macOS and Windows. The change uses only standard C++ and the already parsed header.

Limits

  • A model without general.name (a file not made by the converter) gives an empty name field, so two such encoders with the same arch and size have the same family. The converter always writes it.
  • The family names the embedding space by architecture and checkpoint name. It does not tell two fine-tunes with the same name apart. A registry that needs that also records a hash of the file, as parakeet.cpp does.

🤖 Generated with Claude Code

A store of enrolled voices has to know which encoder made its embeddings.
Equal embedding sizes do not mean the same embedding space: ECAPA-TDNN
and CAM++ both give 192 values. A host that binds libvoicedetect could
not tell the encoders apart, because the context did not report what the
loader had read from the GGUF header.

Add voicedetect_capi_encoder_arch (voicedetect.arch),
voicedetect_capi_encoder_name (general.name) and
voicedetect_capi_encoder_family, which returns
"voicedetect:<arch>:<name>:<embedding_dim>". The pointers belong to the
context and stay valid until voicedetect_capi_free; a NULL context gives
NULL. The values are read once at load time, so they work for the path,
memory and prefixed memory loaders. A prefixed load reads the prefixed
keys, so a bundle component gives the same string as the standalone file
it came from. The config gains name and family for C++ callers.

The string is defined the same way as speaker_encoder_family in
parakeet.cpp: a missing key gives an empty field, the embedding size is
empty when it is 0, and the string is empty when general.architecture is
not "voicedetect". The two must stay identical, and the header says so.
A wrongly typed general.name reads as empty and does not fail a load that
succeeded before.

This is additive: the ABI version stays 1.

test_encoder_family needs no model files: it checks the exact format,
equal strings from the three loaders, missing and odd keys, bundle header
keys that must not leak into a component, NULL and concurrent reads.
test_encoder_family_models (set VOICEDETECT_TEST_GGUF_LIST) checks the
ECAPA-TDNN, CAM++, WeSpeaker ResNet34 and ERes2Net files through the
three loaders against the formula applied to the file header.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit cf9e1d5 into master Oct 5, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants