Skip to content

feat: fingerprint the speaker encoder in the registry and refuse a family mismatch before naming - #90

Merged
mudler merged 3 commits into
masterfrom
feat/registry-encoder-fingerprint
Oct 4, 2026
Merged

mudler merged 3 commits into
masterfrom
feat/registry-encoder-fingerprint

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

Problem

Equal embedding dimensions do not mean the same embedding space. The speaker registry file stored only the dimension, so a swap of the encoder was accepted without a word. ECAPA and CAM++ both give 192 values: a registry enrolled with one and used with the other loaded fine, and names were assigned from the wrong space.

What the fingerprint is

The registry now records which encoder made its voices:

  • Family: voicedetect:<voicedetect.arch>:<general.name>:<dim>, read from the encoder GGUF header. It names the embedding space. A mismatch is a hard error.
  • Weights: sha256: of the encoder file. For a voice component of a bundle it is sha256: plus the source_sha256 recorded in the bundle header, the same identity the context reports. An encoder enrolled with the standalone file and later used from a bundle (unchanged copy) therefore gives the same weights identity and no warning. A mismatch is a soft warning.

PKSR v2 byte layout

Little-endian:

"PKSR", u32 2, i32 dim, u32 n, u32 family_len, family, u32 weights_len, weights, then n speaker records

A registry with no fingerprint is still written as v1, byte for byte. v1 files load unchanged. A reader that only knows v1 refuses a v2 file with "unsupported version".

Behaviour

Registry and encoder Result
Other embedding size Error
Other family Error naming both families, before any name is assigned
Same family, other weights Warning, names are still assigned
Registry has no fingerprint Warning; error in strict mode
Empty registry No check

Enrolment rules:

  • A registry without a fingerprint that already has voices is never stamped silently.
  • A fingerprinted registry refuses an embedding that comes without a fingerprint, and one from another family.
  • Behaviour change worth a release note: the existing parakeet_capi_speaker_registry_add_embedding now fails on a fingerprinted registry. Use the _fp variant below.

API and CLI

C-API, additive (ABI stays 10): parakeet_capi_speaker_registry_add_embedding_fp, parakeet_capi_speaker_registry_encoder_family, parakeet_capi_speaker_registry_encoder_weights, parakeet_capi_speaker_registry_set_strict, parakeet_capi_speaker_encoder_family, parakeet_capi_speaker_last_warning.

CLI:

  • enroll stamps the registry.
  • scene --strict-registry refuses a registry with no fingerprint.
  • parakeet-cli registry <file> prints the version, fingerprint and speakers. --restamp --encoder <gguf> stamps a v1 registry with the fingerprint of the encoder that made it.

Docs: docs/diarization.md and docs/speaker.md.

Tests and results

New or extended: test_speaker_registry, test_capi_speaker_registry, test_speaker_fingerprint, test_cli_registry.sh. Covered:

  • A real swap: ECAPA and CAM++ (both 192 values) is refused with both families named.
  • A simulated swap (same weights, renamed general.name) is refused.
  • Same family, other weights: warning, names assigned (also in strict mode).
  • v1 compatibility, no fingerprint, strict mode, --restamp.
  • The bundle case: CAM++ built into a bundle with scripts/bundle_gguf.py gives the same family and weights identity as the standalone file, and a registry enrolled with one is used with the other with no warning, both ways (PARAKEET_TEST_VD_BUNDLE).
  • test_capi_speaker checked "expects" in the identify size mismatch error; the message now reads "registry holds N-value embeddings, this model produces M" and the test follows.

Results after merging master (Linux x86-64 CPU, CED and voice-detect ON):

  • ctest -LE model: 37 of 37 passed (5 skipped for missing fixtures).
  • Model tests (ctest -L model, 87 tests; 57 skipped because their baseline files were not set): 5 not passing in the first run. test_relpos_attention_local_chunked, test_capi_timestamps and test_combined_offline fail, as they do without this change. test_tdt_beam and test_scene_stream timed out at 900 s with three jobs in parallel and passed when run alone (39 s and 53 s).
  • The speaker tests with real encoders (CAM++, ECAPA, WeSpeaker, diarization, CED): test_speaker_encoder, test_speaker_identify, test_capi_speaker, test_speaker_fingerprint, test_cli_registry, test_capi_diarize_named, test_sound_capi and test_scene_stream all pass.

Limits

  • No quantised voice-detect GGUF was tested. Whether the converter keeps general.name across quantisation is not known; if it changes, the family changes and the registry is refused.
  • The PARAKEET_WITH_VOICEDETECT=OFF build was not compiled.
  • GPU was not tested.
  • --restamp cannot tell two encoders with the same dimension apart, so it trusts the caller.
  • test_bundle_models and test_bundle_full were not run (they need the full set of bundle fixtures).
  • LocalAI is not integrated yet. It builds registries with add_embedding plus a file-basename tag. It needs to store the family and weights identity per embedding and call add_embedding_fp.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

🤖 Generated with Claude Code

mudler added 3 commits October 4, 2026 11:32
…naming

Equal embedding sizes do not mean the same embedding space (ECAPA and
CAM++ both give 192 values), yet a registry only stored the size, so a
swapped encoder was accepted and names came from the wrong space.

PKSR version 2 stores an encoder fingerprint: a family
("voicedetect:<arch>:<name>:<dim>", from the GGUF metadata) and a
weights hash (sha256 of the encoder GGUF bytes, the same string as
parakeet_capi_speaker_identity). A registry without a fingerprint is
still written as version 1, and version 1 files load unchanged. A
reader that only knows version 1 refuses version 2.

Enrolment records the fingerprint of the encoder used. Where an encoder
meets a registry (scene, identify, named diarization, profiles,
speaker-attributed ASR, scene stream) the same check runs before any
name is assigned: another family is an error naming both families,
other weights of the same family and a missing fingerprint are
warnings, and strict mode turns the missing fingerprint into an error.

New C-API (additive, ABI stays 10): add_embedding_fp, registry
encoder_family and encoder_weights, set_strict, speaker encoder_family
and last_warning. New CLI: `registry <file>` shows the fingerprint and
`--restamp --encoder <gguf>` stamps a registry that has none, only on
request. `scene` gains --strict-registry.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Conflicts were in src/speaker_encoder.cpp and examples/cli/main.cpp.

SpeakerEncoder::load keeps the bundle component argument and now sets the
weights identity itself: the sha256 of the file bytes (hashed on both sides
of the load) for a standalone file, and the bundle header's source_sha256
for a voice component. The registry fingerprint therefore carries the same
weights identity for an encoder enrolled from the standalone file and used
from a bundle. The context identity in the C-API reads the same value.

The CLI usage strings keep the bundle component options and the registry
options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
test_speaker_fingerprint takes an optional bundle (PARAKEET_TEST_VD_BUNDLE)
whose voice component was built from the standalone encoder file. It checks
that both give the same family and weights identity, and that a registry
enrolled with one is used with the other with no warning.

test_capi_speaker checked for "expects" in the size mismatch error of
identify; that message now reads "registry holds N-value embeddings, this
model produces M".

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 4bc34c5 into master Oct 4, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants