Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
113 changes: 109 additions & 4 deletions docs/diarization.md
Original file line number Diff line number Diff line change
Expand Up @@ -243,10 +243,12 @@ protect a memory-mapped model from later writes either.
An application must gate profile export and enrollment with its recognition
permissions. Only after explicit user confirmation should it validate version,
unavailable status, finite/nonzero vector, trusted identity and dimension, then
call the existing raw-vector registration path. The native registry validates
vectors/dimensions but does **not** store model identity: callers must enforce
identity and avoid mixing same-dimension encoders. Profiles are not signed and
are not proof of identity. Treat exported voice vectors as sensitive data.
call the registration path. Pass the trusted encoder family and identity with
the vector (`parakeet_capi_speaker_registry_add_embedding_fp`, see "Encoder
fingerprint" below), so the registry records which encoder made it and every
named call checks it. The plain `add_embedding` still works and records
nothing. Profiles are not signed and are not proof of identity. Treat exported
voice vectors as sensitive data.

The native registry is **name-keyed and aggregating**:
`parakeet_capi_speaker_registry_add_embedding` calls `SpeakerRegistry::enroll`,
Expand All @@ -263,3 +265,106 @@ application. This is a downstream application responsibility, not functionality
implemented by this backend's profile export. Relabel only after registration
succeeds. Profile export adds no persistence or automatic enrollment; it does not
change an application's global, in-memory registry lifecycle.

### Encoder fingerprint

Equal embedding sizes do not mean the same embedding space: ECAPA and CAM++
both give 192 values, and a registry enrolled with one names the wrong people
when the other is used. A registry therefore records which encoder made its
voices, and the encoder in use is checked against it before any name is
assigned. Two strings make the fingerprint:

- **Family**: `voicedetect:<voicedetect.arch>:<general.name>:<voicedetect.embedding_dim>`,
read from the encoder GGUF metadata, for example
`voicedetect:ecapa_tdnn:speechbrain/spkrec-ecapa-voxceleb:192`. It names the
embedding space. `parakeet_capi_speaker_encoder_family` returns it.
- **Weights**: `sha256:<64 hex>` of the exact bytes of the encoder GGUF file,
the same string as `parakeet_capi_speaker_identity`. A GGUF has no recorded
source hash, so the file bytes are the definition. Another quantization of the
same encoder has the same family and another weights hash. (A `voice`
component of a bundle reports `sha256:` plus the `source_sha256` of its header,
the hash of the single-model GGUF it came from, so the two agree when the
component is an unchanged copy of that file.)

What the check does, the same everywhere (`parakeet-cli scene`,
`parakeet_capi_speaker_identify_pcm_json`, `parakeet_capi_diarize_named_pcm_json`,
`parakeet_capi_diarize_profiles_pcm_json`,
`parakeet_capi_transcribe_and_diarize_named_json`,
`parakeet_capi_scene_stream_begin_speaker`), with the same message text:

| Registry vs encoder | Result |
|---|---|
| Other embedding size | Error: `registry holds N-value embeddings, this model produces M`. |
| Other family | Error naming both families. No name is assigned. |
| Same family, other weights | Warning only: logged to stderr, kept in `parakeet_capi_speaker_last_warning`. Names are assigned. |
| No fingerprint (a version 1 file, or embeddings added without one) | Accepted with a warning: the encoder is unverified. |
| No fingerprint and strict mode (`parakeet_capi_speaker_registry_set_strict`, `parakeet-cli scene --strict-registry`) | Error. |
| Empty registry | Nothing to check. |

An error is reported like any other failure of that call: NULL (or nonzero)
with the message on the speaker context, and a non-zero exit in the CLI.

Enrolment records the fingerprint of the encoder that computed the voice:
`parakeet-cli enroll`, `parakeet_capi_speaker_enroll` and
`SpeakerIdentifier` enrolment do it themselves. An empty registry takes the
fingerprint of its first voice. Enrolling with another family is refused. A
registry that has voices but no fingerprint refuses a fingerprinted voice, and a
fingerprinted registry refuses a voice with none: it is never stamped
silently, because nothing can verify what made the old voices. A caller that
builds registries from stored embeddings passes the family and identity it
stored with them to `parakeet_capi_speaker_registry_add_embedding_fp`.

To stamp a registry that has no fingerprint, say which encoder made it:

```
parakeet-cli registry reg.bin # show the file
parakeet-cli registry reg.bin --restamp --encoder speaker.gguf # stamp a version 1 file
```

`registry` prints the format version, the embedding size, the family, the
weights hash and the speaker names. `--restamp` only works on a registry with no
fingerprint, checks the embedding size against the encoder, and trusts you for
the rest: it cannot tell ECAPA from CAM++ when both give 192 values.

#### Registry file format

Little-endian. `PKSR` magic, then:

```
version 1 (no fingerprint; also what a registry without one is saved as)
offset size field
0 4 "PKSR"
4 4 u32 version = 1
8 4 i32 dim
12 4 u32 n, the number of speakers
16 ... n speaker records

version 2 (with a fingerprint)
0 4 "PKSR"
4 4 u32 version = 2
8 4 i32 dim
12 4 u32 n
16 4 u32 family length (at most 4096)
20 ... family bytes (UTF-8, not terminated)
... 4 u32 weights length (at most 4096)
... ... weights bytes
... ... n speaker records

speaker record (both versions)
4 u32 name length (1 to 4096)
... name bytes
4 i32 count of enrolled clips (at least 1)
4*dim f32 sum of the L2-normalized embeddings
```

A version 2 file has at least one non-empty fingerprint string. Readers refuse
an unknown version, a truncated file, trailing bytes and implausible lengths. A
reader from before this change only knows version 1, so it refuses a version 2
file with "unsupported version" instead of reading it wrong. A version 1 file
loads unchanged.

The C-API additions are `parakeet_capi_speaker_registry_add_embedding_fp`,
`parakeet_capi_speaker_registry_encoder_family`, `..._encoder_weights`,
`parakeet_capi_speaker_registry_set_strict`, `parakeet_capi_speaker_encoder_family`
and `parakeet_capi_speaker_last_warning`. They are additive; the ABI version
stays 10 and no signature changed.
18 changes: 13 additions & 5 deletions docs/speaker.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,9 +59,14 @@ well. Sizes are for the f32 files. ERes2Net has not been run through any of
the tests here.

A registry belongs to the encoder that made it. The embedding sizes differ, and
even two encoders with the same size do not share a space, so enroll again if
you switch models. `scene` checks the size and stops if it does not match; it
cannot tell two encoders of the same size apart.
even two encoders with the same size do not share a space (ECAPA and CAM++ both
give 192 values), so enroll again if you switch models. The registry records
the encoder's family and weights hash, and `scene` and the C-API check them
before they assign a name: another family stops with an error that names both,
another quantization of the same family only warns, and a registry from before
this check (no fingerprint) warns, or stops with `--strict-registry`. See
"Encoder fingerprint" in [diarization.md](diarization.md) for the rules and the
file format. `parakeet-cli registry <file>` shows what a registry records.

The speaker-model weights have their own licences (WeSpeaker, 3D-Speaker and
SpeechBrain each publish theirs). voice-detect.cpp's own licence does
Expand Down Expand Up @@ -104,7 +109,7 @@ the margin. Names are compared exactly, so near-duplicate names (`Ada` and
```
parakeet-cli scene --model <asr.gguf> --diar <diar.gguf> \
--speakers <speaker.gguf> --registry <file> [--speaker-threshold F] \
--input <wav>
[--strict-registry] --input <wav>
```

`--speakers` needs `--diar` and `--registry`. Real output on the same fixture
Expand Down Expand Up @@ -179,7 +184,10 @@ speaker GGUF loads through `parakeet_capi_load` into a context of kind
parakeet_capi_speaker_dim # embedding size, -1 if not a speaker ctx
parakeet_capi_speaker_registry_new / _free / _size / _last_error
parakeet_capi_speaker_enroll # embed PCM and add it under a name
parakeet_capi_speaker_registry_save / _load # binary file
parakeet_capi_speaker_registry_save / _load # binary file (version 2 holds the encoder fingerprint)
parakeet_capi_speaker_registry_add_embedding_fp # add_embedding plus the encoder family and weights
parakeet_capi_speaker_registry_encoder_family / _encoder_weights / _set_strict
parakeet_capi_speaker_encoder_family / _last_warning
parakeet_capi_speaker_identify_pcm_json # {"name":"alice","score":0.71}
parakeet_capi_scene_stream_begin_speaker # scene stream with a speaker ctx + registry
parakeet_capi_transcribe_and_diarize_named_json # offline speaker-attributed ASR with names
Expand Down
121 changes: 107 additions & 14 deletions examples/cli/main.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -1732,7 +1732,9 @@ static std::string registry_read_error(const std::string& path, int e) {

static const char* kEnrollUsage =
"usage: parakeet-cli enroll --model <speaker.gguf|bundle.gguf> [--component NAME] --name <name> "
"--input <wav> [--input <wav> ...] --registry <file>\n";
"--input <wav> [--input <wav> ...] --registry <file>\n"
" The registry records the encoder (family and weights hash). An existing\n"
" registry with speakers and no fingerprint is refused: see `parakeet-cli registry`.\n";

// parakeet-cli enroll --model <speaker.gguf> --name <name> --input <wav> [--input <wav> ...]
// --registry <file>
Expand Down Expand Up @@ -1781,6 +1783,7 @@ static int cmd_enroll(int argc, char** argv) {
}
}
int clips = 0;
bool warned = false;
for (const std::string& in : inputs) {
pk::Audio audio;
if (!load_audio_arg_16k_mono(in, audio)) {
Expand All @@ -1794,8 +1797,14 @@ static int cmd_enroll(int argc, char** argv) {
enc->last_error().c_str());
return 1;
}
try { reg.enroll(name, emb); }
catch (const std::exception& e) {
try {
// Records the encoder that made the voice (family and weights hash).
const pk::FingerprintVerdict v = reg.enroll(name, emb, enc->fingerprint());
if (v.is_warning() && !warned) {
std::fprintf(stderr, "parakeet-cli enroll: warning: %s\n", v.message.c_str());
warned = true;
}
} catch (const std::exception& e) {
std::fprintf(stderr, "parakeet-cli enroll: %s\n", e.what());
return 1;
}
Expand All @@ -1813,15 +1822,90 @@ static int cmd_enroll(int argc, char** argv) {
return 0;
}

static const char* kRegistryUsage =
"usage: parakeet-cli registry <file> [--restamp --encoder <speaker.gguf>]\n"
" Prints the speaker registry's format version, embedding size, encoder\n"
" fingerprint and speakers. With --restamp it writes the fingerprint of\n"
" the given encoder into a registry that has none (a version 1 file).\n"
" Only restamp with the encoder that made the voices: nothing can verify it.\n";

// parakeet-cli registry <file> [--restamp --encoder <speaker.gguf>]
static int cmd_registry(int argc, char** argv) {
std::string path, encoder;
bool restamp = false;
for (int i = 0; i < argc; ++i) {
if (std::strcmp(argv[i], "--restamp") == 0) restamp = true;
else if (std::strcmp(argv[i], "--encoder") == 0 && i + 1 < argc) encoder = argv[++i];
else if (argv[i][0] != '-' && path.empty()) path = argv[i];
else { std::fprintf(stderr, "%s", kRegistryUsage); return 2; }
}
if (path.empty() || (restamp && encoder.empty()) || (!restamp && !encoder.empty())) {
std::fprintf(stderr, "%s", kRegistryUsage);
return 2;
}
std::string blob;
const int rerr = read_file_bytes(path, blob);
if (rerr != 0) {
std::fprintf(stderr, "parakeet-cli registry: %s\n", registry_read_error(path, rerr).c_str());
return 1;
}
pk::SpeakerRegistry reg;
try { reg = pk::SpeakerRegistry::deserialize(blob); }
catch (const std::exception& e) {
std::fprintf(stderr, "parakeet-cli registry: %s is not a speaker registry: %s\n", path.c_str(), e.what());
return 1;
}
if (restamp) {
if (!reg.fingerprint().empty()) {
std::fprintf(stderr, "parakeet-cli registry: %s already has a fingerprint (%s); "
"enroll again to change the encoder\n", path.c_str(), reg.fingerprint().family.c_str());
return 1;
}
if (!pk::SpeakerEncoder::available()) {
std::fprintf(stderr, "parakeet-cli: built without speaker identification (PARAKEET_WITH_VOICEDETECT=OFF)\n");
return 2;
}
auto enc = pk::SpeakerEncoder::load(encoder);
if (!enc) {
std::fprintf(stderr, "parakeet-cli registry: failed to load speaker model %s\n", encoder.c_str());
return 1;
}
if (reg.dim() != 0 && reg.dim() != enc->dim()) {
std::fprintf(stderr, "parakeet-cli registry: registry holds %d-value embeddings, %s produces %d; "
"this is not the encoder that made them\n", reg.dim(), encoder.c_str(), enc->dim());
return 1;
}
reg.set_fingerprint(enc->fingerprint());
std::string werr;
if (!pk::write_file_atomic(path, reg.serialize(), &werr)) {
std::fprintf(stderr, "parakeet-cli registry: %s\n", werr.c_str());
return 1;
}
std::printf("restamped %s with the fingerprint of %s\n", path.c_str(), encoder.c_str());
}
uint32_t ver = 0;
std::memcpy(&ver, blob.data() + 4, 4);
if (restamp) ver = 2;
const pk::EncoderFingerprint& fp = reg.fingerprint();
std::printf("format version: %u\n", ver);
std::printf("embedding size: %d\n", reg.dim());
std::printf("encoder family: %s\n", fp.family.empty() ? "(none)" : fp.family.c_str());
std::printf("encoder weights: %s\n", fp.weights.empty() ? "(none)" : fp.weights.c_str());
std::printf("speakers: %zu\n", reg.size());
for (const std::string& n : reg.names()) std::printf(" %s\n", n.c_str());
return 0;
}

static const char* kSceneUsage =
"usage: parakeet-cli scene [--model <m.gguf>] [--diar <diar.gguf>] "
"[--sound <ced.gguf>] [--speakers <speaker.gguf> --registry <file> "
"[--speaker-threshold F]] --input <wav|-> "
"[--speaker-threshold F] [--strict-registry]] --input <wav|-> "
"[--latency model|low|very_low|ultra_low] [--chunk-ms N] "
"[--show-speech] [--json]\n"
" each model may be a bundle GGUF (the only component of the right kind is used); name another with\n"
" --asr-component, --diar-component, --sound-component or --speakers-component\n"
" --speaker-threshold: default 0.5; ECAPA needs about 0.7, see docs/speaker.md\n";
" --speaker-threshold: default 0.5; ECAPA needs about 0.7, see docs/speaker.md\n"
" --strict-registry: refuse a registry with no encoder fingerprint\n";

// parakeet-cli scene [--model <m.gguf>] [--diar <diar.gguf>] [--sound <ced.gguf>]
// [--speakers <speaker.gguf> --registry <file> [--speaker-threshold F]]
Expand All @@ -1836,6 +1920,7 @@ static const char* kSceneUsage =
static int cmd_scene(int argc, char** argv) {
std::string model, diar, sound, input, latency_str;
std::string speakers, registry_path;
bool strict_registry = false;
std::string asr_comp_arg, diar_comp_arg, sound_comp_arg, speakers_comp_arg;
bool have_threshold = false;
float speaker_threshold = 0.0f;
Expand All @@ -1861,6 +1946,8 @@ static int cmd_scene(int argc, char** argv) {
speakers = argv[++i];
} else if (std::strcmp(argv[i], "--registry") == 0 && i + 1 < argc) {
registry_path = argv[++i];
} else if (std::strcmp(argv[i], "--strict-registry") == 0) {
strict_registry = true;
} else if (std::strcmp(argv[i], "--speaker-threshold") == 0 && i + 1 < argc) {
char* end = nullptr;
const char* txt = argv[++i];
Expand Down Expand Up @@ -1921,8 +2008,8 @@ static int cmd_scene(int argc, char** argv) {
std::fprintf(stderr, "parakeet-cli: built without sound tagging (PARAKEET_WITH_CED=OFF)\n");
return 2;
}
if (speakers.empty() && (!registry_path.empty() || have_threshold)) {
std::fprintf(stderr, "parakeet-cli scene: --registry and --speaker-threshold need --speakers\n");
if (speakers.empty() && (!registry_path.empty() || have_threshold || strict_registry)) {
std::fprintf(stderr, "parakeet-cli scene: --registry, --speaker-threshold and --strict-registry need --speakers\n");
return 2;
}
pk::SpeakerIdOpts speaker_opts;
Expand Down Expand Up @@ -2001,13 +2088,16 @@ static int cmd_scene(int argc, char** argv) {
registry_path.c_str(), e.what());
return 1;
}
if (registry.dim() != speaker_enc->dim()) {
std::fprintf(stderr,
"parakeet-cli scene: registry %s holds %d-dim voices but %s makes %d-dim embeddings "
"(enroll again with this model)\n",
registry_path.c_str(), registry.dim(), speakers.c_str(), speaker_enc->dim());
// Same check and same text as the C-API, before any name is assigned.
const pk::FingerprintVerdict v = pk::check_registry_for_encoder(
registry, speaker_enc->dim(), speaker_enc->fingerprint(), strict_registry);
if (v.is_error()) {
std::fprintf(stderr, "parakeet-cli scene: %s (registry %s, speaker model %s)\n",
v.message.c_str(), registry_path.c_str(), speakers.c_str());
return 1;
}
if (v.is_warning())
std::fprintf(stderr, "parakeet-cli scene: warning: %s\n", v.message.c_str());
}

pk::Audio audio;
Expand Down Expand Up @@ -2256,6 +2346,8 @@ int main(int argc, char** argv) {
return run_and_shutdown(cmd_bench, argc - 2, argv + 2);
if (argc >= 2 && std::strcmp(argv[1], "enroll") == 0)
return run_and_shutdown(cmd_enroll, argc - 2, argv + 2);
if (argc >= 2 && std::strcmp(argv[1], "registry") == 0)
return run_and_shutdown(cmd_registry, argc - 2, argv + 2);
if (argc >= 2 && std::strcmp(argv[1], "scene") == 0)
return run_and_shutdown(cmd_scene, argc - 2, argv + 2);
if (argc >= 2 && std::strcmp(argv[1], "vad") == 0)
Expand Down Expand Up @@ -2286,12 +2378,13 @@ int main(int argc, char** argv) {
"[--batch-sizes 1,4,8,16] [--threads N] [--reps R] [--json <out>]\n"
" parakeet-cli scene [--model <m.gguf>] [--diar <diar.gguf>] "
"[--sound <ced.gguf>] [--speakers <speaker.gguf> --registry <file> "
"[--speaker-threshold F]] --input <wav|-> "
"[--speaker-threshold F] [--strict-registry]] --input <wav|-> "
"[--latency model|low|very_low|ultra_low] [--chunk-ms N] "
"[--show-speech] [--json]\n"
" --speaker-threshold: default 0.5; ECAPA needs about 0.7, see docs/speaker.md\n"
" each model may be a bundle (--asr-component, --diar-component, --sound-component, --speakers-component)\n"
" parakeet-cli enroll --model <speaker.gguf|bundle.gguf> [--component NAME] --name <name> "
"--input <wav> [--input <wav> ...] --registry <file>\n");
"--input <wav> [--input <wav> ...] --registry <file>\n"
" parakeet-cli registry <file> [--restamp --encoder <speaker.gguf>]\n");
return 2;
}
Loading
Loading