Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -189,6 +189,7 @@ The frozen symbol set:
- `voicedetect_capi_abi_version` - integer LocalAI checks for compatibility; bump it on any breaking change to the frozen symbols (additive functions do not require a bump).
- `voicedetect_capi_load` / `voicedetect_capi_free` - load a GGUF model, free the context.
- `voicedetect_capi_load_from_memory` / `voicedetect_capi_load_from_memory_prefixed` / `voicedetect_capi_last_load_error` - load a model from a buffer instead of a path, and read why a load failed (see "Loading from memory" below).
- `voicedetect_capi_encoder_arch` / `voicedetect_capi_encoder_name` / `voicedetect_capi_encoder_family` - which encoder is loaded, and a string that names its embedding space (see "Encoder identity" below).
- `voicedetect_capi_last_error` - human-readable last error on a context.
- `voicedetect_capi_embed_path` / `voicedetect_capi_embed_pcm` - L2-normalized embedding from a WAV file or in-memory mono float PCM (PCM is linearly resampled to 16 kHz if needed).
- `voicedetect_capi_verify_paths` - cosine distance (`1 - cosine_similarity`) between two clips plus a same-speaker verdict against a threshold.
Expand Down Expand Up @@ -238,6 +239,29 @@ free(bytes); // allowed: the loader has already copied what it needs

Limits: the buffer must hold the complete GGUF (the loader does not stream). The tensors are copied once, so the peak memory is the buffer plus the model.

### Encoder identity

Two encoders can give embeddings of the same size and still not share an embedding space: ECAPA-TDNN and CAM++ both give 192 values. A store of enrolled voices (a speaker registry) must therefore record which encoder made its embeddings and refuse embeddings from another one. Three accessors report what the loader already read from the GGUF header:

```c
const char* voicedetect_capi_encoder_arch(const voicedetect_ctx*); // voicedetect.arch, e.g. "ecapa_tdnn"
const char* voicedetect_capi_encoder_name(const voicedetect_ctx*); // general.name
const char* voicedetect_capi_encoder_family(const voicedetect_ctx*); // see below
```

The family string has exactly this format:

```
voicedetect:<voicedetect.arch>:<general.name>:<voicedetect.embedding_dim>
```

For example `voicedetect:ecapa_tdnn:speechbrain/spkrec-ecapa-voxceleb:192`. The values are copied as they are, with no escaping. A missing key gives an empty field and the colons stay. The size is in decimal and is empty when it is 0 (an analyze model). The string is empty when `general.architecture` is not `voicedetect`. It names the embedding space, so another quantization of the same encoder has the same family.

- The pointers belong to the context, never change, and stay valid until `voicedetect_capi_free`. Do not free them. They are NULL for a NULL context. They can be read from several threads at once.
- They work for a context from every load function. A model loaded with a prefix reads the prefixed keys, so a component of a bundle gives the same family as the standalone file it was made from.
- [parakeet.cpp](https://github.com/mudler/parakeet.cpp) defines the encoder fingerprint of its speaker registry with the same formula (`speaker_encoder_family`). The two definitions must stay identical: change neither without the other.
- This is an additive change: the ABI version stays 1. C++ code can read `vd::Model::config().arch`, `.name` and `.family`.

---

## Model coverage
Expand Down
36 changes: 36 additions & 0 deletions include/voicedetect_capi.h
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,42 @@ char* voicedetect_capi_analyze_path_json(voicedetect_ctx* ctx,
// a NULL ctx. Additive: no ABI version bump.
int voicedetect_capi_embedding_dim(voicedetect_ctx* ctx);

// Identity of the loaded encoder, read from the GGUF header the loader already
// parsed. For a model loaded with a prefix the keys are the prefixed ones, so a
// component of a bundle reports the same values as the standalone file it was
// made from. All three work for every load path (path, memory, prefixed memory).
// Each returns NULL for a NULL ctx. The returned pointer is owned by the context,
// never changes, and stays valid until voicedetect_capi_free; do not free it.
// Reading it from several threads at once is safe. Additive: no ABI version bump.
//
// voicedetect_capi_encoder_arch: the `voicedetect.arch` value, for example
// "ecapa_tdnn", "campplus", "wespeaker_resnet34", "eres2net". "" if the key
// is absent.
const char* voicedetect_capi_encoder_arch(const voicedetect_ctx* ctx);

// voicedetect_capi_encoder_name: the `general.name` value (the checkpoint the
// converter was run on, for example "speechbrain/spkrec-ecapa-voxceleb"). ""
// if the key is absent.
const char* voicedetect_capi_encoder_name(const voicedetect_ctx* ctx);

// voicedetect_capi_encoder_family: a string that names the embedding space, for
// a registry of voices to record next to its embeddings and compare before it
// matches a new embedding. Equal embedding sizes do not mean the same space
// (ECAPA-TDNN and CAM++ both give 192 values). The format is exactly
//
// voicedetect:<voicedetect.arch>:<general.name>:<voicedetect.embedding_dim>
//
// for example "voicedetect:ecapa_tdnn:speechbrain/spkrec-ecapa-voxceleb:192".
// A missing key gives an empty field (the colons stay), and no escaping is
// done: the values are copied as they are. The embedding_dim is written in
// decimal, and is an empty field when it is 0 (an analyze model). If
// `general.architecture` is not "voicedetect" the string is "" (not a
// voicedetect model). Another quantization of the same encoder has the same
// family. parakeet.cpp defines its speaker-registry encoder fingerprint with
// the same formula (speaker_encoder_family); the two MUST stay identical, so
// change neither without the other.
const char* voicedetect_capi_encoder_family(const voicedetect_ctx* ctx);

#ifdef __cplusplus
} // extern "C"
#endif
Expand Down
18 changes: 18 additions & 0 deletions src/model_loader.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,14 @@ static bool kv_bool(const Kv& kv, const char* k, bool d=false){
static std::string kv_str(const Kv& kv, const char* k, const char* d=""){
int64_t id = kv.find(k, GGUF_TYPE_STRING); return id<0 ? std::string(d) : std::string(gguf_get_val_str(kv.g,id));
}
// A string key that is absent or of another type reads as "". Unlike kv_str it
// never records a load error: the identity keys are descriptive, and a model
// that loaded before must still load.
static std::string kv_str_soft(const Kv& kv, const char* k){
int64_t id = gguf_find_key(kv.g, (kv.prefix + k).c_str());
if(id < 0 || gguf_get_kv_type(kv.g, id) != GGUF_TYPE_STRING) return std::string();
return std::string(gguf_get_val_str(kv.g, id));
}
static std::vector<std::string> kv_str_arr(const Kv& kv, const char* k){
std::vector<std::string> out;
int64_t id = kv.find_arr(k, GGUF_TYPE_STRING);
Expand Down Expand Up @@ -235,6 +243,16 @@ bool ModelLoader::finish_(const Kv& kv, bool ctx_names){
cfg_.arch = kv_str(kv, "voicedetect.arch");
cfg_.embedding_dim = kv_u32(kv, "voicedetect.embedding_dim");
cfg_.l2_normalize = kv_bool(kv, "voicedetect.l2_normalize", true);
// Encoder identity (see VoiceDetectConfig::family). Keys come from the same
// prefix as every other key, so a bundle component and the standalone file
// it was made from give the same string.
cfg_.name = kv_str_soft(kv, "general.name");
if(kv_str_soft(kv, "general.architecture") == "voicedetect"){
cfg_.family = "voicedetect:" + kv_str_soft(kv, "voicedetect.arch") + ":" + cfg_.name + ":" +
(cfg_.embedding_dim > 0 ? std::to_string(cfg_.embedding_dim) : std::string());
} else {
cfg_.family.clear();
}
// FBank front end
cfg_.sample_rate = kv_u32(kv, "voicedetect.fbank.sample_rate", 16000);
cfg_.n_mels = kv_u32(kv, "voicedetect.fbank.n_mels", 80);
Expand Down
7 changes: 7 additions & 0 deletions src/model_loader.hpp
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,13 @@ struct VoiceDetectConfig {
// "eres2net" 3D-Speaker ERes2Net (ONNX-direct)
// "campplus" 3D-Speaker CAM++ (ONNX-direct)
std::string arch;
// `general.name` of the GGUF ("" if absent or not a string).
std::string name;
// Encoder family string "voicedetect:<arch>:<name>:<embedding_dim>" (the
// embedding_dim is empty when it is 0), or "" when `general.architecture`
// is not "voicedetect". parakeet.cpp defines the same string for its
// speaker registry (speaker_encoder_family); the two must stay identical.
std::string family;
// Embedding head
uint32_t embedding_dim = 0; // 192 for ECAPA-TDNN; 256 for WeSpeaker/CAM++
bool l2_normalize = true; // L2-normalize the output embedding (cosine space)
Expand Down
15 changes: 15 additions & 0 deletions src/voicedetect_capi.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,21 @@ extern "C" int voicedetect_capi_embedding_dim(voicedetect_ctx* ctx) {
return ctx->model->embedding_dim();
}

extern "C" const char* voicedetect_capi_encoder_arch(const voicedetect_ctx* ctx) {
if (!ctx || !ctx->model) return nullptr;
return ctx->model->config().arch.c_str();
}

extern "C" const char* voicedetect_capi_encoder_name(const voicedetect_ctx* ctx) {
if (!ctx || !ctx->model) return nullptr;
return ctx->model->config().name.c_str();
}

extern "C" const char* voicedetect_capi_encoder_family(const voicedetect_ctx* ctx) {
if (!ctx || !ctx->model) return nullptr;
return ctx->model->config().family.c_str();
}

extern "C" voicedetect_ctx* voicedetect_capi_load(const char* gguf_path) {
if (!gguf_path) { g_load_error = "model path is NULL"; return nullptr; }
try {
Expand Down
4 changes: 3 additions & 1 deletion tests/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,8 @@ vd_add_test(test_blocked)
vd_add_test(test_cache_safety)
vd_add_test(test_load_memory)
vd_add_test(test_load_memory_embed)
vd_add_test(test_encoder_family)
vd_add_test(test_encoder_family_models)

# Manual benchmark target (not registered as a ctest; run directly).
add_executable(bench_wespeaker bench_wespeaker.cpp)
Expand All @@ -30,7 +32,7 @@ target_include_directories(bench_eres2net PRIVATE ${PROJECT_SOURCE_DIR}/src ${PR

# Model/baseline-dependent tests read fixtures/baselines via paths relative to
# the project root, and skip (exit 77) when the baseline env var is unset.
set_tests_properties(test_fbank test_encoder test_embed test_verify test_analyze test_age_gender test_cache_safety test_load_memory_embed PROPERTIES
set_tests_properties(test_fbank test_encoder test_embed test_verify test_analyze test_age_gender test_cache_safety test_load_memory_embed test_encoder_family_models PROPERTIES
LABELS "model"
WORKING_DIRECTORY ${PROJECT_SOURCE_DIR})

Expand Down
204 changes: 204 additions & 0 deletions tests/test_encoder_family.cpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,204 @@
// Encoder identity accessors (voicedetect_capi_encoder_arch / _name / _family).
// No model files needed: the GGUFs are synthesized in memory, so this runs in CI.
//
// * the family string has exactly the documented format
// * path, memory and prefixed (bundle) loads give identical strings
// * a bundle's own header keys never leak into a component
// * missing keys, a wrong-typed name and odd characters
// * NULL context, pointer lifetime and concurrent reads
#include "gguf_bundle.hpp"

#include "voicedetect_capi.h"

#include <atomic>
#include <cstdio>
#include <cstdlib>
#include <random>
#include <string>
#include <thread>

using namespace vdtest;

static int g_fail = 0;
#define CHECK(c) do { if (!(c)) { std::fprintf(stderr, "FAIL %s:%d: %s\n", __FILE__, __LINE__, #c); ++g_fail; } } while (0)
#define CHECK_STR(got, want) do { const char* g_ = (got); const std::string w_ = (want); \
if (!g_ || w_ != g_) { std::fprintf(stderr, "FAIL %s:%d: got '%s' want '%s'\n", __FILE__, __LINE__, g_ ? g_ : "(null)", w_.c_str()); ++g_fail; } } while (0)

static std::string write_tmp(const std::vector<uint8_t>& b) {
std::string p = std::string("vd_test_encoder_family_") + std::to_string((long)std::random_device()()) + ".gguf";
FILE* f = std::fopen(p.c_str(), "wb");
if (f) { std::fwrite(b.data(), 1, b.size(), f); std::fclose(f); }
return p;
}

static void set_identity(Tiny& t, const char* arch_key, const char* name) {
gguf_set_val_str(t.g, "general.architecture", "voicedetect");
if (name) gguf_set_val_str(t.g, "general.name", name);
if (arch_key) gguf_set_val_str(t.g, "voicedetect.arch", arch_key); // replaces the value
}

struct Loaded {
voicedetect_ctx* path = nullptr;
voicedetect_ctx* mem = nullptr;
voicedetect_ctx* bundled = nullptr;
~Loaded() {
voicedetect_capi_free(path);
voicedetect_capi_free(mem);
voicedetect_capi_free(bundled);
}
};

// Load `t` through the three paths. The bundle holds an unrelated component and a
// header with its own general.* keys, next to `t` under "voice.".
static void load_all(Tiny& t, Loaded& L) {
Tiny other(7, "other");
set_identity(other, "ecapa_tdnn", "someone/else");
gguf_context* hdr = gguf_init_empty();
gguf_set_val_str(hdr, "general.architecture", "bundle-kind");
gguf_set_val_str(hdr, "general.name", "the-bundle");
auto solo = make_bundle({{"", t.g, t.ctx}});
auto bundle = make_bundle({{"", hdr, nullptr}, {"other.", other.g, other.ctx}, {"voice.", t.g, t.ctx}});
gguf_free(hdr);

const std::string path = write_tmp(solo);
L.path = voicedetect_capi_load(path.c_str());
std::remove(path.c_str());
L.mem = voicedetect_capi_load_from_memory(solo.data(), solo.size());
L.bundled = voicedetect_capi_load_from_memory_prefixed(bundle.data(), bundle.size(), "voice.");
}

static void test_format_and_equivalence() {
Tiny t(1, "stem");
set_identity(t, nullptr, "speechbrain/spkrec-ecapa-voxceleb");
Loaded L;
load_all(t, L);
CHECK(L.path && L.mem && L.bundled);
for (voicedetect_ctx* c : {L.path, L.mem, L.bundled}) {
if (!c) continue;
CHECK_STR(voicedetect_capi_encoder_arch(c), "wespeaker_resnet34");
CHECK_STR(voicedetect_capi_encoder_name(c), "speechbrain/spkrec-ecapa-voxceleb");
CHECK_STR(voicedetect_capi_encoder_family(c),
"voicedetect:wespeaker_resnet34:speechbrain/spkrec-ecapa-voxceleb:256");
}
// Same pointer on every call, and the same for repeated reads.
if (L.path) {
const char* a = voicedetect_capi_encoder_family(L.path);
const char* b = voicedetect_capi_encoder_family(L.path);
CHECK(a == b);
CHECK(voicedetect_capi_encoder_arch(L.path) == voicedetect_capi_encoder_arch(L.path));
}
}

static void test_missing_and_odd_keys() {
{ // no general.name: empty field, colons stay
Tiny t(1, "stem");
set_identity(t, nullptr, nullptr);
Loaded L;
load_all(t, L);
for (voicedetect_ctx* c : {L.path, L.mem, L.bundled}) {
CHECK(c != nullptr);
if (!c) continue;
CHECK_STR(voicedetect_capi_encoder_name(c), "");
CHECK_STR(voicedetect_capi_encoder_family(c), "voicedetect:wespeaker_resnet34::256");
}
}
{ // general.architecture other than voicedetect (or absent): no family, still loads
Tiny t(1, "stem");
gguf_set_val_str(t.g, "general.name", "x");
Loaded L; // general.architecture absent
load_all(t, L);
for (voicedetect_ctx* c : {L.path, L.mem, L.bundled}) {
CHECK(c != nullptr);
if (!c) continue;
CHECK_STR(voicedetect_capi_encoder_family(c), "");
CHECK_STR(voicedetect_capi_encoder_name(c), "x");
CHECK_STR(voicedetect_capi_encoder_arch(c), "wespeaker_resnet34");
}
Tiny u(1, "stem");
gguf_set_val_str(u.g, "general.architecture", "parakeet");
Loaded M;
load_all(u, M);
CHECK(M.path && M.mem && M.bundled);
if (M.path) CHECK_STR(voicedetect_capi_encoder_family(M.path), "");
}
{ // a wrong-typed general.name must not make a loadable model fail
Tiny t(1, "stem");
gguf_set_val_str(t.g, "general.architecture", "voicedetect");
gguf_set_val_u32(t.g, "general.name", 5);
Loaded L;
load_all(t, L);
for (voicedetect_ctx* c : {L.path, L.mem, L.bundled}) {
CHECK(c != nullptr);
if (!c) continue;
CHECK_STR(voicedetect_capi_encoder_name(c), "");
CHECK_STR(voicedetect_capi_encoder_family(c), "voicedetect:wespeaker_resnet34::256");
}
}
{ // values are copied as they are: colons, spaces and UTF-8 are not escaped
Tiny t(1, "stem");
set_identity(t, nullptr, "a:b c/\xC3\xA9");
Loaded L;
load_all(t, L);
for (voicedetect_ctx* c : {L.path, L.mem, L.bundled}) {
CHECK(c != nullptr);
if (!c) continue;
CHECK_STR(voicedetect_capi_encoder_family(c), "voicedetect:wespeaker_resnet34:a:b c/\xC3\xA9:256");
}
}
}

static void test_prefixed_reads_own_keys() {
// Two components with different identities in one bundle: each prefix reports
// its own, and a prefix never picks up the bundle header or the sibling.
Tiny a(1, "stem"), b(2, "blk");
set_identity(a, "ecapa_tdnn", "model-a");
set_identity(b, "campplus", "model-b");
gguf_context* hdr = gguf_init_empty();
gguf_set_val_str(hdr, "general.architecture", "bundle-kind");
gguf_set_val_str(hdr, "general.name", "the-bundle");
auto bundle = make_bundle({{"", hdr, nullptr}, {"a.", a.g, a.ctx}, {"b.", b.g, b.ctx}});
gguf_free(hdr);
voicedetect_ctx* ca = voicedetect_capi_load_from_memory_prefixed(bundle.data(), bundle.size(), "a.");
voicedetect_ctx* cb = voicedetect_capi_load_from_memory_prefixed(bundle.data(), bundle.size(), "b.");
CHECK(ca && cb);
if (ca) CHECK_STR(voicedetect_capi_encoder_family(ca), "voicedetect:ecapa_tdnn:model-a:256");
if (cb) CHECK_STR(voicedetect_capi_encoder_family(cb), "voicedetect:campplus:model-b:256");
voicedetect_capi_free(ca);
voicedetect_capi_free(cb);
}

static void test_null_and_threads() {
CHECK(voicedetect_capi_encoder_arch(nullptr) == nullptr);
CHECK(voicedetect_capi_encoder_name(nullptr) == nullptr);
CHECK(voicedetect_capi_encoder_family(nullptr) == nullptr);

Tiny t(1, "stem");
set_identity(t, nullptr, "n");
auto solo = make_bundle({{"", t.g, t.ctx}});
voicedetect_ctx* c = voicedetect_capi_load_from_memory(solo.data(), solo.size());
CHECK(c != nullptr);
if (!c) return;
std::atomic<int> bad{0};
std::vector<std::thread> th;
for (int i = 0; i < 8; ++i)
th.emplace_back([&] {
for (int k = 0; k < 2000; ++k) {
const char* f = voicedetect_capi_encoder_family(c);
const char* a = voicedetect_capi_encoder_arch(c);
const char* n = voicedetect_capi_encoder_name(c);
if (!f || !a || !n || std::string(f) != "voicedetect:wespeaker_resnet34:n:256") ++bad;
}
});
for (auto& x : th) x.join();
CHECK(bad == 0);
voicedetect_capi_free(c);
}

int main() {
test_format_and_equivalence();
test_missing_and_odd_keys();
test_prefixed_reads_own_keys();
test_null_and_threads();
std::fprintf(stderr, g_fail ? "FAILED (%d)\n" : "PASS\n", g_fail);
return g_fail ? 1 : 0;
}
Loading
Loading