Skip to content

feat: load ced and voice bundle components from memory (no temp file), bump ced.cpp and voice-detect.cpp - #92

Merged
mudler merged 1 commit into
masterfrom
feat/bundle-load-from-memory
Oct 4, 2026
Merged

mudler merged 1 commit into
masterfrom
feat/bundle-load-from-memory

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

What

The ced and voice components of a bundle are now loaded straight from memory, with no standalone copy and no temporary file.

  • Bump third_party/ced.cpp to 736a4ee (ced_capi_load_from_memory_prefixed) and third_party/voice-detect.cpp to bca46bc (voicedetect_capi_load_from_memory_prefixed). Both are on the default branch of their repositories. .gitmodules still points at the localai-org repositories; CMake and CI (submodules: recursive) need no change.
  • CedTagger::load and SpeakerEncoder::load map the bundle read-only (src/bundle_map.cpp: mmap on Linux and macOS, MapViewOfFile on Windows) and pass the map and the prefix (ced., voice.) to the hook. The hook parses the header and copies only the tensors of that component, then the map is released.
  • src/bundle_extract.*, pk::ComponentFile and PARAKEET_BUNDLE_NO_MEMFD are removed (about 250 lines). Nothing needs them on any platform.
  • Both loaders check that the component exists and has the right kind before touching a tensor, and add ced_capi_last_error(NULL) or voicedetect_capi_last_load_error() to the error message. The speaker encoder family is read from the prefixed keys of the bundle header, so the fingerprint and identity are unchanged.

What bytes the hooks get

The prefixed hooks need a buffer that parses as a GGUF holding the component's keys and tensor table. That is the whole bundle image, not only the component. Reading the whole image is not acceptable (1.1 GB for the standard bundle), so the buffer is a read-only map of the file: the hook touches only the header and the tensors under the prefix, and the operating system reads only those pages. The other components are never read. No sparse image or minimal GGUF writer is needed.

Memory and I/O

Standard bundle, 1100.8 MB (asr 940.5, diar 108.6, ced-small Q8_0 23.6, voice WeSpeaker ResNet34 F32 26.5, vad), one component loaded in a fresh process:

Component Size Peak resident memory above idle Kept after load Read from storage (page cache dropped)
ced 23.6 MB 49 MB 25 MB 24.2 MB
voice 26.5 MB 54 MB 27 MB 26.5 MB

The peak is the copy of the component (kept) plus the touched pages of the map (file cache, released with the map). It depends on the component, not on the bundle: the 338 MB small bundle gives the same figures. I did not measure the old memfd path; it held one more copy of the component during the load.

Behaviour

No memfd, no temporary file, no environment variable; the bundle is opened read-only for the moment it takes to map it and closed again. The same code runs on every platform. Windows and macOS were not run (standard mmap and MapViewOfFile code). CED scores and speaker embeddings are bitwise equal to the standalone files.

Tests

  • ctest -LE model: 38 of 38 pass (CED and voice-detect ON, Release).
  • ctest -L model with real files (tdt_ctc-110m Q8_0, Nemotron diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34 F32, silero-vad-f16, all five in a bundle built by scripts/bundle_gguf.py): 86 of 87 pass; the only failure is test_relpos_attention_local_chunked, known and unrelated.
  • test_bundle_full (also with the published parakeet-bundle-small.gguf): identical transcript, diarization JSON, CED scores (bitwise), speaker embedding (bitwise), identity, Silero probabilities and named diarization. The bundle block of test_speaker_fingerprint, test_bundle_models and check_bundle.py pass.
  • New checks: the loads run with a read-only TMPDIR; the descriptor count, the working directory and /proc/self/maps show no temporary file, memory file or descriptor; the bytes read from storage per component (page cache dropped first) stay near the component size and under half of the bundle. The old bytes-read check used rchar, which does not count mapped reads, so the expectation changed: now about 1x the component size instead of up to 3x. test_bundle adds unit checks for the map (content, missing file, empty file) and the new component and kind errors.

Limits

  • A model file must not change while it loads: if it is shortened while mapped, a read past the end ends the process. Model files are meant to be immutable in service.
  • The whole bundle needs address space, not memory; a 32-bit process cannot map a bundle larger than its address space (the load fails with "cannot map").
  • Prefixed tensor names stay under the 63-byte ggml limit, as before.

docs/bundle.md ("ced and voice components", "Partial loading") and AGENTS.md are updated.

🤖 Generated with Claude Code

…and voice-detect.cpp

ced.cpp and voice-detect.cpp can now load a model from memory, with a
prefix for a model stored inside a larger GGUF. Move both submodules to
the commits that add this (ced.cpp 736a4ee, voice-detect.cpp bca46bc).

Use it for the ced and voice components of a bundle. The bundle is
mapped read-only (mmap, MapViewOfFile on Windows) and handed to the
prefixed loader, which copies only the tensors of that component during
the call. No standalone copy, temporary file, memory file or environment
variable is needed any more, so bundle_extract and
PARAKEET_BUNDLE_NO_MEMFD are removed. Pages of the other components are
never touched. Loading ced-small Q8_0 from the 1.1 GB standard bundle
peaks 49 MB above an idle process and reads 24 MB from storage.

Both loaders now check that the component exists and has the right kind,
and add the loader's own error text to the message. The speaker encoder
family is read from the prefixed keys of the bundle.

The full-bundle test runs with a read-only TMPDIR, checks that the
descriptor count and the working directory do not change, and measures
the bytes read from storage per component with the page cache dropped.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 06eb8f9 into master Oct 4, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants