Skip to content

Load a model from a memory buffer - #2

Merged
mudler merged 1 commit into
masterfrom
feat/load-from-memory
Oct 4, 2026
Merged

mudler merged 1 commit into
masterfrom
feat/load-from-memory

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown

What and why

A model can only be loaded from a file path today. A host that already has the GGUF in memory, for example one component of a larger bundle file that holds several models, has to write a copy to disk (or to an in-memory file that only exists on Linux) and pass that path. This PR adds loaders that take the model from memory. They need no file and no file descriptor, and they work the same on every platform.

This is additive. The path loader keeps its behaviour and shares the config parsing with the new loaders.

API

// Complete GGUF file in memory.
voicedetect_ctx* voicedetect_capi_load_from_memory(const void* data, size_t size);

// One model inside a larger GGUF: every key and tensor name is stored as prefix + name.
voicedetect_ctx* voicedetect_capi_load_from_memory_prefixed(const void* data, size_t size,
                                                            const char* prefix);

// Why the last load on this thread failed ("" if it succeeded). Also set by voicedetect_capi_load.
const char* voicedetect_capi_last_load_error(void);

C++: vd::Model::load_from_memory(data, size, &err), vd::Model::load_from_memory(data, size, prefix, &err), and vd::ModelLoader::load_from_memory(...). vd::Model::load(path, &err) gets the same optional error out-parameter.

Ownership. The loader copies the tensor data into its own memory during the call. The buffer is only read inside the call. The caller can free or overwrite it as soon as the function returns, on success and on failure. While the call runs, the buffer and the model are in memory together.

Prefixed loading. The caller passes the whole bundle and the prefix (for example voice.). Only the tensors under the prefix are copied, so no standalone copy of the component is needed. Tensor names stored as string values in the metadata manifests are not prefixed and are used as they are. A prefix that matches no tensor is an error.

Threads. Loads are independent and may run on several threads. A context is still used by one thread at a time.

Robustness

The buffer comes from the caller, so a bad buffer must give an error:

  • A small pre-check walks the header, the metadata and the tensor table with bounds checks before ggml reads the buffer. ggml aborts the process on some malformed input (an empty metadata key, for example), so this check turns it into an error.
  • Every tensor range is checked against the buffer size without integer overflow,.
  • Metadata values are type-checked before they are read (ggml aborts on a wrong type).
  • ggml's own no_alloc=false memory reader is not used. It sizes its allocation from the tensor table before checking it against the buffer, so a hostile table can abort. The loader parses with no_alloc=true, validates, and copies the tensors itself.

Tests

  • test_load_memory (no model files, runs in CI): builds GGUF files and bundles in memory. Checks that the path, memory and prefixed loads give the same config and tensors, that the other component of a bundle is not loaded, that the buffer is poisoned and freed after the call, NULL, empty and wrong-prefix input, a load at every truncation length, corrupt headers and tensor tables, an empty key name, a wrong key type, 6000 seeded random bit flips, the C API, and concurrent loads.
  • test_load_memory_embed (needs VOICEDETECT_TEST_GGUF and VOICEDETECT_TEST_AUDIO, skips otherwise): embeddings from the path loader, the memory loader, the prefixed loader (inside a bundle next to another model) and the C API are bitwise identical. The buffer is poisoned and freed before the first embedding.

Verification

Run on Linux x86_64 (CPU, GCC):

  • ctest -LE model (CI command): 3 passed, 1 skipped (test_capi_dim, needs a model), 0 failed.
  • AddressSanitizer + UndefinedBehaviorSanitizer build: test_load_memory clean (leak detection on), test_load_memory_embed clean with WeSpeaker ResNet34, ECAPA-TDNN, CAM++ and ERes2Net f32 models. (The leak check is off for the embedding runs: the existing direct-conv state caches are not freed at exit, which is not related to this change.)
  • Bitwise-identical embeddings for those four models.

Not run: macOS and Windows. The code uses only standard C++ and ggml's memory reader and does not use a file, a file descriptor or a platform API, so there is nothing platform-specific to test; CI here runs on Linux only.

Limits

  • The buffer must hold the complete GGUF (no streaming).
  • Peak memory during the call is the buffer plus the model.
  • The wav2vec2 analyze models load the same way; only the speaker encoders were run end to end.

🤖 Generated with Claude Code

A model could only be loaded from a file path. A host that already holds
the GGUF in memory (for example one component of a larger bundle file)
had to write a copy to disk first.

Add voicedetect_capi_load_from_memory(data, size) and
voicedetect_capi_load_from_memory_prefixed(data, size, prefix), plus the
C++ equivalents vd::Model::load_from_memory and
vd::ModelLoader::load_from_memory. The prefixed form reads one component
of a bundle in place: every key and tensor name of the model carries the
prefix, and only the tensors under it are copied.

The loader copies the tensor data during the call, so the caller may free
the buffer when the call returns. The buffer is parsed with ggml's memory
reader after a structural pre-check (ggml aborts on some malformed
input, such as an empty metadata key), and every tensor range is checked
against the buffer size, so a truncated or corrupt buffer gives an error
and never reads out of bounds. Metadata values are type-checked before
they are read for the same reason.

voicedetect_capi_last_load_error() returns the reason for the last failed
load on the calling thread; the path loader sets it too. The path loader
is otherwise unchanged and shares the config parsing with the new loaders.

Tests build GGUF files and bundles in memory, so they run in CI without
model files. A second test compares path, memory and prefixed loads of a
real model bitwise.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit bca46bc into master Oct 4, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants