Repository navigation
Load a model from a memory buffer - #2
Merged
Merged
Conversation
A model could only be loaded from a file path. A host that already holds the GGUF in memory (for example one component of a larger bundle file) had to write a copy to disk first. Add voicedetect_capi_load_from_memory(data, size) and voicedetect_capi_load_from_memory_prefixed(data, size, prefix), plus the C++ equivalents vd::Model::load_from_memory and vd::ModelLoader::load_from_memory. The prefixed form reads one component of a bundle in place: every key and tensor name of the model carries the prefix, and only the tensors under it are copied. The loader copies the tensor data during the call, so the caller may free the buffer when the call returns. The buffer is parsed with ggml's memory reader after a structural pre-check (ggml aborts on some malformed input, such as an empty metadata key), and every tensor range is checked against the buffer size, so a truncated or corrupt buffer gives an error and never reads out of bounds. Metadata values are type-checked before they are read for the same reason. voicedetect_capi_last_load_error() returns the reason for the last failed load on the calling thread; the path loader sets it too. The path loader is otherwise unchanged and shares the config parsing with the new loaders. Tests build GGUF files and bundles in memory, so they run in CI without model files. A second test compares path, memory and prefixed loads of a real model bitwise. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why
A model can only be loaded from a file path today. A host that already has the GGUF in memory, for example one component of a larger bundle file that holds several models, has to write a copy to disk (or to an in-memory file that only exists on Linux) and pass that path. This PR adds loaders that take the model from memory. They need no file and no file descriptor, and they work the same on every platform.
This is additive. The path loader keeps its behaviour and shares the config parsing with the new loaders.
API
C++:
vd::Model::load_from_memory(data, size, &err),vd::Model::load_from_memory(data, size, prefix, &err), andvd::ModelLoader::load_from_memory(...).vd::Model::load(path, &err)gets the same optional error out-parameter.Ownership. The loader copies the tensor data into its own memory during the call. The buffer is only read inside the call. The caller can free or overwrite it as soon as the function returns, on success and on failure. While the call runs, the buffer and the model are in memory together.
Prefixed loading. The caller passes the whole bundle and the prefix (for example
voice.). Only the tensors under the prefix are copied, so no standalone copy of the component is needed. Tensor names stored as string values in the metadata manifests are not prefixed and are used as they are. A prefix that matches no tensor is an error.Threads. Loads are independent and may run on several threads. A context is still used by one thread at a time.
Robustness
The buffer comes from the caller, so a bad buffer must give an error:
no_alloc=falsememory reader is not used. It sizes its allocation from the tensor table before checking it against the buffer, so a hostile table can abort. The loader parses withno_alloc=true, validates, and copies the tensors itself.Tests
test_load_memory(no model files, runs in CI): builds GGUF files and bundles in memory. Checks that the path, memory and prefixed loads give the same config and tensors, that the other component of a bundle is not loaded, that the buffer is poisoned and freed after the call, NULL, empty and wrong-prefix input, a load at every truncation length, corrupt headers and tensor tables, an empty key name, a wrong key type, 6000 seeded random bit flips, the C API, and concurrent loads.test_load_memory_embed(needsVOICEDETECT_TEST_GGUFandVOICEDETECT_TEST_AUDIO, skips otherwise): embeddings from the path loader, the memory loader, the prefixed loader (inside a bundle next to another model) and the C API are bitwise identical. The buffer is poisoned and freed before the first embedding.Verification
Run on Linux x86_64 (CPU, GCC):
ctest -LE model(CI command): 3 passed, 1 skipped (test_capi_dim, needs a model), 0 failed.test_load_memoryclean (leak detection on),test_load_memory_embedclean with WeSpeaker ResNet34, ECAPA-TDNN, CAM++ and ERes2Net f32 models. (The leak check is off for the embedding runs: the existing direct-conv state caches are not freed at exit, which is not related to this change.)Not run: macOS and Windows. The code uses only standard C++ and ggml's memory reader and does not use a file, a file descriptor or a platform API, so there is nothing platform-specific to test; CI here runs on Linux only.
Limits
🤖 Generated with Claude Code