Repository navigation
fix: Silero VAD must not leave the global backend at one thread - #91
Merged
Merged
Conversation
Silero VAD runs each 32 ms chunk with run_graph(0, 1, ...). With no --threads override, run_graph wrote that count to the process-global backend and kept it, so every later graph, including the ASR decode after a VAD pass, ran on one thread. Restore the previous count after the compute. A global override still wins, and pooled backends are unaffected. Add pk::backend_thread_count() for tests and a regression test that checks the count across a plain graph and a Silero pass. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
mudler
approved these changes
Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Silero VAD computes each 32 ms chunk with
run_graph(0, 1, ...). When--threadsis not given,run_graphwrote that count of 1 to the process-global backend and kept it. Every later graph then ran on one thread, including the ASR decode after a VAD pass. This affectsparakeet-cli transcribe --vad --vad-model silero.ggufwithout--threads, and any host that does not set a thread count.Fix
A per-call
n_threadsinrun_graphnow applies to that call only: the previous backend count is restored after the compute. Silero still runs its tiny graphs on one thread. A positive global override (--threads) still wins, and leased pool backends are unaffected.pk::backend_thread_count()is added so tests can read the count.The only other caller that passes a positive count on the global path and is not a test is the CTC head (4 threads). It had the same side effect and is now scoped too. On a 67 s clip with the 110M hybrid model and
--decoder ctc, wall time was 0.91 to 0.95 s before and 0.94 to 0.98 s after, which is within the noise of that machine.Measurement
Build: CPU only, Release. Model: 110M hybrid, Q8_0. Input: a 66.9 s speech clip (a 7.4 s LibriSpeech utterance repeated). Silero model: F16. Wall time of the whole
transcribecommand, 3 runs interleaved. The machine was not quiet (load average about 9 to 11), so these are not measurements from a quiet machine; the gap between before and after is much larger than the noise.Backend thread count after the Silero pass,
--threadsunset: 1 before the fix, 8 after.--vad, no--threads--vad --threads 8Before the fix,
--vadwithout--threadstook about 2.7 times as long as the same run with--threads 8, and the process stayed at about 100 percent CPU. After the fix the two match.Tests
test_run_graph_threads(no model needed): checks the backend count across per-call counts and, whenPARAKEET_TEST_SILERO_GGUFis set, across a Silero pass. It fails on the old code (3 failures) and passes now.ctest -LE model: 35 of 35 pass.test_silero_vad,test_silero_load_negative,test_silero_framer,test_capi_vad_silero,test_transcribe_vad_silero,test_vad_options: pass.Limits
🤖 Generated with Claude Code