Skip to content

fix: Silero VAD must not leave the global backend at one thread - #91

Merged
mudler merged 1 commit into
masterfrom
fix/silero-vad-thread-count
Oct 4, 2026
Merged

mudler merged 1 commit into
masterfrom
fix/silero-vad-thread-count

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

Problem

Silero VAD computes each 32 ms chunk with run_graph(0, 1, ...). When --threads is not given, run_graph wrote that count of 1 to the process-global backend and kept it. Every later graph then ran on one thread, including the ASR decode after a VAD pass. This affects parakeet-cli transcribe --vad --vad-model silero.gguf without --threads, and any host that does not set a thread count.

Fix

A per-call n_threads in run_graph now applies to that call only: the previous backend count is restored after the compute. Silero still runs its tiny graphs on one thread. A positive global override (--threads) still wins, and leased pool backends are unaffected. pk::backend_thread_count() is added so tests can read the count.

The only other caller that passes a positive count on the global path and is not a test is the CTC head (4 threads). It had the same side effect and is now scoped too. On a 67 s clip with the 110M hybrid model and --decoder ctc, wall time was 0.91 to 0.95 s before and 0.94 to 0.98 s after, which is within the noise of that machine.

Measurement

Build: CPU only, Release. Model: 110M hybrid, Q8_0. Input: a 66.9 s speech clip (a 7.4 s LibriSpeech utterance repeated). Silero model: F16. Wall time of the whole transcribe command, 3 runs interleaved. The machine was not quiet (load average about 9 to 11), so these are not measurements from a quiet machine; the gap between before and after is much larger than the noise.

Backend thread count after the Silero pass, --threads unset: 1 before the fix, 8 after.

Run Before (s, RTF) After (s, RTF)
no VAD 1.05, 1.10, 1.13 (0.016 to 0.017) 0.97, 1.00, 1.16 (0.014 to 0.017)
--vad, no --threads 2.78, 2.78, 2.83 (0.042) 0.79, 0.85, 0.91 (0.012 to 0.014)
--vad --threads 8 1.00, 0.98, 0.94 (0.015) 0.82, 0.94, 0.95 (0.012 to 0.014)

Before the fix, --vad without --threads took about 2.7 times as long as the same run with --threads 8, and the process stayed at about 100 percent CPU. After the fix the two match.

Tests

  • New test_run_graph_threads (no model needed): checks the backend count across per-call counts and, when PARAKEET_TEST_SILERO_GGUF is set, across a Silero pass. It fails on the old code (3 failures) and passes now.
  • ctest -LE model: 35 of 35 pass.
  • test_silero_vad, test_silero_load_negative, test_silero_framer, test_capi_vad_silero, test_transcribe_vad_silero, test_vad_options: pass.

Limits

  • The Ultra and Redux head VAD tests were not run: those models were not available. The head path computes through the fused encoder with a count of 0 and does not use a per-call count, so it does not trigger this bug.
  • Only CPU was measured.

🤖 Generated with Claude Code

Silero VAD runs each 32 ms chunk with run_graph(0, 1, ...). With no
--threads override, run_graph wrote that count to the process-global
backend and kept it, so every later graph, including the ASR decode
after a VAD pass, ran on one thread.

Restore the previous count after the compute. A global override still
wins, and pooled backends are unaffected. Add pk::backend_thread_count()
for tests and a regression test that checks the count across a plain
graph and a Silero pass.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 91b120b into master Oct 4, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants