Skip to content

docs: why the VAD head fires on noise (root-cause study, mitigations) and a Silero plus head fusion experiment - #89

Merged
mudler merged 2 commits into
masterfrom
docs/vad-head-noise-and-fusion
Oct 4, 2026
Merged

mudler merged 2 commits into
masterfrom
docs/vad-head-noise-and-fusion

Conversation

@localai-org-maint-bot

@localai-org-maint-bot localai-org-maint-bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

No library code changed. This PR adds documentation, small scripts and small result files.

What it says

The Parakeet VAD head (Ultra and Redux) gives false alarms on audio that has no speech. Silero does not. Measured at threshold 0.5 on synthetic data (LibriSpeech test-clean with added white and pink noise):

  • On speech-free 30 s noise clips (96 clips), the head calls 99.4 percent (Ultra) and 97.8 percent (Redux) of the frames speech. Silero: 0.0 percent. The cause is studied below.
  • Over a 30 s noise stretch inside a file that has speech (96 files), the false-alarm frame rate was 17.7 percent for Ultra and 55 percent for Redux, and for Ultra it grows with the noise level (0.0, 0.6, 14.3 and 55.8 percent at 20, 10, 5 and 0 dB). Silero: 0 percent.

So: prefer Silero as the always-on gate, or when the audio can have long stretches without speech. Use the head on audio known to be mostly speech (for example before transcription of recorded talks), or where its higher recall matters. This is added to docs/vad.md, the "When to use which" section of docs/vad-benchmarks.md, and the README VAD section.

Root-cause study of the noise false alarms

Added to docs/vad.md and docs/vad-benchmarks.md, with scripts and small results in scripts/vad_bench/noise_dive/ (no audio, models or large files).

  • Not a bug in parakeet.cpp. A reference built from Hugging Face transformers and the documented head matches parakeet-cli vad --probabilities to a maximum of about 2.6e-4 on F16 files. The only real difference is digital silence.
  • It is a property of the model. The head scores each 80 ms frame by its level and texture relative to the file's own average. Steady loud noise, or a noise-only file, sits at logit +1.5 to +2 (p about 0.8 to 0.9); speech at +5 to +12; pauses at -3 to -8. Per-file normalisation is one input, not the whole cause. Amplitude invariance does not hold: the output is flat above about -40 dBFS and falls off below it. The cards say the head exists to cut recordings at pauses into segments of at most 30 s; what it was trained on is not stated.
  • A table of what the head does with digital silence, a tone, a sweep, clicks, hum, music-like chords and brown noise, and how far below the speech a noise stretch must be before it stops firing. Redux reacts to flat-spectrum noise; Ultra to level, and also to clicks and music. The earlier statement that false alarms grow with the noise level is corrected per model.
  • User-visible cost with transcribe --vad (60 s inserted block in a real talk): 0 hallucinated words without VAD in 24 runs; with VAD up to 15 words (Ultra, music), 2 (Redux, white at -20 dB), 20 for Silero on one file. A hard 30 s segment cut can keep most of a noise gap (54 to 60 s of a 60 s block) even where the head called none of it speech; this is a segmenter behaviour, not fixed.
  • Mitigations tested offline (thresholds, minimum speech, energy gate, median-logit run gate, Silero AND head, two-stage). Thresholds trade recall for fewer false alarms; the energy gate and minimum speech do not help; the median-logit run gate is the best cheap option but is not validated on WER and not implemented; two-stage needs Silero.
  • Proposed, not in this PR: a guard for digital silence or a constant tone. Defaults are unchanged.
  • Limits: the model's own runtime was not available (wiring inferred), synthetic noise plus one real talk, the mitigation numbers use a Python segmenter, hallucination counts come from one talk and one insertion point, and the black-box check against the released runtime was skipped.

Fusion experiment (offline)

A new section in docs/vad-benchmarks.md, "Fusing Silero and the head (offline experiment)", and the scripts in scripts/vad_bench/fusion/ (a Python port of the segmenter, with a README that says how to reproduce it and which inputs are missing). Fusion is not implemented in parakeet.cpp.

Setup: 342 speech clips (38 sets of 5 LibriSpeech test-clean utterances in 9 conditions: clean, and white and pink noise at 20, 10, 5 and 0 dB), plus 96 noise-only 30 s clips. Silero ONNX 6.2.3, Ultra Q8_0 and Redux packed heads, a 10 ms grid, the same post-processing for all rules, speaker-disjoint 4-fold cross-validation, 95 percent bootstrap intervals over clips.

Pooled, percent P R F1
Silero, threshold 0.5 99.6 85.6 92.1
Ultra head 98.6 89.4 93.8
Redux head 97.6 92.0 94.7
Two-stage (Redux) 98.9 91.2 94.9

After cross-validated tuning, F1 is 94.5 for Silero, 94.4 for the Ultra head, 94.8 for the Redux head, 95.3 for the two-stage rule with Ultra (+0.75) and 95.9 with Redux (+1.13, interval +1.00 to +1.28). At precision of at least 99 percent, recall is 91.5 for two-stage Redux, 89.4 for the Redux head and 88.6 for Silero.

In plain terms, per 100 s of real speech Silero misses 14.4 s, the Redux head 8.0 s and the two-stage rule 8.8 s; per 100 s that a detector calls speech, Silero is wrong for 0.4 s, the Redux head for 2.4 s and the two-stage rule for 1.1 s.

The two-stage rule: Silero decides what is speech. The head only moves boundaries outward (160 to 320 ms before a start, 160 ms after an end) and fills gaps shorter than 300 to 600 ms. It cannot add speech where Silero finds none, so it has no false alarms on noise. Running it costs about the sum of both detectors (about 80x real time with both).

Rules that did not help: OR, max, mean, switching on an estimated noise level, logistic regression and gradient boosting either inherit the head's noise false alarms or do not transfer (a learned model fitted without noise-only clips called about 90 percent of a 600 s talk speech, against 77.5 percent for Silero).

Limits

  • Synthetic read English speech with synthetic white and pink noise. Babble, music, reverberation and other languages were not tested.
  • Part of every tuned gain comes from the reference convention of the synthetic clips.
  • The word-time reference of the TED talks was lost, so the rules were not scored on real talks.
  • The results come from one corpus and one machine.

🤖 Generated with Claude Code

mudler added 2 commits October 4, 2026 11:25
… fusion experiment

Measured on synthetic LibriSpeech with added noise: on a speech-free file
the Parakeet VAD head calls about 99 percent of the frames speech, and in
a 30 s noise stretch inside a speech file it called 17.7 percent (Ultra)
and 55 percent (Redux) of the frames speech, against 0 percent for Silero.
Say so in docs/vad.md, docs/vad-benchmarks.md and the README, and prefer
Silero as the always-on gate.

Add the offline experiment that fuses Silero and the head, with the
scripts and small result files under scripts/vad_bench/fusion/. Fusion is
not implemented in parakeet.cpp; no code changed.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…igations

A reference built from Hugging Face transformers matches the CLI, so the
false alarms are a property of the model: the head scores each frame by
its level and texture relative to the file. Document the verified
mechanism, the noise types and levels that trigger it, the cost in
transcribe --vad, the mitigations tested and the limits. Add the study's
scripts and small result files under scripts/vad_bench/noise_dive.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@localai-org-maint-bot localai-org-maint-bot changed the title docs: the Parakeet VAD head false-alarms on noise-only audio; Silero plus head fusion experiment docs: why the VAD head fires on noise (root-cause study, mitigations) and a Silero plus head fusion experiment Oct 4, 2026
@mudler
mudler merged commit 781a973 into master Oct 4, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants