Repository navigation
docs: why the VAD head fires on noise (root-cause study, mitigations) and a Silero plus head fusion experiment - #89
Merged
Conversation
… fusion experiment Measured on synthetic LibriSpeech with added noise: on a speech-free file the Parakeet VAD head calls about 99 percent of the frames speech, and in a 30 s noise stretch inside a speech file it called 17.7 percent (Ultra) and 55 percent (Redux) of the frames speech, against 0 percent for Silero. Say so in docs/vad.md, docs/vad-benchmarks.md and the README, and prefer Silero as the always-on gate. Add the offline experiment that fuses Silero and the head, with the scripts and small result files under scripts/vad_bench/fusion/. Fusion is not implemented in parakeet.cpp; no code changed. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
…igations A reference built from Hugging Face transformers matches the CLI, so the false alarms are a property of the model: the head scores each frame by its level and texture relative to the file. Document the verified mechanism, the noise types and levels that trigger it, the cost in transcribe --vad, the mitigations tested and the limits. Add the study's scripts and small result files under scripts/vad_bench/noise_dive. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No library code changed. This PR adds documentation, small scripts and small result files.
What it says
The Parakeet VAD head (Ultra and Redux) gives false alarms on audio that has no speech. Silero does not. Measured at threshold 0.5 on synthetic data (LibriSpeech test-clean with added white and pink noise):
So: prefer Silero as the always-on gate, or when the audio can have long stretches without speech. Use the head on audio known to be mostly speech (for example before transcription of recorded talks), or where its higher recall matters. This is added to
docs/vad.md, the "When to use which" section ofdocs/vad-benchmarks.md, and the README VAD section.Root-cause study of the noise false alarms
Added to
docs/vad.mdanddocs/vad-benchmarks.md, with scripts and small results inscripts/vad_bench/noise_dive/(no audio, models or large files).parakeet-cli vad --probabilitiesto a maximum of about 2.6e-4 on F16 files. The only real difference is digital silence.transcribe --vad(60 s inserted block in a real talk): 0 hallucinated words without VAD in 24 runs; with VAD up to 15 words (Ultra, music), 2 (Redux, white at -20 dB), 20 for Silero on one file. A hard 30 s segment cut can keep most of a noise gap (54 to 60 s of a 60 s block) even where the head called none of it speech; this is a segmenter behaviour, not fixed.Fusion experiment (offline)
A new section in
docs/vad-benchmarks.md, "Fusing Silero and the head (offline experiment)", and the scripts inscripts/vad_bench/fusion/(a Python port of the segmenter, with a README that says how to reproduce it and which inputs are missing). Fusion is not implemented in parakeet.cpp.Setup: 342 speech clips (38 sets of 5 LibriSpeech test-clean utterances in 9 conditions: clean, and white and pink noise at 20, 10, 5 and 0 dB), plus 96 noise-only 30 s clips. Silero ONNX 6.2.3, Ultra Q8_0 and Redux packed heads, a 10 ms grid, the same post-processing for all rules, speaker-disjoint 4-fold cross-validation, 95 percent bootstrap intervals over clips.
After cross-validated tuning, F1 is 94.5 for Silero, 94.4 for the Ultra head, 94.8 for the Redux head, 95.3 for the two-stage rule with Ultra (+0.75) and 95.9 with Redux (+1.13, interval +1.00 to +1.28). At precision of at least 99 percent, recall is 91.5 for two-stage Redux, 89.4 for the Redux head and 88.6 for Silero.
In plain terms, per 100 s of real speech Silero misses 14.4 s, the Redux head 8.0 s and the two-stage rule 8.8 s; per 100 s that a detector calls speech, Silero is wrong for 0.4 s, the Redux head for 2.4 s and the two-stage rule for 1.1 s.
The two-stage rule: Silero decides what is speech. The head only moves boundaries outward (160 to 320 ms before a start, 160 ms after an end) and fills gaps shorter than 300 to 600 ms. It cannot add speech where Silero finds none, so it has no false alarms on noise. Running it costs about the sum of both detectors (about 80x real time with both).
Rules that did not help: OR, max, mean, switching on an estimated noise level, logistic regression and gradient boosting either inherit the head's noise false alarms or do not transfer (a learned model fitted without noise-only clips called about 90 percent of a 600 s talk speech, against 77.5 percent for Silero).
Limits
🤖 Generated with Claude Code