Repository navigation
docs: the Redux trim regression was noise (enlarged measurement), keep the 0.3 default - #95
Merged
Merged
Conversation
The first trim run (PR 94) read a small Redux loss on talks and on pink noise at 0 dB. It was repeated on 9 talks and about 375 noisy files with paired bootstrap intervals: no interval for talks or noisy speech excludes zero for the Ultra head, the Redux head or Silero with TDT v3. The earlier sets were too small (one word is 0.23 points) and Redux changes its text for tiny pad shifts. Correct docs/vad-benchmarks.md and docs/vad.md, keep the first numbers labelled as superseded, add the follow-up tables, and commit the compact scripts and results. No code changes. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Docs only. No code changed.
What was wrong
PR #94 made the VAD segment trim (0.3 s) the default. Its docs said the Redux head got slightly
worse with it: talks 4.39 to 4.51 WER, and speech in noise at pink 0 dB 12.35 to 13.29. Those sets
were small (3 talks, and 429 words per noisy condition, where one word is 0.23 points and the 0.93
was four words), so the claim was not supported.
What was measured now
The comparison was repeated on 9 talks (21,540 words) and about 375 noisy LibriSpeech files, with a
paired bootstrap 95 percent interval. Delta is trim 0.3 minus the old cuts, in WER points (negative
means trim is better):
No interval for talks or noisy speech excludes zero. Redux is the most volatile model: trim 0.3
changes the text of 15.5 percent of its talk segments (Ultra 10.0, v3 7.6), and moving the pad from
0.30 to 0.32 s changes 13.5 percent of them. Other paddings (0.5/0.5, 0.5 before and 0.3 after,
trimming only edges of at least 0.5 s) are within the intervals of 0.3. There is also a real
benefit: with Silero, the old cuts drop whole sentences on clean speech when a long segment starts
with silence (held-out clean WER 6.37 to 3.19 with trim, interval -7.76..-0.14).
The default stays 0.3 s.
What changed
docs/vad-benchmarks.md: the enlarged measurement and its limits replace the regression claim; thefirst numbers stay, labelled as small sets, superseded.
docs/vad.md: the trim section no longer says the trim costs about 0.1 WER point.scripts/vad_bench/decoder_guards/README.md: a note thatresults/tables.txtis the first run;results/trim_regression_followup.mdholds the follow-up tables.tables.txtis unchanged.scripts/vad_bench/trim_regression/: the compact scripts, results and a README on how toregenerate (about 260 KB, no audio, no caches).
throwaway_cli_patch.diffis for measurement onlyand is not part of the product.
Limits
Synthetic noise over read speech; the models differ from those of PR 94 (Redux packed, Ultra Q8_0,
v3 Q8_0 with Silero F16, against F16 files); the truth for speech start and end is the utterance
span, which is loose; only 5 talks are held out; no correction for multiple comparisons.
Not done
No code, flag or option was added. A minimum edge length for trimming (
trim_min_sec) wasconsidered and not added.
🤖 Generated with Claude Code