Skip to content

docs: the Redux trim regression was noise (enlarged measurement), keep the 0.3 default - #95

Merged
mudler merged 1 commit into
masterfrom
docs/trim-regression-correction
Oct 5, 2026
Merged

mudler merged 1 commit into
masterfrom
docs/trim-regression-correction

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

Docs only. No code changed.

What was wrong

PR #94 made the VAD segment trim (0.3 s) the default. Its docs said the Redux head got slightly
worse with it: talks 4.39 to 4.51 WER, and speech in noise at pink 0 dB 12.35 to 13.29. Those sets
were small (3 talks, and 429 words per noisy condition, where one word is 0.23 points and the 0.93
was four words), so the claim was not supported.

What was measured now

The comparison was repeated on 9 talks (21,540 words) and about 375 noisy LibriSpeech files, with a
paired bootstrap 95 percent interval. Delta is trim 0.3 minus the old cuts, in WER points (negative
means trim is better):

Set Redux head Ultra head v3 + Silero
9 talks +0.08 (-0.03..+0.20) +0.01 (-0.06..+0.08) +0.01 (-0.06..+0.10)
5 held-out talks +0.16 (-0.01..+0.36) +0.02 (-0.06..+0.11) +0.03 (-0.08..+0.20)
noisy speech (white 5, pink 5, pink 0) -0.12 (-0.57..+0.45) -0.27 (-0.57..+0.01) -0.26 (-0.57..+0.02)
pink 0 dB alone, Redux -0.15 (-0.90..+0.71)

No interval for talks or noisy speech excludes zero. Redux is the most volatile model: trim 0.3
changes the text of 15.5 percent of its talk segments (Ultra 10.0, v3 7.6), and moving the pad from
0.30 to 0.32 s changes 13.5 percent of them. Other paddings (0.5/0.5, 0.5 before and 0.3 after,
trimming only edges of at least 0.5 s) are within the intervals of 0.3. There is also a real
benefit: with Silero, the old cuts drop whole sentences on clean speech when a long segment starts
with silence (held-out clean WER 6.37 to 3.19 with trim, interval -7.76..-0.14).

The default stays 0.3 s.

What changed

  • docs/vad-benchmarks.md: the enlarged measurement and its limits replace the regression claim; the
    first numbers stay, labelled as small sets, superseded.
  • docs/vad.md: the trim section no longer says the trim costs about 0.1 WER point.
  • scripts/vad_bench/decoder_guards/README.md: a note that results/tables.txt is the first run;
    results/trim_regression_followup.md holds the follow-up tables. tables.txt is unchanged.
  • scripts/vad_bench/trim_regression/: the compact scripts, results and a README on how to
    regenerate (about 260 KB, no audio, no caches). throwaway_cli_patch.diff is for measurement only
    and is not part of the product.

Limits

Synthetic noise over read speech; the models differ from those of PR 94 (Redux packed, Ultra Q8_0,
v3 Q8_0 with Silero F16, against F16 files); the truth for speech start and end is the utterance
span, which is loose; only 5 talks are held out; no correction for multiple comparisons.

Not done

No code, flag or option was added. A minimum edge length for trimming (trim_min_sec) was
considered and not added.

🤖 Generated with Claude Code

The first trim run (PR 94) read a small Redux loss on talks and on
pink noise at 0 dB. It was repeated on 9 talks and about 375 noisy
files with paired bootstrap intervals: no interval for talks or noisy
speech excludes zero for the Ultra head, the Redux head or Silero with
TDT v3. The earlier sets were too small (one word is 0.23 points) and
Redux changes its text for tiny pad shifts.

Correct docs/vad-benchmarks.md and docs/vad.md, keep the first numbers
labelled as superseded, add the follow-up tables, and commit the
compact scripts and results. No code changes.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit da8bb4c into master Oct 5, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants