Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
78 changes: 74 additions & 4 deletions docs/vad-benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -952,7 +952,76 @@ claimed.
"Seconds decoded" is the overlap of the cuts with the block, from `vad --mode segments`.
- Noise alone: 63 files of 30 s (seven noise types, three levels), decoded whole.

### Trimming: word error rate
### Trimming: word error rate, enlarged measurement

The first run (below) had small sets and read a small loss for the Redux head as real. It was
repeated on a larger set, with the same trim of 0.3 s against the old cuts (`--vad-trim 0`):

- 9 whole TED-LIUM talks (21,540 reference words), 4 used to tune and 5 held out.
- About 375 noisy LibriSpeech files (6 utterances with gaps each, white and pink noise at 20, 10, 5
and 0 dB SNR, 3 noise seeds, split by speaker into tune and held out). The rows below use white 5
dB, pink 5 dB and pink 0 dB (15,642 reference words for Ultra and v3, 17,946 for Redux).
- Models: Redux packed, Ultra Q8_0, and TDT 0.6B v3 Q8_0 with Silero F16. The first run used F16
files, so the two runs differ in the models as well as in the size.
- The delta is trim 0.3 minus the old cuts, in WER points (negative means trim is better), with a
95 percent interval from a paired bootstrap over talks and utterances.

| Set | Redux head | Ultra head | v3 + Silero |
| --- | ---: | ---: | ---: |
| 9 talks | +0.08 (-0.03..+0.20) | +0.01 (-0.06..+0.08) | +0.01 (-0.06..+0.10) |
| 5 held-out talks | +0.16 (-0.01..+0.36) | +0.02 (-0.06..+0.11) | +0.03 (-0.08..+0.20) |
| 4 tune talks | -0.01 (-0.14..+0.13) | +0.00 (-0.10..+0.12) | -0.02 (-0.09..+0.04) |
| noisy speech (white 5, pink 5, pink 0) | -0.12 (-0.57..+0.45) | -0.27 (-0.57..+0.01) | -0.26 (-0.57..+0.02) |
| pink 0 dB alone (5,982 words) | -0.15 (-0.90..+0.71) | | |

No interval for talks or for noisy speech excludes zero. The Redux loss of the first run (4.39 to
4.51 on talks, 12.35 to 13.29 on pink 0 dB) does not hold up: on 9 talks it is +0.08, and on pink 0
dB it has the other sign. The first run looked worse for two reasons. It was small: its four noisy
sets had 429 words each, so one word is 0.23 points and the +0.93 was four words. And Redux is the
most volatile of the three: trim 0.3 changes the text of 15.5 percent of its talk segments, against
10.0 percent for Ultra and 7.6 percent for v3, and moving the pad from 0.30 to 0.32 s (a change that
cannot matter) changes the text of 13.5 percent of them and the talk WER by -0.05 (0.35 s: -0.07,
16.6 percent). That is the noise floor of this comparison. Of the +18 net errors that Redux gains on
talks, none are at the cut edges (0 net) and 18 are in the interior of the segments, where the
decoder sees a slightly different input; only 1 of 21,459 words is cut away.

Redux with other trims, against the old cuts (talks / noisy speech with white and pink noise at
four levels, 47,856 words), is not monotonic in the trim: 0.1 gives +0.03 / -0.01, 0.2 gives +0.02 /
-0.11, 0.3 gives +0.08 / -0.12, 0.5 gives -0.02 / +0.06, and 1.0 gives -0.00 / -0.01. Three other
rules (0.5 s before and after, 0.5 s before and 0.3 s after, and trimming only the edges of at least
0.5 s) are within the intervals of 0.3 on every set, and none is better than the default. Padding is
nearly free in noise: with trim 1.0 the decoder still gets only 5.7 s (Redux), 13.3 s (Ultra) and
0.6 s (v3) of a 60 s noise block, against 4.6, 12.1 and 0.0 s with 0.3 (old cuts: 33.2, 37.2, 27.0).
The 0.3 pad cuts away few words: on noisy speech 14 correct words of 47.8k (Redux), 14 of 16.1k
(Ultra) and 19 of 15.5k (v3), counted as words whose time lies outside the kept audio (a lower
bound). The median lateness of the detected speech start against the true start is -5 ms for
Redux, 79 ms for Ultra and 237 ms for Silero.

Trim also has a real benefit that the first run could not see. With Silero the old cuts drop whole
sentences on clean speech when a long segment starts with silence: on the held-out clean files the v3
WER goes from 6.37 to 3.19 with trim 0.3 (-3.19, interval -7.76..-0.14). That interval and the
Redux clean interval (-0.37, -0.80..-0.07 on held-out clean) are the only ones at 0.3 that exclude
zero, and both favour the trim.

Decision: the default stays 0.3 s. No code or option changed with this measurement; a minimum edge
length for trimming (`trim_min_sec`) was considered and not added.

Limits:

- The noise is synthetic and added to read speech. Real room noise and overlapping speech were not
tested.
- The models are not the ones of the first run (see above).
- The truth for where speech starts and ends is loose: it is the span of each utterance in the
synthetic file, so a pause inside an utterance counts as speech.
- Only 5 talks are held out, and the clean sets are small (1,350 words held out).
- There is no correction for the many comparisons in the tables. Read a single interval as a
range, not as a test.

Tables, the scripts and a note on how to regenerate them:
[scripts/vad_bench/trim_regression](../scripts/vad_bench/trim_regression/README.md) and
[results/trim_regression_followup.md](../scripts/vad_bench/decoder_guards/results/trim_regression_followup.md).

### Trimming: word error rate, first run (small sets, superseded)

| Set | Detector | Old | Trim 0.3 | Change |
| --- | --- | ---: | ---: | ---: |
Expand All @@ -963,9 +1032,10 @@ claimed.
| white noise 5 dB | Ultra / Redux / v3 + Silero | 4.90 / 7.23 / 5.59 | 4.43 / 6.76 / 5.13 | -0.47 / -0.47 / -0.47 |
| pink noise 0 dB | Ultra / Redux / v3 + Silero | 6.53 / 12.35 / 7.93 | 5.59 / 13.29 / 7.69 | -0.93 / +0.93 / -0.23 |

On talks the change is within 0.12 points; the Redux head loses a little on two of the three talks
and on pink noise. The speech-in-noise sets are small (one word is 0.23 points), so read them as
"neutral", not as a gain.
These are the numbers of the first run, kept for the record. The text that went with them said the
Redux head loses a little on talks and on pink noise. The enlarged measurement above shows that
this was noise: the sets are too small to tell (one word is 0.23 points), so read them as
"neutral".

### Trimming: the noise block

Expand Down
10 changes: 7 additions & 3 deletions docs/vad.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,9 +124,13 @@ whole file. `trim` 0 (`--vad-trim 0`) gives the previous cuts exactly.
This is a change of default behaviour for `transcribe --vad`,
`parakeet_capi_transcribe_path_json_vad*` and the `segments` mode of the VAD
functions, for the head, for Silero and for VAD-only slices (they share the
segmenter). On talks, transcripts of long audio can shift slightly (a word WER
cost of about 0.1 point in our runs); on audio with long noisy stretches the
decoder sees much less noise. Numbers: [vad-benchmarks.md](vad-benchmarks.md#trimming-segments-and-the-word-filter).
segmenter). Transcripts of long audio can shift slightly. An enlarged measurement
(9 talks and about 375 noisy files) found no change in word error rate whose
interval excludes zero, for the Ultra head, the Redux head or Silero with TDT v3;
an earlier small run that read a loss for the Redux head was noise. On audio with
long noisy stretches the decoder sees much less noise, and with Silero the trim
stops whole sentences from being dropped on clean speech. Numbers:
[vad-benchmarks.md](vad-benchmarks.md#trimming-segments-and-the-word-filter).

## Word filter (opt-in)

Expand Down
6 changes: 6 additions & 0 deletions scripts/vad_bench/decoder_guards/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,3 +36,9 @@ What the files are:
`--vad-trim 0` gives the old output byte for byte, the WER with the old cuts, with trim 0.3 and with
trim 0.3 plus `--min-local-conf 0.5`, the seconds of the noise block that the decoder gets, and the
words it returns inside the block.

Note on the trim results: `results/tables.txt` is the output of the first run, which has small
sets. Its Redux rows (talks 4.39 to 4.51, pink noise 0 dB 12.35 to 13.29) read as a small loss for
the Redux head. An enlarged measurement showed that this was noise (no interval excludes zero); the
file is kept unchanged. See `results/trim_regression_followup.md` and
[../trim_regression](../trim_regression/README.md).
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
# Trim regression follow-up: tables

Tables behind the section "Trimming: word error rate, enlarged measurement" of
[docs/vad-benchmarks.md](../../../../docs/vad-benchmarks.md). Every table is copied from a file in
[../../trim_regression/results](../../trim_regression/results) (named in each heading); the scripts
and the way to regenerate them are in [../../trim_regression](../../trim_regression/README.md).
`results/tables.txt` in this directory is the first run (small sets) and is not changed. Its Redux
rows (talks 4.39 to 4.51, pink noise 0 dB 12.35 to 13.29) were noise.

Deltas are trim 0.3 minus old cuts (`--vad-trim 0`) in WER points, with a 95 percent paired
bootstrap interval. T0 is the old cuts, Tx is trim x seconds.

## Trim 0.3 against the old cuts (sdi.txt, grid_redux_*.md, cand_*_heldout.md, cand_*_tune.md, grid_*_talks_all.txt)

| Set | Redux head | Ultra head | v3 + Silero |
| --- | ---: | ---: | ---: |
| 9 talks (21,540 words) | +0.08 (-0.03..+0.20) | +0.01 (-0.06..+0.08) | +0.01 (-0.06..+0.10) |
| 5 held-out talks (11,490 words) | +0.16 (-0.01..+0.36) | +0.02 (-0.06..+0.11) | +0.03 (-0.08..+0.20) |
| 4 tune talks (10,050 words) | -0.01 (-0.14..+0.13) | +0.00 (-0.10..+0.12) | -0.02 (-0.09..+0.04) |
| noisy speech: white 5, pink 5, pink 0 (Redux 17,946 words, others 15,642) | -0.12 (-0.57..+0.45) | -0.27 (-0.57..+0.01) | -0.26 (-0.57..+0.02) |
| pink 0 dB alone, Redux (5,982 words) | -0.15 (-0.90..+0.71) | | |
| clean speech, held out (1,350 words) | -0.37 (-0.80..-0.07) | -0.22 (-0.76..+0.28) | -3.19 (-7.76..-0.14) |

Held-out clean WER for v3 + Silero: 6.37 with the old cuts, 3.19 with trim 0.3.

## Redux, trim variants against the old cuts (grid_redux_all.md, 9 talks and all noisy files)

noisy-all is white and pink noise at 20, 10, 5 and 0 dB (47,856 words).

| Variant | talks | noisy-all |
| --- | ---: | ---: |
| T0.1 | +0.03 (-0.12..+0.18) | -0.01 (-0.22..+0.22) |
| T0.2 | +0.02 (-0.13..+0.19) | -0.11 (-0.40..+0.16) |
| T0.3 | +0.08 (-0.03..+0.20) | -0.12 (-0.42..+0.17) |
| T0.5 | -0.02 (-0.08..+0.04) | +0.06 (-0.13..+0.25) |
| T1.0 | -0.00 (-0.01..+0.00) | -0.01 (-0.13..+0.12) |

## Text changes when the trim changes (segchange.txt)

Share of segments whose text changes, and the net change in errors.

| Detector | T0 to T0.3, talks | T0 to T0.3, noisy | T0.3 to T0.5, talks |
| --- | ---: | ---: | ---: |
| Redux | 46 of 296 (15.5%), net +17 | 439 of 860 (51.0%), net -56 | 49 of 296 (16.6%), net -22 |
| Ultra | 30 of 299 (10.0%), net +1 | 149 of 291 (51.2%), net -40 | 29 of 299 (9.7%), net +0 |
| v3 + Silero | 21 of 276 (7.6%), net +2 | 151 of 279 (54.1%), net -32 | 20 of 276 (7.2%), net -3 |

Redux, pad 0.30 to 0.32 s: 40 of 296 segments (13.5%) change text; 0.30 to 0.35 s: 49 of 296
(16.6%). On the 9 talks (weighted from the tune and held-out sets of cand_redux_*.md) the talk WER
is 4.87 with 0.30, 4.82 with 0.32 and 4.80 with 0.35, that is -0.05 and -0.07.

## Where the errors move, Redux, T0 to T0.3 (boundary_T0_T0.3.txt)

Errors by zone of the segment (substitutions, deletions, insertions).

| Set | cut away | start edge | end edge | interior |
| --- | ---: | ---: | ---: | ---: |
| talks | 2 to 1 | 18 to 18 | 21 to 21 | 992 to 1010 (+18) |
| noisy speech | 4 to 5 | 21 to 15 | 31 to 13 | 3208 to 3169 (-39) |

Words cut away on noisy speech that were correct, by trim 0.3 (clip_count.txt): Redux 14 of 47,791,
Ultra 14 of 16,133, v3 19 of 15,506.

## Noise block (noise_seconds_90_inserts.txt, noise_words_20_inserts.md)

Seconds of a 60 s noise block that the decoder gets (90 insert files), and invented words in the
block (20 insert files).

| Detector | T0 | T0.3 | T0.5 | T1.0 | words, T0 | words, T0.3 |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Redux | 33.2 | 4.6 | 4.9 | 5.7 | 0 | 0 |
| Ultra | 37.2 | 12.1 | 12.4 | 13.3 | 18 | 0 |
| v3 + Silero | 27.0 | 0.0 | 0.2 | 0.6 | 14 | 0 |

## Detected speech start against the true start (vad_edges.txt, all noisy files)

Median lateness: Redux -5 ms, Ultra 79 ms, Silero 237 ms.

## Other padding rules (per_detector_default_table.md)

Held-out WER (talks, clean and the three noisy conditions together) and the delta against T0.3:
no rule is better than 0.3 outside the intervals.

| Detector | no trim | 0.3 / 0.3 | 0.5 / 0.5 | 0.5 / 0.3 | 0.3 / 0.3, edges of at least 0.5 s only |
| --- | ---: | ---: | ---: | ---: | ---: |
| Redux | 6.34 | 6.40 | 6.32 | 6.33 | 6.31 |
| Ultra | 5.10 | 5.01 | 5.05 | 5.05 | 5.04 |
| v3 + Silero | 5.70 | 5.40 | 5.52 | 5.51 | 5.41 |
48 changes: 48 additions & 0 deletions scripts/vad_bench/trim_regression/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Trim regression follow-up

Scripts and result files behind the section "Trimming: word error rate, enlarged measurement" of
[docs/vad-benchmarks.md](../../../docs/vad-benchmarks.md). The first run of the trim change (see
[../decoder_guards](../decoder_guards/README.md)) read a small loss for the Redux head. This
follow-up repeated the comparison on 9 talks and about 375 noisy files with paired bootstrap
intervals and found that the loss was noise.

No audio, no model files, no probability files and no decode cache are committed. The tables that
the page quotes are in [../decoder_guards/results/trim_regression_followup.md](../decoder_guards/results/trim_regression_followup.md).

## What is here

| Path | What |
| --- | --- |
| `scripts/` | Python scripts. `common.py` holds the paths and the segmenter settings, `seg.py` a Python port of the segmenter (with pad before, pad after, minimum edge and minimum length options), `dec.py` the segment decode cache, `lib.py` the scoring and the paired bootstrap |
| `make_results.sh` | Rebuilds the tables in `results/` from the decode cache; it runs no ASR |
| `throwaway_cli_patch.diff` | A patch of `src/model.cpp` for measurement only. It is not part of the product. It adds three environment variables to `parakeet-cli` (dump the VAD probabilities, decode a given list of segments, print the words of each slice) |
| `results/` | The output of `make_results.sh`, plus `validate_cli_vs_harness.txt` |

Some tables (`results/sdi.txt`, `results/per_detector_default_table.md`) were made by one-off
commands on the same cache; their scripts were not kept, so `make_results.sh` does not rebuild
them. The other files in `results/` are rebuilt by it.

## How to regenerate

Environment variables: `PK_WORK` is the work directory (default: the current directory; it holds
`data/`, `probs/`, `cache/` and `tmp/`), `PK_CLI` is the patched `parakeet-cli`
(default `$PK_WORK/build/examples/cli/parakeet-cli`), `PK_GGUF` is a directory with
`ultra-q8_0.gguf`, `redux-keep.gguf` (the Redux head, packed), `tdt-0.6b-v3-q8_0.gguf` and
`silero-vad-f16.gguf`. Python packages: `numpy soundfile jiwer librosa datasets`.

```
git apply throwaway_cli_patch.diff # in a scratch checkout; build parakeet-cli from it
export PK_WORK=/path/to/work PK_CLI=/path/to/patched/parakeet-cli PK_GGUF=/path/to/models
python3 scripts/fetch.py $PK_WORK/data # TED-LIUM long-form talks and LibriSpeech test-clean (streamed)
python3 scripts/mkcorpus.py $PK_WORK/data # speech in noise sets, noise inserts, manifest.json (fixed seeds)
python3 scripts/probe.py 4 # VAD probabilities of every file, per detector
python3 scripts/run.py T0,T0.1,T0.2,T0.3,T0.5,T1.0 talk,sinr,insert # decode the segments of each variant (hours of CPU, resumes)
python3 scripts/validate.py # the harness against the patched CLI: segment bounds and text must match
sh make_results.sh # tables into $PK_WORK/results
rm -r $PK_WORK/data # audio
```

`scripts/variants.py` defines the variants: `Tx` is trim x seconds, `T0` the old cuts, `Px/y` a
pad of x before and y after, `Ex` trim only edges of at least x seconds, `Lx` trim only segments of
at least x seconds. `results/validate_cli_vs_harness.txt` shows 30 of 30 runs with the same segment
bounds and the same text as the CLI.
23 changes: 23 additions & 0 deletions scripts/vad_bench/trim_regression/make_results.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
#!/bin/sh
# Regenerates every table in results/ from the decode cache (no ASR run).
# Run it from the work directory (or set PK_WORK); the scripts are next to this file.
# PYTHON is the interpreter (needs numpy, soundfile, jiwer).
HERE="$(cd "$(dirname "$0")" && pwd)"
export PK_WORK="${PK_WORK:-$(pwd)}"
P="${PYTHON:-python3}"; S="$HERE/scripts"; O="$PK_WORK/results"; mkdir -p "$O"
$P $S/report_grid.py redux all T0,T0.1,T0.2,T0.3,T0.5,T1.0 > $O/grid_redux_all.md
for s in tune heldout; do $P $S/report_grid.py redux $s T0,T0.1,T0.2,T0.3,T0.5,T1.0 > $O/grid_redux_$s.md; done
for d in ultra v3; do for s in tune heldout; do $P $S/report_cand.py $d $s T0,T0.3,T0.5,P0.5/0.3,E0.5 T0 white5,pink5,pink0 > $O/cand_${d}_$s.md; done; done
for s in tune heldout; do $P $S/report_cand.py redux $s T0,T0.1,T0.2,T0.3,T0.32,T0.35,T0.5,T1.0,P0.5/0.3,E0.5,E1.0,E0.5p0.5,L3 T0 white5,white0,pink5,pink0 > $O/cand_redux_$s.md; done
for d in ultra v3; do $P $S/table_trim.py T0,T0.1,T0.2,T0.3,T0.5,T1.0 all $d | grep talks > $O/grid_${d}_talks_all.txt; done
for d in redux ultra v3; do for k in talk sinr; do $P $S/boundary.py $d T0 T0.3 all $k; done; done > $O/boundary_T0_T0.3.txt
for d in redux ultra v3; do for p in "T0 T0.3" "T0.3 T0.5"; do $P $S/segchange.py $d $p talk,sinr; done; done > $O/segchange.txt 2>&1
$P $S/segchange.py redux T0.3 T0.32 talk >> $O/segchange.txt; $P $S/segchange.py redux T0.3 T0.35 talk >> $O/segchange.txt
$P $S/clip_count.py T0.1,T0.2,T0.3,T0.5,T1.0,P0.5/0.3 > $O/clip_count.txt 2>&1
$P $S/audio_clip.py T0.1,T0.2,T0.3,T0.4,T0.5,T1.0,P0.5/0.3,P0.5/0.4 ultra,redux,v3 > $O/audio_clip.txt 2>&1
$P $S/vad_edges.py > $O/vad_edges.txt 2>&1
$P $S/vad_cover.py > $O/vad_cover.txt 2>&1
$P $S/removed.py > $O/removed.txt 2>&1
$P $S/noise_ins.py T0,T0.1,T0.3,T0.5,T1.0,P0.5/0.3,E0.5,E1.0,E0.5p0.5,L3,L6 > $O/noise_seconds_90_inserts.txt 2>&1
$P $S/noise_words.py T0,T0.3,T0.5,P0.5/0.3,E0.5,L3 redux,ultra,v3 > $O/noise_words_20_inserts.md 2>&1
$P $S/pooled.py > $O/pooled.md
Loading
Loading