Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,14 @@ ternary ones.
functions, as a stream (`parakeet_capi_vad_stream_*`), and as the cutter for any ASR model:
`parakeet-cli transcribe --vad --vad-model silero.gguf`. See [`docs/vad.md`](docs/vad.md).
- VAD-only slices of Ultra and Redux (6 to 10 MB, `redux-vad.gguf` and `ultra-vad-q8_0.gguf` in the same repo) hold just the head and its front end. They run `vad` and the `parakeet_capi_vad_*` calls and cannot transcribe. See [`docs/vad.md`](docs/vad.md).
- The head alone gives false alarms on audio without speech: on speech-free noise it calls about
99 percent of the frames speech (Ultra 99.4, Redux 97.8), and over a 30 s noise stretch inside a
file with speech the false-alarm frame rate was 17.7 percent for Ultra and 55 percent for Redux,
against 0 percent for Silero (synthetic LibriSpeech with added noise). Prefer Silero as the
always-on gate or when the audio can have long non-speech stretches; use the head on audio known
to be mostly speech, or where its higher recall matters. An offline experiment that lets Silero
decide and the head move the edges (not implemented here) is in
[`docs/vad-benchmarks.md`](docs/vad-benchmarks.md#fusing-silero-and-the-head-offline-experiment).
- Accuracy, speed and size of both detectors, with the method and the scripts to repeat
them: [`docs/vad-benchmarks.md`](docs/vad-benchmarks.md).
---
Expand Down
390 changes: 388 additions & 2 deletions docs/vad-benchmarks.md

Large diffs are not rendered by default.

40 changes: 40 additions & 0 deletions docs/vad.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,46 @@ transcription segments. The segmenter takes the frame period as a parameter
Accuracy and speed numbers for both detectors, with the method and the scripts to
repeat them, are in [vad-benchmarks.md](vad-benchmarks.md).

## Which detector to use

The Parakeet VAD head alone gives false alarms on audio without speech. On a speech-free
file the head calls about 99 percent of the frames speech (Ultra 99.4 percent, Redux 97.8
percent). Inside a file that has speech, over a 30 s stretch of noise at threshold 0.5,
the false-alarm frame rate was 17.7 percent for the Ultra head and 55 percent for the
Redux head. Silero showed 0 percent in both tests. This was measured on synthetic data:
LibriSpeech with added white and pink noise, see
[the experiment](vad-benchmarks.md#the-head-on-noise-only-audio).

This is not a bug in parakeet.cpp. An independent reference built from Hugging Face
transformers and the documented head structure matches `parakeet-cli vad --probabilities`
to a maximum probability difference of 2e-4 on F16 files. It is a property of the model.
The head judges each 80 ms frame by its level and texture relative to the average of the
file it is in. Steady loud noise, or a file that is only noise, sits at a logit of about
+1.5 to +2 (probability 0.8 to 0.9). Speech sits at +5 to +12 and pauses at -3 to -8. The
cards of the model say the head exists so that recordings can be cut at pauses into
segments of at most 30 s. What it was trained on is not stated. It is not a noise or
music rejector.

So use Silero, or Silero first with the head extending the boundaries, whenever long
stretches without speech are possible (always-on gates, calls, recordings with music or
noise). Use the head alone on audio that is known to be mostly speech (for example
before the transcription of recorded talks), or where its higher recall matters. If you
use the head alone and want fewer false alarms, raising the threshold to 0.9 removes most
noise false alarms at a cost of 1.5 to 2.6 F1 points on speech. Details, the noise types
that do and do not trigger it, and the mitigations we tested are in
[the benchmark page](vad-benchmarks.md#the-head-on-noise-only-audio).

For `transcribe --vad` the user-visible cost is small: some wasted decoding, and rarely a
hallucinated phrase. The 30 s hard cut of the segmenter can also keep a large part of a
long noise gap, even where the head called none of it speech. This is a segmenter
behaviour, not a head false alarm, and it has not been fixed.

An offline experiment also combined the two: Silero decides what is speech, and the head
only moves the edges. It scored 0.75 to 1.13 F1 points above the best single detector on
the synthetic clips and gave no false alarms on noise. It is not implemented in
parakeet.cpp. See
[Fusing Silero and the head](vad-benchmarks.md#fusing-silero-and-the-head-offline-experiment).

## Standalone VAD API

One call returns the speech regions of a clip as JSON, for either detector. Load
Expand Down
70 changes: 70 additions & 0 deletions scripts/vad_bench/fusion/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# Silero plus head fusion experiment

Offline experiment behind the section "Fusing Silero and the head (offline experiment)"
in [docs/vad-benchmarks.md](../../../docs/vad-benchmarks.md). It is not part of
parakeet.cpp: the fusion rules are a Python port of the segmenter, and no C++ code
implements them.

No audio, no probability files and no pickles are committed. The four small result
files are:

| File | What |
| --- | --- |
| `results.md` | All cross-validated tables for the Ultra head and the Redux head (precision, recall, F1 with bootstrap intervals, false alarms on noise-only clips, paired differences, noise-shift split, operating points at precision 99 percent or more) |
| `gap_results.md` | False alarms inside a 30 s noise stretch that sits in a file with speech. It is the printed output of `gapeval.py` with a longer heading |
| `pr_points.csv` | Points of the precision and recall frontiers (the data of the plot that `pr_curves.py` draws) |
| `ted_results.txt` | Output of `ted.py` on a 600 s talk: share of speech and agreement with Silero and with the head |

## How it works

Every script runs from one working directory (the "work directory") and reads and writes
files there with relative names: `data/`, `clips/`, `clips.json`, `gap.json`,
`probs.npz`, `probs_gap.npz`, `results.pkl`. All scripts need `numpy`, `soundfile`;
`run.py` and `gapeval.py` also need `scikit-learn`; `collect*.py` need `onnxruntime`;
`prep_libri.py` needs `datasets`; `pr_curves.py` needs `matplotlib`. Run them with
the work directory as the current directory, for example `python3 /path/to/fusion/run.py`
after copying or linking the scripts into it (the scripts import each other by name).

The models are given by environment variables, read by `collect.py`, `collect_gap.py`
and `ted.py`:

| Variable | Value |
| --- | --- |
| `SILERO_ONNX` | `silero_vad.onnx` of Silero VAD 6.2.3 |
| `PARAKEET_CLI` | built `parakeet-cli` (it must support `vad --probabilities`) |
| `ULTRA_GGUF` | ASR GGUF of `moondream/parakeet-ultra`, Q8_0 |
| `REDUX_GGUF` | ASR GGUF of `moondream/parakeet-redux`, packed (`keep`) |

## Steps

```
python3 prep_libri.py # every 13th utterance of LibriSpeech test-clean -> data/libri.npy, data/spk.npy
python3 mkcorpus.py # 342 speech clips (38 sets of 5 utterances x 9 conditions) + 96 noise-only 30 s clips
python3 mkgap.py # 96 clips with a 30 s noise stretch between speech
python3 collect.py # per-frame probabilities of Silero and both heads -> probs.npz
python3 collect_gap.py # the same for the gap clips -> probs_gap.npz
python3 run.py # all rules, thresholds and the learned models -> results.pkl
python3 report.py # results.md (reads results.pkl)
python3 gapeval.py # prints the gap tables
python3 pr_curves.py # pr_curves.png and pr_points.csv
python3 detail.py # boundary error and missed segments -> detail.json
python3 ted.py talk.wav # speech share on one long talk (needs the clips and probs of the steps above)
```

The speaker folds are fixed by the seed in `mkcorpus.py`. `rules.py` holds the fusion
rules and the two-stage rule, `fl.py` the 10 ms grid, the post-processing and the
metrics, `analyze.py` the cross-validation and the bootstrap, `lib.py` the clip
builder and the noise.

## Missing inputs

- The TED talk used by `ted.py` is not part of this directory and is not fetched by
these scripts; `../fetch_data.py` fetches one TED-LIUM talk. The word-time reference of
the talk used in the main benchmark was lost, so the experiment could not score the
fusion rules against a TED reference. `ted.py` only reports how much of the talk each
rule calls speech and how well the rules agree with each other.
- `ted.py` fits its logistic regression on the synthetic clips, so it needs `clips.json`,
`clips/` and `probs.npz` from the steps above.
- `analyze.py` and `report.py` read `results.pkl` (about 50 MB), which is not committed.
Run `run.py` to make it. The random seeds are fixed, so a re-run should give the same tables
on the same probabilities, but this was not checked on a second machine.
69 changes: 69 additions & 0 deletions scripts/vad_bench/fusion/analyze.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
import pickle,json,sys,numpy as np
from fl import prf
R=pickle.load(open("results.pkl","rb")); CFG=R["cfg"]; LR=R["lr"]; M=R["meta"]
N=len(M["names"]); fold=np.array(M["fold"]); noz=np.array(M["noise_only"]); snr=np.array(M["snr_true"]); dur=np.array(M["dur"])
sp=~noz
LEVELS=[("clean",99),("20 dB",20),("10 dB",10),("5 dB",5),("0 dB",0)]
rng=np.random.default_rng(0); B=2000
def pooled(A,idx): c=A[idx,:4].sum(0); return prf(c)
def key(rule,p,h): return (rule,tuple(p),h,0.1)
def A_of(rule,p,h): return CFG[key(rule,p,"ultra" if rule=="silero" else h)]
HEADS=("ultra","redux")
# ---- candidate sets: dict name -> (list of params, array [ncfg,N,5]) ----
def cand(rule,h):
hh="ultra" if rule=="silero" else h
ps=[k[1] for k in CFG if k[0]==rule and k[2]==hh and k[3]==0.1]
# keep insertion order
ps=[tuple(float(x) for x in p) for p in dict.fromkeys(ps)]; return ps,np.stack([CFG[key(rule,p,hh)] for p in ps])
def lrcand(h,kind,split): o,c=LR[(h,kind,split)]; return [(float(t),) for t in __import__("rules").THR],o
def score(arr,idx,crit):
"""arr [ncfg,N,5]; returns best cfg index on clips idx"""
c=arr[:,idx,:4].sum(1); tp,fp,fn=c[:,0],c[:,1],c[:,2]
P=tp/np.maximum(1,tp+fp);Rr=tp/np.maximum(1,tp+fn);F=2*P*Rr/np.maximum(1e-12,P+Rr)
if crit=="f1": return int(np.argmax(F))
ok=P>=crit
return int(np.argmax(np.where(ok,Rr,-1+P*0)) if ok.any() else np.argmax(P))
def crossfit(getarr,crit,mode="speaker"):
"""returns (rows [N,5] with each clip from the cfg chosen without its own fold/shift side, chosen params per split)"""
rows=np.zeros((N,5)); chosen={}
if mode=="speaker":
for f in range(4):
ps,arr=getarr("f%d"%f); tr=np.flatnonzero(sp&(fold!=f)); te=np.flatnonzero(fold==f)
i=score(arr,tr,crit); rows[te]=arr[i][te]; chosen[f]=ps[i]
else:
ps,arr=getarr("shift"); tr=np.flatnonzero(sp&(snr>=10)); te=np.flatnonzero(snr<10)
i=score(arr,tr,crit); rows[te]=arr[i][te]; chosen["shift"]=ps[i]
return rows,chosen
def fixed(arr_ps,p): ps,arr=arr_ps; return arr[ps.index(tuple(p))]
def static(rule,h):
ps,arr=cand(rule,h); return lambda s:(ps,arr)
def lrget(h,kind): return lambda s:lrcand(h,kind,s)
# ---- bootstrap ----
def boot_idx(idx): return rng.integers(0,len(idx),(B,len(idx)))
def ci(vals): return np.percentile(vals,[2.5,97.5])
def prf_b(rows,idx,bi):
c=rows[idx][:,:4][bi].sum(1).astype(float) # B,4
P=c[:,0]/np.maximum(1,c[:,0]+c[:,1]);Rr=c[:,0]/np.maximum(1,c[:,0]+c[:,2]);F=2*P*Rr/np.maximum(1e-12,P+Rr)
return P,Rr,F
def fmt(x,lo,hi): return f"{100*x:.1f} [{100*lo:.1f},{100*hi:.1f}]"
BI={}
def bidx(name,idx):
if name not in BI: BI[name]=(idx,rng.integers(0,len(idx),(B,len(idx))))
return BI[name]
def level_idx(lv): return np.flatnonzero(sp&(snr==lv))
def stats(rows):
"""pooled + per level point estimates and CIs (paired resample indices fixed per index set)"""
out={}
for nm,idx in [("all",np.flatnonzero(sp))]+[(l,level_idx(v)) for l,v in LEVELS]:
idx,bi=bidx(nm,idx); P,Rr,F=prf_b(rows,idx,bi); p,r,f=pooled(rows,idx)
out[nm]=dict(P=p,R=r,F=f,Pci=ci(P),Rci=ci(Rr),Fci=ci(F),F_b=F)
return out
def fa(rows):
"""false-alarm segments per hour and speech-frame FP rate on noise-only clips, per noise level"""
o={}
for l,v in LEVELS[1:]:
idx=np.flatnonzero(noz&(snr==v)); o[l]=(rows[idx,4].sum()/(dur[idx].sum()/3600), rows[idx,1].sum()/max(1,rows[idx,:4].sum()))
idx=np.flatnonzero(noz); o["all"]=(rows[idx,4].sum()/(dur[idx].sum()/3600), rows[idx,1].sum()/rows[idx,:4].sum())
return o
if __name__=="__main__":
pass
27 changes: 27 additions & 0 deletions scripts/vad_bench/fusion/collect.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
import json,os,subprocess,numpy as np,soundfile as sf,onnxruntime as ort
from multiprocessing import Pool
# Paths come from the environment: SILERO_ONNX (silero_vad.onnx, 6.2.3), PARAKEET_CLI (parakeet-cli),
# ULTRA_GGUF and REDUX_GGUF (the ASR GGUFs of the two heads).
M=os.environ["SILERO_ONNX"]; CLI=os.environ["PARAKEET_CLI"]
HEADS={"ultra":os.environ["ULTRA_GGUF"],"redux":os.environ["REDUX_GGUF"]}
meta=json.load(open("clips.json"))
def sil(y,sess):
n=(len(y)+511)//512; y=np.pad(y,(0,n*512-len(y)))
st=np.zeros((2,1,128),np.float32); ctx=np.zeros(64,np.float32); out=[]
for i in range(n):
x=np.concatenate([ctx,y[i*512:(i+1)*512]])[None].astype(np.float32)
o,st=sess.run(None,{"input":x,"state":st,"sr":np.array(16000,np.int64)})
out.append(float(o[0,0])); ctx=x[0,-64:]
return np.array(out,np.float32)
def work(m):
so=ort.SessionOptions(); so.intra_op_num_threads=1; so.inter_op_num_threads=1
sess=ort.InferenceSession(M,so,providers=["CPUExecutionProvider"])
f=f"clips/{m['name']}.wav"; y=sf.read(f,dtype="float32")[0]; r={"silero":sil(y,sess)}
for k,g in HEADS.items():
j=json.loads(subprocess.run([CLI,"vad","--model",g,"--input",f,"--probabilities","--threads","1"],capture_output=True,text=True,check=True).stdout)
r[k]=np.array(j["probabilities"],np.float32)
return m["name"],r
if __name__=="__main__":
with Pool(4) as p: res=p.map(work,meta,chunksize=4)
flat={f"{n}|{k}":v for n,r in res for k,v in r.items()}
np.savez_compressed("probs.npz",**flat); print("done",len(res))
27 changes: 27 additions & 0 deletions scripts/vad_bench/fusion/collect_gap.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
import json,os,subprocess,numpy as np,soundfile as sf,onnxruntime as ort
from multiprocessing import Pool
# Paths come from the environment: SILERO_ONNX (silero_vad.onnx, 6.2.3), PARAKEET_CLI (parakeet-cli),
# ULTRA_GGUF and REDUX_GGUF (the ASR GGUFs of the two heads).
M=os.environ["SILERO_ONNX"]; CLI=os.environ["PARAKEET_CLI"]
HEADS={"ultra":os.environ["ULTRA_GGUF"],"redux":os.environ["REDUX_GGUF"]}
meta=json.load(open("gap.json"))
def sil(y,sess):
n=(len(y)+511)//512; y=np.pad(y,(0,n*512-len(y)))
st=np.zeros((2,1,128),np.float32); ctx=np.zeros(64,np.float32); out=[]
for i in range(n):
x=np.concatenate([ctx,y[i*512:(i+1)*512]])[None].astype(np.float32)
o,st=sess.run(None,{"input":x,"state":st,"sr":np.array(16000,np.int64)})
out.append(float(o[0,0])); ctx=x[0,-64:]
return np.array(out,np.float32)
def work(m):
so=ort.SessionOptions(); so.intra_op_num_threads=1; so.inter_op_num_threads=1
sess=ort.InferenceSession(M,so,providers=["CPUExecutionProvider"])
f=f"clips/{m['name']}.wav"; y=sf.read(f,dtype="float32")[0]; r={"silero":sil(y,sess)}
for k,g in HEADS.items():
j=json.loads(subprocess.run([CLI,"vad","--model",g,"--input",f,"--probabilities","--threads","1"],capture_output=True,text=True,check=True).stdout)
r[k]=np.array(j["probabilities"],np.float32)
return m["name"],r
if __name__=="__main__":
with Pool(4) as p: res=p.map(work,meta,chunksize=4)
flat={f"{n}|{k}":v for n,r in res for k,v in r.items()}
np.savez_compressed("probs_gap.npz",**flat); print("done",len(res))
35 changes: 35 additions & 0 deletions scripts/vad_bench/fusion/detail.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# Boundary error / missed segments / TED agreement for the cross-fitted systems
import json,numpy as np,soundfile as sf,subprocess
import run
from run import *
from analyze import crossfit,static,lrget,HEADS
folds=np.array([d["fold"] for d in D]); sp_idx=[i for i,d in enumerate(D) if not d["noise_only"]]
def build_systems(h):
S={}
for nm,(r,p) in {"Silero .5":("silero",(0.5,)),"head .5":("head",(0.5,)),"OR .5/.5":("or",(0.5,0.5)),"two-stage .5/.5/160/160/300":("two",(0.5,0.5,16,16,30))}.items():
S[nm]=lambda d,f,r=r,p=p:raw_mask(r,p,d,h)
for r,nm in (("silero","Silero tuned"),("head","head tuned"),("or","OR tuned"),("and","AND tuned"),("mean","mean tuned"),("two","two-stage tuned")):
_,ch=crossfit(static(r,h),"f1"); S[nm]=lambda d,f,r=r,ch=ch:raw_mask(r,ch[f],d,h)
models={}
for kind in("lr6","hgb6","lr6nz","hgb6nz"):
base=kind.replace("nz",""); _,ch=crossfit(lrget(h,kind),"f1")
for f in range(4):
tr=[i for i in range(len(D)) if folds[i]!=f and (kind.endswith("nz") or not D[i]["noise_only"])]
X=np.concatenate([feats(D[i],h,base)[::10] for i in tr]); y=np.concatenate([D[i]["ref_mask"][::10] for i in tr])
mdl=HistGradientBoostingClassifier(max_iter=120,learning_rate=0.1,random_state=0) if base=="hgb6" else LogisticRegression(C=1.0,max_iter=300)
models[(kind,f)]=mdl.fit(X,y)
S[f"{kind} tuned"]=lambda d,f,kind=kind,base=base,ch=ch:models[(kind,f)].predict_proba(feats(d,h,base))[:,1]>=ch[f][0]
return S
def metrics(fn,h):
st=[];en=[];miss=0;nref=0;nseg=0;cnt=np.zeros(4)
for i in sp_idx:
d=D[i]; m,regs=post(fn(d,d["fold"]),**POST_HEAD); a,b,ms=bnd(regs,d["ref"]); st+=a;en+=b;miss+=ms;nref+=len(d["ref"]);nseg+=len(regs); cnt+=counts(m,d["ref_mask"])
q=lambda x:(np.median(x),np.percentile(x,90))
return dict(start=q(st),end=q(en),miss=100*miss/nref,regions_per_ref=nseg/nref,PRF=[100*x for x in prf(cnt)])
out={}
for h in HEADS:
out[h]={}
for nm,fn in build_systems(h).items():
out[h][nm]=metrics(fn,h); o=out[h][nm]
print(f"{h:5s} {nm:30s} start {o['start'][0]:.0f}/{o['start'][1]:.0f} end {o['end'][0]:.0f}/{o['end'][1]:.0f} miss {o['miss']:.2f}% regions/ref {o['regions_per_ref']:.2f} PRF {o['PRF'][0]:.1f}/{o['PRF'][1]:.1f}/{o['PRF'][2]:.1f}",flush=True)
json.dump(out,open("detail.json","w"),default=float)
Loading
Loading