▌
▛▘▌▌▛▌▛▌█▌▛▌
▄▌▙▌▙▌▙▌▙▖▌▌
▄▌
A GPU-accelerated batch subtitle generator for Linux, macOS, and WSL. It transcribes media folders to .srt using FFmpeg and whisper.cpp
subgen/
├── subgen.sh ← entry point
├── .gitmodules ← whisper.cpp declared as a submodule
└── whisper.cpp/ ← git submodule (ggml-org/whisper.cpp)
├── build/
│ └── bin/
│ └── whisper-cli ← compiled in the build step
└── models/
├── download-ggml-model.sh
├── download-vad-model.sh
├── ggml-large-v3.bin ← main transcription model
└── ggml-silero-v6.2.0.bin ← VAD model (optional but recommended)
- Project Structure
- Prerequisites
- Setup
- Hardware Acceleration (CUDA & Metal)
- Usage
- Add to PATH
- Configuration
- Model Reference
- Updating whisper.cpp
- Supported Video Formats
- Robustness Features
- Troubleshooting
- Future Work
- License
| Requirement | Notes |
|---|---|
| OS | Linux, macOS, or Windows Subsystem for Linux (WSL 2) |
| GPU (Optional) | NVIDIA GPU (requires CUDA) or Apple Silicon (uses Metal) |
| CUDA Toolkit | Required for NVIDIA GPUs (see Hardware Acceleration for Linux/WSL setup) |
cmake >= 3.14 |
sudo apt install cmake build-essential (Ubuntu/WSL) or brew install cmake (macOS) |
ffmpeg |
sudo apt install ffmpeg (Ubuntu/WSL) or brew install ffmpeg (macOS) |
git |
For cloning and submodule init |
fd-find (optional) |
sudo apt install fd-find or brew install fd (Massively speeds up file discovery on large drives) |
git clone --recurse-submodules https://github.com/abstraction/subgen.git
cd subgenIf you already cloned without --recurse-submodules:
git submodule update --init --recursiveFor NVIDIA GPUs (Linux/WSL):
cd whisper.cpp
cmake -B build -DGGML_CUDA=1
cmake --build build --config Release -j$(nproc)
cd ..For macOS (Apple Silicon), Metal is enabled by default:
cd whisper.cpp
cmake -B build
cmake --build build --config Release -j$(sysctl -n hw.ncpu)
cd ..Verify:
whisper.cpp/build/bin/whisper-clishould now exist.
If the build crashes on Linux/WSL: NVCC generates large memory structures for CUDA templates (
fattn,ggml-cuda), peaking at 3-4 GB per compiler thread. On RAM-limited machines this kills the build. Two options:Use a memory-aware job count instead of
-j$(nproc):cmake --build build -j $(free -g | awk '/^Mem:/{j=int($7/3); print j<1?1:j}') --config ReleaseOr add swap on the Windows host. Create/edit
%USERPROFILE%\.wslconfig:[wsl2] swap=8GBThen apply with
wsl --shutdownin PowerShell.
sudo apt install ffmpegcd whisper.cpp/models
bash download-ggml-model.sh large-v3
cd ../..Downloads ggml-large-v3.bin into whisper.cpp/models/.
Whisper large-v3 is OpenAI's most accurate speech recognition model. It requires a moderate amount of VRAM (usually around 4-5 GB) during inference, fitting comfortably on 6 GB cards like the RTX 3050. For smaller GPUs, you can use quantized variants like
q5_0.
Verify the model and GPU are working with the bundled JFK sample:
cd whisper.cpp
./build/bin/whisper-cli -m models/ggml-large-v3.bin -l en -f samples/jfk.wavYou should see a transcript. ggml_cuda_init in the output confirms the GPU was picked up.
Silero VAD is a small neural Voice Activity Detection model (~864 KB). With it enabled, whisper.cpp uses Silero VAD to detect speech regions before transcription, reducing unnecessary processing on long silent sections. This saves time on videos with long pauses, intros, or music.
The model is distributed in GGML format by ggml-org and comes with whisper.cpp's own download helper.
# From the project root:
bash whisper.cpp/models/download-vad-model.sh silero-v6.2.0Downloads ggml-silero-v6.2.0.bin (~864 KB) into whisper.cpp/models/.
Tip: VAD is on by default (
USE_VAD=true). To skip the download and run without it, setUSE_VAD=falseinsubgen.sh.
On Apple Silicon Macs, whisper.cpp automatically uses the Metal framework for GPU acceleration. No extra drivers or toolkits are needed.
On native Linux, install the NVIDIA driver and CUDA toolkit provided by your distribution or from NVIDIA's official repository.
NVIDIA GPU acceleration in Windows Subsystem for Linux works through GPU Paravirtualization (GPU-PV):
- Windows host: install the NVIDIA Graphics Driver from the official vendor portal. The Windows kernel-mode driver surfaces the GPU to WSL automatically.
- WSL2: don't install a Linux graphics driver (
.runor.deb) inside the WSL instance. It corrupts the paravirtualization layer and breaks the passthrough.
Verify the passthrough from inside WSL:
nvidia-smiExpected output: GPU name, driver version, and CUDA version pulled from the Windows host.
If CMake can't find the CUDA compiler (No CMAKE_CUDA_COMPILER could be found), install the CUDA toolkit from NVIDIA's official WSL repository:
wget https://developer.download.nvidia.com/compute/cuda/repos/wsl-ubuntu/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt -y install cuda-toolkitThen expose the compiler to your shell. Add these to ~/.zshrc or ~/.bashrc:
export PATH=/usr/local/cuda/bin:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATHReload and verify:
source ~/.zshrc # or ~/.bashrc
nvcc --version./subgen.sh [options] "/mnt/c/Users/YourName/Videos/ProjectFolder"-a, --auto: Enable multilingual auto-detection (sets language toauto)-l, --lang <lang>: Set specific language (e.g.,es,fr)-f, --force: Force regeneration even if an.srtfile already exists--no-postprocess: Disable the Python post-processor and retain raw Whisper output--cpl <number>: Override the maximum characters per line for subtitles (default: 40)
The script will:
- Discover all supported media files (video & audio) in the given directory (recursive).
- Skip any media file that already has a matching
.srtbeside it (resume-safe), unless--forceis passed. - Extract a mono 16 kHz WAV to
/tmp/via ffmpeg. - Transcribe with
whisper-clion your NVIDIA GPU. - Write the
.srtnext to the original file and clean up temp files. - Log any failures to
transcription_errors.loginside the input directory.
# Standard English run:
./subgen.sh "/mnt/c/Users/Alice/Videos/Lectures"
# output:
# /mnt/c/Users/Alice/Videos/Lectures/lecture_01.srt
# /mnt/c/Users/Alice/Videos/Lectures/lecture_02.srt
# ...
# Multilingual run (auto-detect language):
./subgen.sh --auto "/mnt/c/Users/Alice/Videos/Lectures"
# Specific language run (e.g., French):
./subgen.sh --lang fr "/mnt/c/Users/Alice/Videos/Lectures"To run subgen from anywhere without specifying the full path, symlink it into ~/.local/bin:
mkdir -p ~/.local/bin
ln -s /path/to/subgen/subgen.sh ~/.local/bin/subgen
chmod +x /path/to/subgen/subgen.shThen verify ~/.local/bin is on your PATH (it usually is on modern Linux distros):
echo $PATH | grep -o "$HOME/.local/bin"If it's missing, add this to your ~/.bashrc or ~/.zshrc:
export PATH="$HOME/.local/bin:$PATH"After that, you can run it from anywhere:
subgen "/mnt/c/Users/YourName/Videos/ProjectFolder"Note: The symlink approach means the script always reflects the latest version — no copying needed.
Most options live at the top of subgen.sh, though language can be overridden at runtime using flags.
| Variable | Default | Description |
|---|---|---|
WHISPER_MODEL_NAME |
large-v3 |
Whisper model variant to use |
LANGUAGE |
en |
Spoken language code |
TASK |
transcribe |
transcribe or translate (to English) |
USE_CUDA |
true |
Enable NVIDIA GPU acceleration |
NVIDIA_GPU_INDEX |
0 |
GPU index from nvidia-smi (set to 1 on dual-GPU systems) |
GPU_FEED_THREADS |
8 |
CPU threads used to feed the GPU |
USE_VAD |
true |
Enable Voice Activity Detection (skip silence) |
Edit WHISPER_MODEL_NAME in subgen.sh to any variant, then download it:
cd whisper.cpp/models && bash download-ggml-model.sh <model-name>See the Model Reference section below for a full breakdown.
.en models are English-only, which gives a marginal accuracy edge on clean English audio. The original Whisper release only provides English-specific .en checkpoints up to medium size. Large models are multilingual only:
.en model |
Multilingual equivalent |
|---|---|
tiny.en |
tiny |
base.en |
base |
small.en |
small |
medium.en |
medium |
large-v3 is multilingual only — there's no official large-v3.en. It handles English fine; pass -l en to lock the language and stop it from guessing wrong.
Quantization swaps 16-bit float weights for smaller integers, cutting file size and VRAM at a small accuracy cost.
Quality order: FP16 > Q8_0 > Q5_0 > Q4
| Quantization | Bits | Accuracy loss | Notes |
|---|---|---|---|
FP16 (none) |
16 | baseline | Full precision, highest VRAM |
Q8_0 |
8 | Negligible | Essentially identical to FP16 in practice |
Q5_0 |
5 | Very small | Best balance, recommended for most GPUs |
Q4 |
4 | Noticeable | Only worth it when VRAM is tight |
The real-world gap between Q8 and Q5 is small. Most files produce identical transcripts. You're more likely to notice a difference with heavy accents, noisy audio, overlapping speakers, or dense technical vocabulary.
VRAM usage varies significantly depending on context size and runtime settings. The below are approximations.
| Model | VRAM Need | Quality | Languages | Best for |
|---|---|---|---|---|
large-v3 (FP16) |
Moderate | Baseline | 99 | Default, best all-rounder |
large-v3-q8_0 |
Moderate | Near-lossless | 99 | Maximum quantized quality |
large-v3-q5_0 |
Low | Excellent | 99 | Good balance for smaller GPUs |
medium.en |
Moderate | Good | English | Unquantized medium baseline |
medium.en-q5_0 |
Low | Good | English | Speed priority or low-VRAM fallback |
large-v3(FP16) fits comfortably on most 6 GB cards like the RTX 3050. If you are tight on VRAM, you can use a quantized model likeq5_0.
Because large-v3 is multilingual, it auto-detects the language from audio. For English-only content, subgen.sh locks the language to English by default (LANGUAGE="en"). This prevents whisper-cli from hallucinating or mis-detecting languages on noisy English audio.
If you want multilingual output or auto-detection, use the runtime flags:
- Pass
--autoor-ato let Whisper guess the language automatically. - Pass
--lang <lang_code>or-l <lang_code>(e.g.,--lang fr) to lock it to a specific non-English language.
Smaller pre-quantized models are available via the download script and are the easiest option if you are low on VRAM:
bash ./models/download-ggml-model.sh large-v3-q5_0To quantize a full model yourself (e.g. to Q8_0):
# Download the full FP16 model first
bash ./models/download-ggml-model.sh large-v3
# Quantize it
./build/bin/whisper-quantize \
models/ggml-large-v3.bin \
models/ggml-large-v3-q8_0.bin \
q8_0Then set WHISPER_MODEL_NAME="large-v3-q8_0" in subgen.sh.
| Goal | Model to use |
|---|---|
| Best overall (default) | large-v3 + LANGUAGE=en |
| Maximum quantized quality | large-v3-q8_0 |
| Speed over accuracy | medium.en-q5_0 |
| Multilingual / language switching | large-v3 (run with --auto) |
The whisper.cpp directory is a Git submodule pinned to a specific upstream commit. To pull in upstream changes:
cd whisper.cpp
git fetch
git checkout master || git checkout main
git pull
cd ..
git add whisper.cpp
git commit -m "chore: bump whisper.cpp upstream commit reference"Rebuild after updating:
cd whisper.cpp && cmake -B build -DGGML_CUDA=1 && cmake --build build --config Release -j$(nproc)subgen includes a zero-dependency Python post-processor (json_to_srt.py) that formats raw Whisper transcriptions into broadcast-compliant SRT files. It automatically enforces rules from the Netflix Timed Text Style Guide and BBC Subtitle Guidelines:
- Characters Per Line (CPL) Limits: Automatically wraps text at a maximum of 40 characters (configurable via
--cpl <number>). - Reading Speed (CPS) Limits: Enforces a maximum reading speed of 20 characters per second (
MAX_CPS). Cues that are too fast are extended so viewers have time to read them. - Syntactic Line Balancing (Inverted Pyramid): Replaces Whisper's arbitrary character-count breaks with NLP parsing. It splits lines symmetrically on natural grammatical boundaries (punctuation, conjunctions, prepositions) while prioritizing the BBC "inverted pyramid" aesthetic (bottom-heavy lines).
- Abbreviation & SDH Bracket Awareness: Protects against mid-sentence splitting on abbreviations (
Dr.,U.S.) and prevents splitting inside Sound Effect / SDH brackets ([dog barking]). - Micro-Gap Snapping: Eliminates subtitle "flicker" by snapping micro-gaps (<120ms) between adjacent cues while preventing overlapping timestamps.
- CJK Support: Handles non-spaced languages (Chinese, Japanese, Korean) where Whisper token boundaries differ from Latin scripts.
This phase runs automatically as Phase 2.5. If you wish to disable it and retain raw Whisper output, pass the --no-postprocess flag.
Video: .mp4 · .mkv · .avi · .webm · .ts · .mov · .m4v · .wmv · .m2ts · .mts · .mxf · .vob · .flv
Audio: .mp3 · .m4a · .aac · .wav · .flac · .ogg · .opus · .wma
Detection is case-insensitive (.MP4, .mp3, etc. all work).
- Recursive batch discovery — uses
fd(if installed) to scan huge directories instantly, falling back tofindif missing. - Resume support — skips files that already have an SRT.
- VAD auto-retry — if whisper-cli crashes with VAD enabled (a known malloc issue), the script retries without VAD before giving up on that file.
- CUDA silent failure detection — inspects the whisper-cli log for
ggml_cuda_initto catch cases where the GPU was silently skipped. - CUDA version mismatch detection — detects
failed to initialize CUDAand exits with a clear message to update host NVIDIA drivers. - Graceful Ctrl+C — temp files are cleaned up even on interrupt.
- Error log — every failed file is recorded in
transcription_errors.logso you can handle failures selectively.
| Symptom | Fix |
|---|---|
whisper-cli: not found |
Run the cmake build step inside whisper.cpp/. |
| Model not found | Run bash whisper.cpp/models/download-ggml-model.sh large-v3. |
VAD model not found |
Run bash whisper.cpp/models/download-vad-model.sh silero-v6.2.0, or set USE_VAD=false in subgen.sh. |
| CUDA silent failure (fell back to CPU) | Not enough VRAM. Try medium.en-q5_0 (~1.0 GB VRAM). |
failed to initialize CUDA |
Update NVIDIA drivers on the Windows host, then reboot. |
No CMAKE_CUDA_COMPILER could be found |
Install the CUDA toolkit and add nvcc to PATH. See Hardware Acceleration. |
| Build crashes or system resets during compilation | NVCC OOM. Use the memory-aware build command or add swap. See Build step 2. |
| Wrong GPU used (iGPU instead of dGPU) | Set NVIDIA_GPU_INDEX=1 (or the correct index from nvidia-smi). |
| Zero-byte WAV file | The media file has no audio track; it is skipped automatically. |
| Scanning directory takes forever | Install fd-find (sudo apt install fd-find or brew install fd). Reading network/mounted drives is slow, and fd handles this much better than find. |
The current pipeline uses whisper.cpp with CUDA acceleration and the large-v3 model. It is already optimized for local GPU inference and includes VAD support, progress reporting, error handling, and automatic resume handling.
One area for improvement is the current process lifecycle. At the moment, each file starts a new whisper-cli process, which means the model has to be loaded into GPU memory again for every file:
File 1 → Start whisper-cli → Load model → Transcribe → Exit
File 2 → Start whisper-cli → Load model → Transcribe → Exit
File 3 → Start whisper-cli → Load model → Transcribe → Exit
For large batches containing many short videos, the repeated model initialization overhead can become significant.
A persistent transcription worker would keep the Whisper model loaded and process multiple files through the same inference session:
Video files → FFmpeg → Persistent Whisper worker → SRT output
|
└── Model loaded once and reused
Approaches to evaluate:
- Use the existing
whisper.cppserver/library API to keep a long-running inference process - Evaluate
faster-whisper(CTranslate2) as an alternative backend - Compare both approaches based on real workloads rather than synthetic benchmarks
The evaluation criteria would be:
- total batch processing time
- model initialization overhead
- GPU utilization
- VRAM usage
- transcription accuracy
- stability across long-running jobs
The current whisper.cpp CLI pipeline will remain the baseline until a measurable improvement is demonstrated.
whisper.cpp is a separate project with its own MIT License.