This repository contains an edge meeting-capture pipeline that uses camera-based speaker localization to steer a microphone-array beamformer before speech-to-text transcription. The system is designed for privacy-preserving meeting capture: video is used locally to infer who is speaking and where they are, while the live pipeline avoids persisting raw or beamformed WAV recordings.
- Detects people with YOLOv8 and tracks face landmarks with MediaPipe.
- Uses mouth openness (
jawOpen) to estimate the active talker. - Maps the active talker's image region to a microphone-array azimuth.
- Polls ReSpeaker USB firmware direction-of-arrival (DoA) as an audio fallback.
- Fuses vision and DoA when both are reliable.
- Applies delay-and-sum beamforming and light spectral enhancement.
- Streams enhanced beamformed audio to Whisper for chunked transcription.
- Saves transcripts, optional annotated video, and latency/resource metrics.
.
|-- vision_guided.py # Main fused vision + audio + Whisper pipeline
|-- vision.py # Camera, YOLO, FaceLandmarker, talker-region logic
|-- audio_dsp.py # DoA tracking, filtering, DAS beamforming, WAV writer
|-- transcribe.py # Whisper worker and fixed-duration audio chunker
|-- shared_state.py # Thread-safe shared state and angle helpers
|-- metrics.py # Runtime latency, CPU, and memory collection
|-- config.py # Hardware paths, camera/audio/DSP/model parameters
|-- record.py # Fixed-angle omni/DAS recorder for offline tests
|-- whisper_benchmarking.py # Whisper WER, latency, CPU, and RAM benchmark
|-- yolo_benchmarking.py # YOLO + FaceLandmarker talking/not-talking benchmark
|-- wer_e2e.py # End-to-end transcript WER comparison helper
|-- test_whisper_wer.py # Whisper WER test helper
|-- build_acm_report.py # Generates the ACM-style project report DOCX
|-- vision_guided_beamforming_acm_report.docx
|-- models/ # YOLO, NCNN, ONNX, and MediaPipe model assets
|-- benchmark_results/ # Saved component benchmark outputs
`-- End to End Results/ # Saved full-pipeline experiments
The default configuration targets a Linux edge device, such as a Raspberry Pi, with:
- USB camera
- ReSpeaker USB microphone array
- 16 kHz, 6-channel audio input
- Four raw microphone channels selected from the device stream
- Local CPU inference for YOLO, MediaPipe, and Whisper
The pipeline expects the ReSpeaker USB device with vendor/product ID 0x2886:0x0018.
Use Python 3.10+ if possible. The project imports the following main packages:
pip install numpy scipy opencv-python mediapipe ultralytics openai-whisper pyaudio pyusb psutil jiwer datasets tqdm python-docxDepending on the platform, pyaudio may require PortAudio system headers:
sudo apt-get update
sudo apt-get install portaudio19-dev python3-pyaudio ffmpegWhisper also requires ffmpeg to be available on the system path.
Model paths are configured in config.py. The current defaults use absolute Raspberry Pi paths:
YOLO_WEIGHTS = "/home/g20/Vision_Guided_Beamforming/models/yolov8n_ncnn_model"
FACE_MODEL = Path("/home/g20/Vision_Guided_Beamforming/models/face_landmarker.task")If you clone or move the repository, update these paths to match your local checkout. For example:
YOLO_WEIGHTS = str(PROJECT_ROOT / "models" / "yolov8n_ncnn_model")
FACE_MODEL = PROJECT_ROOT / "models" / "face_landmarker.task"The repository already includes several model artifacts under models/, including YOLOv8 NCNN/ONNX exports and MediaPipe landmarker task files.
Most tunable parameters live in config.py.
Important settings include:
DEVICE_INDEX: PyAudio input device index for the ReSpeaker.INPUT_CHANNELS: number of channels exposed by the audio device.RAW_MIC_CHANNELS: raw microphone channels used for beamforming.CAM_W,CAM_H: requested camera capture resolution.YOLO_EVERY_N: how often YOLO person detection runs.JAW_THRESHOLD: mouth-open threshold used to classify talking.REGION_AZIMUTH: mapping from camera region to microphone-array azimuth.WHISPER_MODEL: Whisper model used by the live pipeline.CHUNK_SECONDS: duration of audio chunks sent to Whisper.
Before running on new hardware, verify the camera index and PyAudio device index.
python3 vision_guided.pyCommon options:
python3 vision_guided.py --camera 0
python3 vision_guided.py --no-video
python3 vision_guided.py --no-yolo-filter
python3 vision_guided.py --video-output vision_debug.aviThe live pipeline:
- Opens the camera and probes for a valid frame.
- Loads YOLO and MediaPipe FaceLandmarker.
- Starts ReSpeaker DoA polling.
- Opens the configured PyAudio input stream.
- Processes audio in fixed blocks.
- Chooses either
beamorwidemode. - Sends beamformed audio chunks to Whisper without saving live WAV files.
- Prints a live status line showing direction, source, RMS, and buffering.
Stop with Ctrl+C. The program will flush remaining audio, wait for Whisper to drain, write metrics, and save final outputs.
By default, the full pipeline writes:
transcript.txt: final Whisper transcript with chunk tagsvision_fused_<timestamp>.avi: annotated video, unless--no-videois usedmetrics_<timestamp>.csv: latency, resource, and end-to-end timing summary
The live pipeline computes omni and beamformed audio internally, but audio recording is disabled in vision_guided.py; it does not save omni.wav or das_enhanced.wav during normal operation.
Transcript lines include tags such as:
[BEAM ang=90deg src=vision] ...
[WIDE src=none] ...
Use record.py when you intentionally want to capture an omni baseline and a fixed-angle DAS beamformed file for offline speech-to-text tests.
python3 record.py --angle 90 --output-prefix recordingThis writes:
recording_omni.wavrecording_das.wav
whisper_benchmarking.py benchmarks Whisper accuracy, latency, real-time factor, CPU, and RAM on samples from MLCommons People's Speech.
Example:
python3 whisper_benchmarking.py \
--model tiny \
--samples 50 \
--out benchmark_results/whisper_tiny_results.csvUseful options:
--model tiny|base|small|...--samples N--max_seconds 15--sample_cache_dir peoples_speech_cache--rebuild_cache
The script caches audio samples locally so repeated model runs can use the same data.
yolo_benchmarking.py evaluates talking/not-talking classification on a YOLO-format dataset with train, valid, and/or test splits.
Expected dataset layout:
dataset/
|-- train/
| |-- images/
| `-- labels/
|-- valid/
| |-- images/
| `-- labels/
`-- test/
|-- images/
`-- labels/
Run:
python3 yolo_benchmarking.py \
--dataset /path/to/dataset \
--split all \
--output-csv mouth_eval_summary_all.csvLabels are interpreted as:
- class
0: mouth closed / not talking - class
1: mouth open / talking
Use wer_e2e.py to compare a reference meeting script against a captured transcript and write WER details to wer_compare.json.
Edit REFERENCE_TEXT, HYPOTHESIS_TEXT, and OUTPUT_FILE in the script, then run:
python3 wer_e2e.pyThe repository includes saved benchmark and end-to-end experiment artifacts:
benchmark_results/Whisper - STT/benchmark_results/Whisper - STT - Clean Environment/benchmark_results/Whisper - STT - Background Noise/benchmark_results/YOLO_Face_LandMark/End to End Results/8n_ONNX_Tiny/End to End Results/8n_ONNX_Base/End to End Results/8n_NCNN_Tiny/End to End Results/8n_NCNN_Base/
These folders contain CSV, JSON, WAV, transcript, and video outputs from prior experiments.
If the camera cannot open:
- Check whether another process is using
/dev/video*. - Try a different
--cameraindex. - Put the camera and ReSpeaker on different USB controllers if bandwidth is an issue.
- Use
--no-videoif video encoding fails but vision inference works.
If the ReSpeaker is not found:
- Confirm the USB device appears with
lsusb. - Check permissions for USB access.
- Verify that the device ID matches
0x2886:0x0018.
If MediaPipe model loading fails:
- Confirm
FACE_MODELpoints to an existing.taskfile. - Make sure the file is not a partial download. The face landmarker model should be larger than 1 MB.
If Whisper is too slow:
- Use
WHISPER_MODEL = "tiny"or"base"inconfig.py. - Increase
CHUNK_SECONDSto reduce queue pressure. - Watch the printed real-time factor and queue wait metrics at shutdown.
- The
REGION_AZIMUTHmapping is hardware-specific. It reflects the camera and microphone mounting used for this project and should be recalibrated for a different physical setup. - The live system runs all perception and transcription locally.
- Annotated video is intended as a debug artifact; it can be disabled with
--no-video.