Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Vision-Guided Beam Steering for Privacy-Preserving Edge Meeting Capture

This repository contains an edge meeting-capture pipeline that uses camera-based speaker localization to steer a microphone-array beamformer before speech-to-text transcription. The system is designed for privacy-preserving meeting capture: video is used locally to infer who is speaking and where they are, while the live pipeline avoids persisting raw or beamformed WAV recordings.

What It Does

  • Detects people with YOLOv8 and tracks face landmarks with MediaPipe.
  • Uses mouth openness (jawOpen) to estimate the active talker.
  • Maps the active talker's image region to a microphone-array azimuth.
  • Polls ReSpeaker USB firmware direction-of-arrival (DoA) as an audio fallback.
  • Fuses vision and DoA when both are reliable.
  • Applies delay-and-sum beamforming and light spectral enhancement.
  • Streams enhanced beamformed audio to Whisper for chunked transcription.
  • Saves transcripts, optional annotated video, and latency/resource metrics.

Repository Layout

.
|-- vision_guided.py          # Main fused vision + audio + Whisper pipeline
|-- vision.py                 # Camera, YOLO, FaceLandmarker, talker-region logic
|-- audio_dsp.py              # DoA tracking, filtering, DAS beamforming, WAV writer
|-- transcribe.py             # Whisper worker and fixed-duration audio chunker
|-- shared_state.py           # Thread-safe shared state and angle helpers
|-- metrics.py                # Runtime latency, CPU, and memory collection
|-- config.py                 # Hardware paths, camera/audio/DSP/model parameters
|-- record.py                 # Fixed-angle omni/DAS recorder for offline tests
|-- whisper_benchmarking.py   # Whisper WER, latency, CPU, and RAM benchmark
|-- yolo_benchmarking.py      # YOLO + FaceLandmarker talking/not-talking benchmark
|-- wer_e2e.py                # End-to-end transcript WER comparison helper
|-- test_whisper_wer.py       # Whisper WER test helper
|-- build_acm_report.py       # Generates the ACM-style project report DOCX
|-- vision_guided_beamforming_acm_report.docx
|-- models/                   # YOLO, NCNN, ONNX, and MediaPipe model assets
|-- benchmark_results/        # Saved component benchmark outputs
`-- End to End Results/       # Saved full-pipeline experiments

Hardware Assumptions

The default configuration targets a Linux edge device, such as a Raspberry Pi, with:

  • USB camera
  • ReSpeaker USB microphone array
  • 16 kHz, 6-channel audio input
  • Four raw microphone channels selected from the device stream
  • Local CPU inference for YOLO, MediaPipe, and Whisper

The pipeline expects the ReSpeaker USB device with vendor/product ID 0x2886:0x0018.

Software Requirements

Use Python 3.10+ if possible. The project imports the following main packages:

pip install numpy scipy opencv-python mediapipe ultralytics openai-whisper pyaudio pyusb psutil jiwer datasets tqdm python-docx

Depending on the platform, pyaudio may require PortAudio system headers:

sudo apt-get update
sudo apt-get install portaudio19-dev python3-pyaudio ffmpeg

Whisper also requires ffmpeg to be available on the system path.

Model Files

Model paths are configured in config.py. The current defaults use absolute Raspberry Pi paths:

YOLO_WEIGHTS = "/home/g20/Vision_Guided_Beamforming/models/yolov8n_ncnn_model"
FACE_MODEL = Path("/home/g20/Vision_Guided_Beamforming/models/face_landmarker.task")

If you clone or move the repository, update these paths to match your local checkout. For example:

YOLO_WEIGHTS = str(PROJECT_ROOT / "models" / "yolov8n_ncnn_model")
FACE_MODEL = PROJECT_ROOT / "models" / "face_landmarker.task"

The repository already includes several model artifacts under models/, including YOLOv8 NCNN/ONNX exports and MediaPipe landmarker task files.

Configuration

Most tunable parameters live in config.py.

Important settings include:

  • DEVICE_INDEX: PyAudio input device index for the ReSpeaker.
  • INPUT_CHANNELS: number of channels exposed by the audio device.
  • RAW_MIC_CHANNELS: raw microphone channels used for beamforming.
  • CAM_W, CAM_H: requested camera capture resolution.
  • YOLO_EVERY_N: how often YOLO person detection runs.
  • JAW_THRESHOLD: mouth-open threshold used to classify talking.
  • REGION_AZIMUTH: mapping from camera region to microphone-array azimuth.
  • WHISPER_MODEL: Whisper model used by the live pipeline.
  • CHUNK_SECONDS: duration of audio chunks sent to Whisper.

Before running on new hardware, verify the camera index and PyAudio device index.

Running the Full Pipeline

python3 vision_guided.py

Common options:

python3 vision_guided.py --camera 0
python3 vision_guided.py --no-video
python3 vision_guided.py --no-yolo-filter
python3 vision_guided.py --video-output vision_debug.avi

The live pipeline:

  1. Opens the camera and probes for a valid frame.
  2. Loads YOLO and MediaPipe FaceLandmarker.
  3. Starts ReSpeaker DoA polling.
  4. Opens the configured PyAudio input stream.
  5. Processes audio in fixed blocks.
  6. Chooses either beam or wide mode.
  7. Sends beamformed audio chunks to Whisper without saving live WAV files.
  8. Prints a live status line showing direction, source, RMS, and buffering.

Stop with Ctrl+C. The program will flush remaining audio, wait for Whisper to drain, write metrics, and save final outputs.

Runtime Outputs

By default, the full pipeline writes:

  • transcript.txt: final Whisper transcript with chunk tags
  • vision_fused_<timestamp>.avi: annotated video, unless --no-video is used
  • metrics_<timestamp>.csv: latency, resource, and end-to-end timing summary

The live pipeline computes omni and beamformed audio internally, but audio recording is disabled in vision_guided.py; it does not save omni.wav or das_enhanced.wav during normal operation.

Transcript lines include tags such as:

[BEAM  ang=90deg  src=vision] ...
[WIDE  src=none] ...

Recording Fixed-Angle Audio

Use record.py when you intentionally want to capture an omni baseline and a fixed-angle DAS beamformed file for offline speech-to-text tests.

python3 record.py --angle 90 --output-prefix recording

This writes:

  • recording_omni.wav
  • recording_das.wav

Whisper Benchmarking

whisper_benchmarking.py benchmarks Whisper accuracy, latency, real-time factor, CPU, and RAM on samples from MLCommons People's Speech.

Example:

python3 whisper_benchmarking.py \
  --model tiny \
  --samples 50 \
  --out benchmark_results/whisper_tiny_results.csv

Useful options:

  • --model tiny|base|small|...
  • --samples N
  • --max_seconds 15
  • --sample_cache_dir peoples_speech_cache
  • --rebuild_cache

The script caches audio samples locally so repeated model runs can use the same data.

Vision Benchmarking

yolo_benchmarking.py evaluates talking/not-talking classification on a YOLO-format dataset with train, valid, and/or test splits.

Expected dataset layout:

dataset/
|-- train/
|   |-- images/
|   `-- labels/
|-- valid/
|   |-- images/
|   `-- labels/
`-- test/
    |-- images/
    `-- labels/

Run:

python3 yolo_benchmarking.py \
  --dataset /path/to/dataset \
  --split all \
  --output-csv mouth_eval_summary_all.csv

Labels are interpreted as:

  • class 0: mouth closed / not talking
  • class 1: mouth open / talking

End-to-End WER

Use wer_e2e.py to compare a reference meeting script against a captured transcript and write WER details to wer_compare.json.

Edit REFERENCE_TEXT, HYPOTHESIS_TEXT, and OUTPUT_FILE in the script, then run:

python3 wer_e2e.py

Existing Results

The repository includes saved benchmark and end-to-end experiment artifacts:

  • benchmark_results/Whisper - STT/
  • benchmark_results/Whisper - STT - Clean Environment/
  • benchmark_results/Whisper - STT - Background Noise/
  • benchmark_results/YOLO_Face_LandMark/
  • End to End Results/8n_ONNX_Tiny/
  • End to End Results/8n_ONNX_Base/
  • End to End Results/8n_NCNN_Tiny/
  • End to End Results/8n_NCNN_Base/

These folders contain CSV, JSON, WAV, transcript, and video outputs from prior experiments.

Troubleshooting

If the camera cannot open:

  • Check whether another process is using /dev/video*.
  • Try a different --camera index.
  • Put the camera and ReSpeaker on different USB controllers if bandwidth is an issue.
  • Use --no-video if video encoding fails but vision inference works.

If the ReSpeaker is not found:

  • Confirm the USB device appears with lsusb.
  • Check permissions for USB access.
  • Verify that the device ID matches 0x2886:0x0018.

If MediaPipe model loading fails:

  • Confirm FACE_MODEL points to an existing .task file.
  • Make sure the file is not a partial download. The face landmarker model should be larger than 1 MB.

If Whisper is too slow:

  • Use WHISPER_MODEL = "tiny" or "base" in config.py.
  • Increase CHUNK_SECONDS to reduce queue pressure.
  • Watch the printed real-time factor and queue wait metrics at shutdown.

Notes

  • The REGION_AZIMUTH mapping is hardware-specific. It reflects the camera and microphone mounting used for this project and should be recalibrated for a different physical setup.
  • The live system runs all perception and transcription locally.
  • Annotated video is intended as a debug artifact; it can be disabled with --no-video.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages