This repository is a voice activity detection (VAD) toolkit that implements multiple models (DNN, bDNN, LSTM, ACAM) for detecting speech versus non-speech in audio. It also provides a recorded dataset in varied real-world settings (e.g. bus stop, construction site, park, room) with ground truth labeling. Acoustic feature extraction (multi-resolution cochleagram, MRCG). Post-processing modules (e.g. smoothing, thresholds). The toolkit supports both MATLAB and Python/TensorFlow components (for feature extraction, classification, postprocessing). Acoustic feature extraction (multi-resolution cochleagram, MRCG). Provided real-world dataset with manual annotations.
Features
- Multi-model VAD: DNN, boosted DNN (bDNN), LSTM, ACAM (adaptive context attention model)
- Acoustic feature extraction (multi-resolution cochleagram, MRCG)
- Training scripts (MATLAB / Python) and pretrained models
- Post-processing modules (e.g. smoothing, thresholds)
- Provided real-world dataset with manual annotations
- Combined MATLAB + TensorFlow interoperability
Categories
Sound/AudioFollow VAD
Other Useful Business Software
Save Up to 91% on Cloud Compute With Spot VMs
Run batch jobs at 60-91% off with Spot VMs. Long-running workloads get automatic discounts with sustained use.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of VAD!