This repository is a voice activity detection (VAD) toolkit that implements multiple models (DNN, bDNN, LSTM, ACAM) for detecting speech versus non-speech in audio. It also provides a recorded dataset in varied real-world settings (e.g. bus stop, construction site, park, room) with ground truth labeling. Acoustic feature extraction (multi-resolution cochleagram, MRCG). Post-processing modules (e.g. smoothing, thresholds). The toolkit supports both MATLAB and Python/TensorFlow components (for feature extraction, classification, postprocessing). Acoustic feature extraction (multi-resolution cochleagram, MRCG). Provided real-world dataset with manual annotations.
Features
- Multi-model VAD: DNN, boosted DNN (bDNN), LSTM, ACAM (adaptive context attention model)
- Acoustic feature extraction (multi-resolution cochleagram, MRCG)
- Training scripts (MATLAB / Python) and pretrained models
- Post-processing modules (e.g. smoothing, thresholds)
- Provided real-world dataset with manual annotations
- Combined MATLAB + TensorFlow interoperability
Categories
Sound/AudioFollow VAD
Other Useful Business Software
Build Data Resilience - Take the Assessment Today
Is your recovery strategy as strong as you think? Take this quick self-assessment to check your recovery readiness and gain tailored insights. In only 2 minutes, you'll learn where you fall on the recovery readiness scale.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of VAD!