Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Robust Speech Recognition via Large-Scale Weak Supervision
End-to-end speech processing toolkit
Translate the video from one language to another and embed dubbing
Automatic Speech Recognition with Word-level Timestamps
Underthesea - Vietnamese NLP Toolkit
Faster Whisper transcription with CTranslate2
Toolkit for conversational AI
A modular voice assistant application for experimenting
Comprehensive Gradio WebUI for audio processing
Persian NLP Toolkit
Use Microsoft Edge's online text-to-speech service from Python
Open Source Speech Language Model
Audio foundation model excelling in audio understanding
Han Language Processing
Generate audiobooks from EPUBs, PDFs and text with captions
Open-source multi-speaker long-form text-to-speech model
AI-powered tool for generating, optimizing, and translating subtitles
Fast multimodal LLM for real-time voice interaction and AI apps
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Framework for building realtime multimodal voice AI agents apps
Training data (data labeling, annotation, workflow) for all data types
Towards Human-Sounding Speech
Voice Recognition to Text Tool
Self-host the powerful Chatterbox TTS model