Qwen3-ASR is an open-source series of ASR models
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Robust Speech Recognition Across Languages, Dialects
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
kaldi-asr/kaldi is the official location of the Kaldi project
Bailing is a voice dialogue robot similar to GPT-4o
StreamSpeech is a seamless model for offline speech recognition
Audio foundation model excelling in audio understanding
Video translation and dubbing tool powered by LLMs
Port of OpenAI's Whisper model in C/C++
Real-time voice interactive digital human
Repo of Qwen2-Audio chat & pretrained large audio language model
An Open Real-time Video-Language Interaction System
Easy-to-use Speech Toolkit including Self-Supervised Learning model
Scalable generative AI framework built for researchers and developers
Speech-AI-Forge is a project developed around TTS generation model
End-to-end speech processing toolkit
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
Conversational voice AI agents
Open-source framework for intelligent speech interaction
Open source AI VTuber platform with voice chat and Live2D avatars
Speech Note Linux app. Note taking, reading and translating
Fast and accurate automatic speech recognition (ASR) for edge devices
Framework for building AI-powered interactive digital humans and agent
Open-source industrial-grade ASR models