A text-to-speech, speech-to-text and speech-to-speech library
Repo of Qwen2-Audio chat & pretrained large audio language model
Audio foundation model excelling in audio understanding
Taming Stable Diffusion for Lip Sync
The official Python library for the Fish Audio API
Automatically translates the text of a video based on a subtitle file
Oobabooga - The definitive Web UI for local AI, with powerful features
MiniMax H3 is a general-purpose, omni-modal generative system
Code for openai.fm, a demo for the OpenAI Speech API
Tokenizer-Free TTS for Multilingual Speech Generation
Generate audiobooks from EPUBs, PDFs and text with captions
super expressive prompting model based on ltx2.3
Speech-to-text, text-to-speech, and speaker recognition
A Family of Open Sourced Music Foundation Models
Miso TTS is an 8 billion, highly emotive text-to-speech model
A free, open source, and extensible speech-to-text application
SOTA Open Source TTS
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Generate blog articles from video or audio
Qwen3-omni is a natively end-to-end, omni-modal LLM
Open source text-to-speech tool, supports extra-long text
High-Quality Voice Cloning TTS for 600+ Languages
Automatic Speech Recognition with Word-level Timestamps
Comprehensive Gradio WebUI for audio processing
The open-source voice synthesis studio powered by Qwen3-TTS