A text-to-speech, speech-to-text and speech-to-speech library
Open-source framework for intelligent speech interaction
Repo of Qwen2-Audio chat & pretrained large audio language model
Audio foundation model excelling in audio understanding
Audio Language Models are Few-Shot Learners
LLM-based Reinforcement Learning audio edit model
Taming Stable Diffusion for Lip Sync
The official Python library for the Fish Audio API
Multi-modal large language model designed for audio understanding
Automatically translates the text of a video based on a subtitle file
Oobabooga - The definitive Web UI for local AI, with powerful features
MiniMax H3 is a general-purpose, omni-modal generative system
Tokenizer-Free TTS for Multilingual Speech Generation
Generate audiobooks from EPUBs, PDFs and text with captions
A Family of Open Sourced Music Foundation Models
Miso TTS is an 8 billion, highly emotive text-to-speech model
SOTA Open Source TTS
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Generate blog articles from video or audio
Qwen3-omni is a natively end-to-end, omni-modal LLM
High-Quality Voice Cloning TTS for 600+ Languages
Comprehensive Gradio WebUI for audio processing
Automatic Speech Recognition with Word-level Timestamps
TTS model capable of streaming conversational audio in realtime
Capable of understanding text, audio, vision, video