Audio foundation model excelling in audio understanding
The official Python library for the Fish Audio API
MiniMax H3 is a general-purpose, omni-modal generative system
Code for openai.fm, a demo for the OpenAI Speech API
Tokenizer-Free TTS for Multilingual Speech Generation
super expressive prompting model based on ltx2.3
A Family of Open Sourced Music Foundation Models
Generate blog articles from video or audio
Open-source multi-speaker long-form text-to-speech model
Industrial-level controllable zero-shot text-to-speech system
A nearly-live implementation of OpenAI's Whisper
TTS model capable of streaming conversational audio in realtime
High-Quality Voice Cloning TTS for 600+ Languages
Open source text-to-speech tool, supports extra-long text
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Multimodal-Driven Architecture for Customized Video Generation
Official Python inference and LoRA trainer package
The python library for real-time communication
Controllable & emotion-expressive zero-shot TTS
Fast multimodal LLM for real-time voice interaction and AI apps
Interface for OuteTTS models
Framework for building real-time voice and multimodal AI agents
Use Microsoft Edge's online text-to-speech service from Python
Robust Speech Recognition via Large-Scale Weak Supervision
A TTS model capable of generating ultra-realistic dialogue