Unified web UI for training and running open models locally
Framework for building real-time voice and multimodal AI agents
SOTA discrete acoustic codec models with 40/75 tokens per second
AudioMuse-AI is an Open Source Dockerized environment
TTS model capable of streaming conversational audio in realtime
Unofficial Python API and agentic skill for Google NotebookLM
Googles NotebookLM but local
Pushing the Frontier of Long Audio-Visual Generation
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Open-source multi-speaker long-form text-to-speech model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Implementation of AudioLM audio generation model in Pytorch
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Qwen3-omni is a natively end-to-end, omni-modal LLM
Let Claude (or any LLM) actually watch a video
Download videos from almost any website
Video player for improving quality of hand-drawn images
Instant voice cloning by MIT and MyShell. Audio foundation model
Download videos from websites like YouTube and many others
Capable of understanding text, audio, vision, video
Sample code and notebooks for Generative AI on Google Cloud
Comprehensive Gradio WebUI for audio processing
Automatically translates the text of a video based on a subtitle file
Interface for OuteTTS models
Swing Music is a beautiful, self-hosted music player