A text-to-speech, speech-to-text and speech-to-speech library
WhatsApp MCP server enabling AI access to chats and messaging
The Triton Inference Server provides an optimized cloud
A nearly-live implementation of OpenAI's Whisper
Mopidy is an extensible music server written in Python
Automated Music Discovery and Collection Manager
Music Assistant is a free, opensource Media library manager
Free, high-quality text-to-speech API endpoint to replace OpenAI
Instant voice cloning by MIT and MyShell. Audio foundation model
Swing Music is a beautiful, self-hosted music player
Ableton Live Model Context Protocol Integration
The music player of today
Googles NotebookLM but local
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Voice Recognition to Text Tool
Interface for OuteTTS models
Easy-to-use Speech Toolkit including Self-Supervised Learning model
Official MiniMax Model Context Protocol (MCP) server
Video editing with Python
A lightweight text-to-speech model with zero-shot voice cloning
Framework for building realtime multimodal voice AI agents apps
A simple native web interface that uses ChatTTS to synthesize text
tensorboard for pytorch (and chainer, mxnet, numpy, etc.)
Build cross-modal and multimodal applications on the cloud