Music player and music library manager for Linux, Windows, and macOS
Cross platform GUI tool for downloading videos from Bilibili sites
Multilingual speech recognition and audio understanding model
Dumb downloader that scrapes the web
A Python library for audio
Audio Normalization for Python/ffmpeg
Taming Stable Diffusion for Lip Sync
Automagically synchronize subtitles with video
Extract audio and video content and organize it into a Markdown note
Automatic Speech Recognition with Word-level Timestamps
Miso TTS is an 8 billion, highly emotive text-to-speech model
AI video generator optimized for low VRAM and older GPUs use
Tokenizer-Free TTS for Multilingual Speech Generation
VoiceStudio is the open-source, fully-local ElevenLabs alternative
Speakr is a personal, self-hosted web application
A Python library for audio data augmentation
Oobabooga - The definitive Web UI for local AI, with powerful features
The official Python SDK for the ElevenLabs API
Multimodal Diffusion with Representation Alignment
Open source AI model for generating full songs from lyrics prompts
Free, high-quality text-to-speech API endpoint to replace OpenAI
Generate audiobooks from EPUBs, PDFs and text with captions
Generate audiobooks from e-books, voice cloning & 1107+ languages
Automatic subtitle synchronization tool
Unified web UI for training and running open models locally