Automated Music Discovery and Collection Manager
Transforming Multimodal Content into Captivating Multilingual Audio
A lightning fast audio upsampler
Python Audio Analysis Library: Feature Extraction, Classification
Cross platform GUI tool for downloading videos from Bilibili sites
Audio Normalization for Python/ffmpeg
Dumb downloader that scrapes the web
Tokenizer-Free TTS for Multilingual Speech Generation
Chinese Financial Trading Framework Based on Multi-Agent LLM
Taming Stable Diffusion for Lip Sync
Multilingual speech recognition and audio understanding model
Automatic Speech Recognition with Word-level Timestamps
Speakr is a personal, self-hosted web application
Automagically synchronize subtitles with video
VoiceStudio is the open-source, fully-local ElevenLabs alternative
The Chrome OS Virtual Machine Monitor
Music player and music library manager for Linux, Windows, and macOS
A Codex dating strategist who first captures emotions
Open source AI model for generating full songs from lyrics prompts
Miso TTS is an 8 billion, highly emotive text-to-speech model
Oobabooga - The definitive Web UI for local AI, with powerful features
A Python library for audio data augmentation
AI video generator optimized for low VRAM and older GPUs use
The official Python SDK for the ElevenLabs API
Framework for building real-time voice and multimodal AI agents