A text-to-speech, speech-to-text and speech-to-speech library
Repo of Qwen2-Audio chat & pretrained large audio language model
Audio Language Models are Few-Shot Learners
Audio foundation model excelling in audio understanding
Open-source framework for intelligent speech interaction
LLM-based Reinforcement Learning audio edit model
Large Audio Language Model built for natural interactions
Multi-modal large language model designed for audio understanding
The official Python library for the Fish Audio API
GUI for a Vocal Remover that uses Deep Neural Networks
Download your Spotify playlists and songs along with album art
MiniMax H3 is a general-purpose, omni-modal generative system
Official Python inference and LoRA trainer package
A cross-platform GUI wrapper for yt-dlp written in PySide6
Python library for audio and music analysis
Dumb downloader that scrapes the web
Tokenizer-Free TTS for Multilingual Speech Generation
A lightning fast audio upsampler
AudioMuse-AI is an Open Source Dockerized environment
Audio Normalization for Python/ffmpeg
A Python library for audio
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Taming Stable Diffusion for Lip Sync
Award-Winning Open Source Video Editing Software
Extract audio and video content and organize it into a Markdown note