A text-to-speech, speech-to-text and speech-to-speech library
Audio foundation model excelling in audio understanding
Audio Language Models are Few-Shot Learners
Open-source framework for intelligent speech interaction
Repo of Qwen2-Audio chat & pretrained large audio language model
Large Audio Language Model built for natural interactions
LLM-based Reinforcement Learning audio edit model
Multi-modal large language model designed for audio understanding
The official Python library for the Fish Audio API
GUI for a Vocal Remover that uses Deep Neural Networks
MiniMax H3 is a general-purpose, omni-modal generative system
Buzz transcribes and translates audio offline
Download your Spotify playlists and songs along with album art
A cross-platform GUI wrapper for yt-dlp written in PySide6
Official Python inference and LoRA trainer package
Award-Winning Open Source Video Editing Software
Python library for audio and music analysis
A Family of Open Sourced Music Foundation Models
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Fast multimodal LLM for real-time voice interaction and AI apps
A lightning fast audio upsampler
Python Audio Analysis Library: Feature Extraction, Classification
Transforming Multimodal Content into Captivating Multilingual Audio
Automated Music Discovery and Collection Manager
Speech recognition module for Python