The music player of today
Data manipulation and transformation for audio signal processing
A speech-text foundation model for real time dialogue
The official Python library for the OpenAI API
MOSS‑TTS Family open‑source speech and sound generation model
AI tool converting video/audio into structured documents instantly
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Mopidy is an extensible music server written in Python
Trying to be a robust, user-friendly and hackable music player
Industrial-level controllable zero-shot text-to-speech system
A TTS model capable of generating ultra-realistic dialogue
Translate the video from one language to another and embed dubbing
Open Source Speech Language Model
A general fine-tuning kit geared toward image/video/audio diffusion
A python tool that uses GPT-4, FFmpeg, and OpenCV
Generate audiobooks from e-books
Foundational video generation model with 13.6B parameters
Code and models for ICML 2024 paper, NExT-GPT
A high-quality rapid TTS voice cloning model
ImageBind One Embedding Space to Bind Them All
GenAI Processors is a lightweight Python library
Make any agent harness multimodal-native
Music Assistant is a free, opensource Media library manager
Qwen3-ASR is an open-source series of ASR models
Voice Recognition to Text Tool