A Python library for audio
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Official Python inference and LoRA trainer package
Multimodal Diffusion with Representation Alignment
Official repository for LTX-Video
Workflow and speech recognition app
Speech-to-text, text-to-speech, and speaker recognition
Ableton Live Model Context Protocol Integration
Build Vision Agents quickly with any model or video provider
MOSS‑TTS Family open‑source speech and sound generation model
Python inference and LoRA trainer package for the LTX-2 audio–video
State-of-the-art TTS model under 25MB
A Telegram bot that integrates with OpenAI's official ChatGPT APIs
Open Source AI Dictation App
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Transform your voice in real-time voxal voice changer
Audiocraft is a library for audio processing and generation
Folo — AI-powered RSS reader for deep noise-free reading
Beauty can be applied to live broadcasts, short videos, and selfies
IPTV/NVR/CCTV/Video cloud https://fastocloud.com
Nash Operating System for Modern Ecommerce
Open source speech models for Julius in English and other languages.
Speech recognition application builder and library