"VideoRAG: Chat with Your Videos
Official Python inference and LoRA trainer package
AI-powered video clipping and highlight generation
Ableton Live Model Context Protocol Integration
MiniMax H3 is a general-purpose, omni-modal generative system
[CVPR 2025 Best Paper Award] VGGT
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Edit videos with Claude Code
Clone a voice in 5 seconds to generate arbitrary speech in real-time
A python tool that uses GPT-4, FFmpeg, and OpenCV
PyTorch code and models for VJEPA2 self-supervised learning from video
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
ImageBind One Embedding Space to Bind Them All
Large Multimodal Models for Video Understanding and Editing
Automatically translates the text of a video based on a subtitle file
Repo of Qwen2-Audio chat & pretrained large audio language model
Open-weight, large-scale hybrid-attention reasoning model
Download, transcribe and convert videos on your PC. No cloud.
Open-source AI video pipeline, fully automated with MCP
MARS5 speech model (TTS) from CAMB.AI
Audiocraft is a library for audio processing and generation
Real-time music generation using stable diffusion techniques AI
Twitch YouTube bot. Automatically make video compilations
A dataset of short, object-centric video clips