A sound cloning tool with a web interface, using your voice
Edit videos with Claude Code
A Systematic Framework for Interactive World Modeling
Marrying Grounding DINO with Segment Anything & Stable Diffusion
State-of-the-art diffusion models for image and audio generation
Controllable and fast Text-to-Speech for over 7000 languages
Data Infrastructure providing an approach to multimodal AI workloads
Build multimodal language agents for fast prototype and production
A fast TTS architecture with conditional flow matching
Make any agent harness multimodal-native
Easy-to-use Speech Toolkit including Self-Supervised Learning model
The data structure for multimodal data
Build AI-powered semantic search applications
A python tool that uses GPT-4, FFmpeg, and OpenCV
HunyuanVideo: A Systematic Framework For Large Video Generation Model
tensorboard for pytorch (and chainer, mxnet, numpy, etc.)
Large Multimodal Models for Video Understanding and Editing
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Official PyTorch Implementation
Open-source Video Translation Skill
Build multimodal AI applications with cloud-native stack
TorchMultimodal is a PyTorch library
LLM Large Model of Selling Anchor
Context data platform for building observable, self-learning AI agents
MARS5 speech model (TTS) from CAMB.AI