Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Official Python inference and LoRA trainer package
Multimodal Diffusion with Representation Alignment
Official repository for LTX-Video
MOSS‑TTS Family open‑source speech and sound generation model
State-of-the-art TTS model under 25MB
Python inference and LoRA trainer package for the LTX-2 audio–video
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Local-First AI for the Web