Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Official Python inference and LoRA trainer package
Multimodal Diffusion with Representation Alignment
Official repository for LTX-Video
MOSS‑TTS Family open‑source speech and sound generation model
Python inference and LoRA trainer package for the LTX-2 audio–video
State-of-the-art TTS model under 25MB
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming