Learning agent trained in a diffusion world model
Fast-stable-diffusion + DreamBooth
dLLM: Simple Diffusion Language Modeling
High-Resolution Image Synthesis with Latent Diffusion Models
Open platform for sharing and discovering Stable Diffusion models
A general fine-tuning kit geared toward image/video/audio diffusion
100–200× Acceleration for Video Diffusion Models
Block Diffusion for Ultra-Fast Speculative Decoding
Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Temporal-Consistent Diffusion Model for Real-World Video
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image
Image generation model with single-stream diffusion transformer
PyTorch implementation of JiT
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Diffusion Transformer with Fine-Grained Chinese Understanding
My ComfyUI workflows collection
UniMate: One Unified Model to Animate Diverse Skeletons
Tokenizer-Free TTS for Multilingual Speech Generation
Multimodal Diffusion with Representation Alignment
Open-source multi-speaker long-form text-to-speech model
RGBD video generation model conditioned on camera input
Cosmos-RL is a flexible and scalable Reinforcement Learning framework
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Collection of CVPR 2026 Papers and Open Source Projects
Official inference repo for FLUX.1 models