Qwen-Image is a powerful image generation foundation model
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Fast-stable-diffusion + DreamBooth
A Customizable Image-to-Video Model based on HunyuanVideo
Official inference repo for FLUX.2 models
Sharp Monocular Metric Depth in Less Than a Second
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Wan2.2: Open and Advanced Large-Scale Video Generative Model
High-Resolution Image Synthesis with Latent Diffusion Models
Text and image to video generation: CogVideoX and CogVideo
Reference PyTorch implementation and models for DINOv3
Diffusion Transformer with Fine-Grained Chinese Understanding
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Capable of understanding text, audio, vision, video
RGBD video generation model conditioned on camera input
Chinese and English multimodal conversational language model
Advancing Open-source World Models
AI Suite for upscaling, interpolating & restoring images/videos
Easy Docker setup for Stable Diffusion with user-friendly UI
A state-of-the-art open visual language model
A latent text-to-image diffusion model