Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Fast-stable-diffusion + DreamBooth
State-of-the-art TTS model under 25MB
Python inference and LoRA trainer package for the LTX-2 audio–video
AlphaFold 3 inference pipeline
Sharp Monocular Metric Depth in Less Than a Second
A Customizable Image-to-Video Model based on HunyuanVideo
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Reference PyTorch implementation and models for DINOv3
Text and image to video generation: CogVideoX and CogVideo
Official inference repo for FLUX.2 models
The official repo of Qwen chat & pretrained large language model
Qwen-Image is a powerful image generation foundation model
High-Resolution Image Synthesis with Latent Diffusion Models
Hackable and optimized Transformers building blocks
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
gpt-oss-120b and gpt-oss-20b are two open-weight language models
OpenTinker is an RL-as-a-Service infrastructure for foundation models
RGBD video generation model conditioned on camera input
Miso TTS is an 8 billion, highly emotive text-to-speech model
Diffusion Transformer with Fine-Grained Chinese Understanding
Advancing Open-source World Models
DeepMind model for tracking arbitrary points across videos & robotics
FAIR Sequence Modeling Toolkit 2