Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Repo for SeedVR2 & SeedVR
Accurate × Fast × Comprehensive
State-of-the-art (SoTA) text-to-video pre-trained model
Repo of Qwen2-Audio chat & pretrained large audio language model
Qwen2.5-VL is the multimodal large language model series
The Clay Foundation Model - An open source AI model and interface
Advancing Open-source World Models
Global weather forecasting model using graph neural networks and JAX
code for Mesh R-CNN, ICCV 2019
Open image model at the forefront of design
26m function call model that runs on incredibly small devices
Qwen3-ASR is an open-source series of ASR models
A Pragmatic VLA Foundation Model
Collection of Gemma 3 variants that are trained for performance
Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
Diffusion Bee is the easiest way to run Stable Diffusion locally
Genome modeling and design across all domains of life
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Diffusion Transformer with Fine-Grained Chinese Understanding
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Large Multimodal Models for Video Understanding and Editing