Qwen-Image is a powerful image generation foundation model
Agentic, Reasoning, and Coding (ARC) foundation models
Text and image to video generation: CogVideoX and CogVideo
Qwen3 is the large language model series developed by Qwen team
A Family of Open Sourced Music Foundation Models
Official repository for LTX-Video
Advanced language and coding AI model
Models for object and human mesh reconstruction
High-Resolution Image Synthesis with Latent Diffusion Models
Convert Google Gemini web into OpenAI-compatible API
Qwen3-Coder is the code version of Qwen3
State-of-the-art TTS model under 25MB
Visual Causal Flow
Programmatic access to the AlphaGenome model
An experimental version of DeepSeek model
LTX-Video Support for ComfyUI
gpt-oss-120b and gpt-oss-20b are two open-weight language models
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Generating Immersive, Explorable, and Interactive 3D Worlds
Recovering the Visual Space from Any Views
Qwen2.5-VL is the multimodal large language model series
Code for running inference with the SAM 3D Body Model 3DB
Accurate × Fast × Comprehensive
Sharp Monocular Metric Depth in Less Than a Second
Renderer for the harmony response format to be used with gpt-oss