Agentic, Reasoning, and Coding (ARC) foundation models
From Images to High-Fidelity 3D Assets
Qwen3-Coder is the code version of Qwen3
Accurate × Fast × Comprehensive
Native and Compact Structured Latents for 3D Generation
Text and image to video generation: CogVideoX and CogVideo
Controllable & emotion-expressive zero-shot TTS
Official inference repo for FLUX.2 models
26m function call model that runs on incredibly small devices
Reference PyTorch implementation and models for DINOv3
Contexts Optical Compression
Open-source multi-speaker long-form text-to-speech model
Code for running inference and finetuning with SAM 3 model
Visual Causal Flow
Multimodal-Driven Architecture for Customized Video Generation
A theoretical reconstruction of the Claude Mythos architecture
AlphaFold 3 inference pipeline
GLM-4 series: Open Multilingual Multimodal Chat LMs
A Family of Open Sourced Music Foundation Models
Qwen-Image is a powerful image generation foundation model
Phi-3.5 for Mac: Locally-run Vision and Language Models
Open-source large language model family from Tencent Hunyuan
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Convert Google Gemini web into OpenAI-compatible API
A Powerful Native Multimodal Model for Image Generation