Fast and Universal 3D reconstruction model for versatile tasks
Unified Multimodal Understanding and Generation Models
State-of-the-art TTS model under 25MB
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
Sharp Monocular Metric Depth in Less Than a Second
A SOTA open-source image editing model
Qwen3-Coder is the code version of Qwen3
Accurate × Fast × Comprehensive
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
Qwen-Image is a powerful image generation foundation model
Controllable & emotion-expressive zero-shot TTS
Contexts Optical Compression
OCR expert VLM powered by Hunyuan's native multimodal architecture
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Designed for text embedding and ranking tasks
GPT4V-level open-source multi-modal model based on Llama3-8B
Global weather forecasting model using graph neural networks and JAX
Provides convenient access to the Anthropic REST API from any Python 3
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
General-purpose image editing model that delivers high-fidelity
GLM-4-Voice | End-to-End Chinese-English Conversational Model
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Renderer for the harmony response format to be used with gpt-oss
A series of math-specific large language models of our Qwen2 series