Instant AI Face Swap
Recovering the Visual Space from Any Views
Contexts Optical Compression
Diffusion Transformer with Fine-Grained Chinese Understanding
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Sharp Monocular Metric Depth in Less Than a Second
Your clothes, extracted and organized with gpt-image
Multimodal model achieving SOTA performance
Official implementation of DreamCraft3D
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Large-language-model & vision-language-model based on Linear Attention
AI-powered tool to quickly remove watermarks from images flawlessly
AI Suite for upscaling, interpolating & restoring images/videos
Detect faces in an image
Software that can generate photos from paintings
Efficient 320B multimodal MoE model for coding and autonomous agents
Google’s flagship dense multimodal model for coding and reasoning
Small 3B-base multimodal model ideal for custom AI on edge hardware
Efficient multimodal MoE model for coding, tools, and reasoning
Omnimodal AI model for agents, coding, and long-context tasks
Compact 8B multimodal instruct model optimized for edge deployment
Efficient 14B multimodal instruct model with edge deployment and FP8