SOTA on-device LLMs, small yet powerful
Convert Google Gemini web into OpenAI-compatible API
New set of lightweight state-of-the-art, open foundation models
Open Frontier Intelligence
Contexts Optical Compression
MiniMax M2.1, a SOTA model for real-world dev & agents.
Strong, Economical, and Efficient Mixture-of-Experts Language Model
MiMo-V2-Flash: Efficient Reasoning, Coding, and Agentic Foundation
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
New family of code large language models (LLMs)
MiniMax-M2, a model built for Max coding & agentic workflows
DeepSeek LLM: Let there be answers
Efficient 13B MoE language model with long context and reasoning modes
Pruned GLM-5.3 model for self-hosted cybersecurity AI and coding
Open agentic coding model optimized for local deployment
685B model with improved agents and consistency
Efficient 14B multimodal instruct model with edge deployment and FP8
Efficient MoE model for reasoning, coding, and AI agent workflows
Lightweight 24B agentic coding model with vision and long context
Compact agentic model for coding, tools, and productivity tasks
Compact 8B multimodal instruct model optimized for edge deployment
4-bit Command A+ model for enterprise agents and multilingual tasks
Hermes 4 FP8: hybrid reasoning Llama-3.1-405B model by Nous Research
Compact hybrid reasoning language model for intelligent responses
Efficient 250B MoE model for agents, coding, and long-context work