GLM-5.3-FlashZ.ai
|
K2 HorizonInstitute of Foundation Models
|
|||||
Related Products
|
||||||
About
GLM-5.3-Flash is Z.ai’s natively multimodal model in the GLM-5 series (previously previewed as Ox Alpha), designed to deliver strong coding, agentic, visual, and knowledge-work performance at relatively low inference cost. It uses 320 billion total parameters with 18 billion active parameters, along with a hybrid architecture that combines sparse and linear attention to reduce the cost of long-context processing. The model supports context lengths of up to one million tokens and was trained on a 30-trillion-token multimodal corpus. GLM-5.3-Flash can reason across text, images, documents, interfaces, dashboards, and other visual information while using that feedback to refine its own outputs. Z.ai reports substantial gains over GLM-5.2 on coding and agentic benchmarks, including DeepSWE and AutomationBench, while approaching higher-cost frontier models on several evaluations.
|
About
K2 Horizon is a connected fleet of six open models spanning 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B, designed to deliver strong performance across reasoning, mathematics, coding, agentic tasks, and general capabilities. The models share core architecture, vocabulary, training methodology, interfaces, evaluation infrastructure, and deployment tooling, making it easier to move between sizes and route workloads dynamically. The 375B-A23B model is the fleet’s most capable option for complex reasoning, software engineering, research, and long-horizon agentic work, while the 32B and 36B-A4B models target powerful local deployment. The 36B-A4B model introduces Mixture-of-Value Attention, combining sparse attention with Mixture-of-Experts layers to activate about 4 billion parameters per token while approaching the performance of the dense 32B model.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
Developers, AI engineers, agent builders, researchers, and organizations that need cost-efficient multimodal reasoning, long-context processing, advanced coding, visual analysis, and autonomous workflow capabilities
|
Audience
AI researchers, developers, and organizations searching for open models for reasoning, coding, tool use, agentic workflows, and deployments
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
$0.15 per 1M tokens (input)
Input: $0.15 per 1M tokens
Output: $0.50 per 1M tokens Cached input: $0.03 per 1M tokens
Free Version
Free Trial
|
Pricing
No information available.
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Pros & Cons from Real UsersPros
Cons
|
||||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationZ.ai
Founded: 2019
China
z.ai
|
Company InformationInstitute of Foundation Models
Founded: 2025
United States
ifm.ai/blog/k2/
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
Cheaper Inference
Claude Code
DeepSeek Harness
GLM Coding Plan
Hermes Agent
OpenClaw
OpenCode Go
OpenCode Zen
OpenRouter
Pi Agent
|
Integrations
Cheaper Inference
Claude Code
DeepSeek Harness
GLM Coding Plan
Hermes Agent
OpenClaw
OpenCode Go
OpenCode Zen
OpenRouter
Pi Agent
|
|||||
|
|
|