GLM-5.3-FlashZ.ai
|
Step 5 PreviewStepFun
|
|||||
Related Products
|
||||||
About
GLM-5.3-Flash is Z.ai’s natively multimodal model in the GLM-5 series (previously previewed as Ox Alpha), designed to deliver strong coding, agentic, visual, and knowledge-work performance at relatively low inference cost. It uses 320 billion total parameters with 18 billion active parameters, along with a hybrid architecture that combines sparse and linear attention to reduce the cost of long-context processing. The model supports context lengths of up to one million tokens and was trained on a 30-trillion-token multimodal corpus. GLM-5.3-Flash can reason across text, images, documents, interfaces, dashboards, and other visual information while using that feedback to refine its own outputs. Z.ai reports substantial gains over GLM-5.2 on coding and agentic benchmarks, including DeepSWE and AutomationBench, while approaching higher-cost frontier models on several evaluations.
|
About
Step 5 Preview is StepFun’s flagship model for agentic work, designed for real-world tasks across software engineering and professional knowledge work, with particular strength in finance. It natively supports text, image, and video input and provides a 1M-token context window, enabling tasks that require large amounts of information, tool calls, and continuous progress toward a deliverable. The model can analyze long documents, multiple source materials, and conversation history for cross-document question answering and research organization. For programming and software engineering, it works across multiple languages and can support troubleshooting, code changes, verification, and test creation. Its multi-step agent capabilities let applications provide tools for retrieving information, processing documents, conducting deep research, and producing analytical reports. Multimodal understanding combines images, video, and text for chart analysis, screenshot question answering, etc.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
Developers, AI engineers, agent builders, researchers, and organizations that need cost-efficient multimodal reasoning, long-context processing, advanced coding, visual analysis, and autonomous workflow capabilities
|
Audience
Developers, engineering teams, and knowledge-work organizations wanting to build multimodal AI agents for coding, research, analysis, and other complex multi-step tasks
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
$0.15 per 1M tokens (input)
Input: $0.15 per 1M tokens
Output: $0.50 per 1M tokens Cached input: $0.03 per 1M tokens
Free Version
Free Trial
|
Pricing
$0.04 per input
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Pros & Cons from Real UsersPros
Cons
|
||||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationZ.ai
Founded: 2019
China
z.ai
|
Company InformationStepFun
Founded: 2023
United States
platform.stepfun.ai/docs/en/guides/models/step-5-preview
|
|||||
Alternatives |
AlternativesNo Alternatives
|
|||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
Hermes Agent
OpenClaw
Alibaba Cloud
Alibaba Cloud Model Studio
Cheaper Inference
Cline
DeepSeek Harness
Happy Shrimp 1.0
Hugging Face
ModelScope
|
Integrations
Hermes Agent
OpenClaw
Alibaba Cloud
Alibaba Cloud Model Studio
Cheaper Inference
Cline
DeepSeek Harness
Happy Shrimp 1.0
Hugging Face
ModelScope
|
|||||
|
|
|