GLM-4.5

GLM-4.5

Z.ai
Mercury 2

Mercury 2

Inception
+
+

Related Products

  • Gemini Enterprise Agent Platform
    984 Ratings
    Visit Website
  • Nexo
    18,395 Ratings
    Visit Website
  • Google Cloud Speech-to-Text
    366 Ratings
    Visit Website
  • GWI
    198 Ratings
    Visit Website
  • Google AI Studio
    30 Ratings
    Visit Website
  • Canditech
    110 Ratings
    Visit Website
  • Zendesk
    7,954 Ratings
    Visit Website
  • Fraud.net
    56 Ratings
    Visit Website
  • LM-Kit.NET
    29 Ratings
    Visit Website
  • Kevel
    96 Ratings
    Visit Website

About

GLM‑4.5 is Z.ai’s latest flagship model in the GLM family, engineered with 355 billion total parameters (32 billion active) and a companion GLM‑4.5‑Air variant (106 billion total, 12 billion active) to unify advanced reasoning, coding, and agentic capabilities in one architecture. It operates in a “thinking” mode for complex, multi‑step reasoning and tool use, and a “non‑thinking” mode for instant responses, supporting up to 128 K token context length and native function calling. Available via the Z.ai chat platform and API, with open weights on HuggingFace and ModelScope, GLM‑4.5 ingests diverse inputs to solve general problem‑solving, common‑sense reasoning, coding from scratch or within existing projects, and end‑to‑end agent workflows such as web browsing and slide generation. Built on a Mixture‑of‑Experts design with loss‑free balance routing, grouped‑query attention, and an MTP layer for speculative decoding, it delivers enterprise‑grade performance.

About

Mercury 2 is the first reasoning model fast enough to pick up the phone, a reasoning diffusion language model built for real-time voice agents. Instead of making callers wait through seconds of dead air while an autoregressive model generates thinking tokens one by one, Mercury 2 uses a diffusion large language model architecture to generate tokens in parallel, decoding 1000+ tokens per second on standard NVIDIA GPUs. That speed is fast enough to run a full reasoning pass and start speaking within the latency budget of a natural conversation, reducing the cost of reasoning from seconds of silence to roughly 300 milliseconds. Mercury models work by corrupting clean text into noise, then training a standard Transformer to reverse the process and predict clean text across all positions simultaneously. Because each denoising pass touches many tokens, generation uses the GPU more efficiently than one-token-at-a-time decoding, making custom-silicon-like speed possible on NVIDIA H100s.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

Developers and AI practitioners wanting a solution providing reasoning, coding and agentic functions for building sophisticated applications

Audience

Voice AI infrastructure teams that need low-latency reasoning models for phone agents, tool-calling workflows, and natural customer conversations

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version
Free Trial

Pricing

No information available.
Free Version
Free Trial

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

Z.ai
Founded: 2019
China
z.ai/blog/glm-4.5

Company Information

Inception
United States
www.inceptionlabs.ai/blog/mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone

Alternatives

Alternatives

Mercury Coder

Mercury Coder

Inception Labs
Mercury Edit 2

Mercury Edit 2

Inception
Inkling

Inkling

Thinking Machines Lab
ByteDance Seed

ByteDance Seed

ByteDance
Sarvam 105B

Sarvam 105B

Sarvam
Uni-1

Uni-1

Luma AI

Categories

Categories

Integrations

ArkClaw
Biela.dev
Cerebras
GPT-4.1
Groq
Hugging Face
Inception Labs
LiveKit
ModelScope
Nebius Token Factory
OpenAI
OpenClaw
Pipecat
Retell AI
SiliconFlow
Trancy
Vapi AI

Integrations

ArkClaw
Biela.dev
Cerebras
GPT-4.1
Groq
Hugging Face
Inception Labs
LiveKit
ModelScope
Nebius Token Factory
OpenAI
OpenClaw
Pipecat
Retell AI
SiliconFlow
Trancy
Vapi AI
Claim GLM-4.5 and update features and information
Claim GLM-4.5 and update features and information
Claim Mercury 2 and update features and information
Claim Mercury 2 and update features and information