GLM-OCR

GLM-OCR

Z.ai
Mercury 2

Mercury 2

Inception
+
+

Related Products

  • PackageX OCR Scanning
    48 Ratings
    Visit Website
  • LogicalDOC
    148 Ratings
    Visit Website
  • Nutrient SDK
    111 Ratings
    Visit Website
  • MyQ
    197 Ratings
    Visit Website
  • Apryse PDF SDK
    157 Ratings
    Visit Website
  • Foxit Document Workflow APIs
    6 Ratings
    Visit Website
  • Google Cloud Speech-to-Text
    366 Ratings
    Visit Website
  • LTX
    182 Ratings
    Visit Website
  • LinkSquares
    724 Ratings
    Visit Website
  • LM-Kit.NET
    29 Ratings
    Visit Website

About

GLM-OCR is a multimodal optical character recognition model and open source repository that provides accurate, efficient, and comprehensive document understanding by combining text and visual modalities into a unified encoder–decoder architecture derived from the GLM-V family. Built with a visual encoder pre-trained on large-scale image–text data and a lightweight cross-modal connector feeding into a GLM-0.5B language decoder, the model supports layout detection, parallel region recognition, and structured output for text, tables, formulas, and complicated real-world document formats. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization, achieving state-of-the-art benchmarks on major document understanding tasks.

About

Mercury 2 is the first reasoning model fast enough to pick up the phone, a reasoning diffusion language model built for real-time voice agents. Instead of making callers wait through seconds of dead air while an autoregressive model generates thinking tokens one by one, Mercury 2 uses a diffusion large language model architecture to generate tokens in parallel, decoding 1000+ tokens per second on standard NVIDIA GPUs. That speed is fast enough to run a full reasoning pass and start speaking within the latency budget of a natural conversation, reducing the cost of reasoning from seconds of silence to roughly 300 milliseconds. Mercury models work by corrupting clean text into noise, then training a standard Transformer to reverse the process and predict clean text across all positions simultaneously. Because each denoising pass touches many tokens, generation uses the GPU more efficiently than one-token-at-a-time decoding, making custom-silicon-like speed possible on NVIDIA H100s.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

Developers, researchers, and engineers wanting a tool to accurately parse and understand complex documents, layouts, and visual-text content at scale

Audience

Voice AI infrastructure teams that need low-latency reasoning models for phone agents, tool-calling workflows, and natural customer conversations

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

Free
Free Version
Free Trial

Pricing

No information available.
Free Version
Free Trial

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

Z.ai
Founded: 2019
China
github.com/zai-org/GLM-OCR

Company Information

Inception
United States
www.inceptionlabs.ai/blog/mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone

Alternatives

CodeT5

CodeT5

Salesforce

Alternatives

Mercury Coder

Mercury Coder

Inception Labs
HunyuanOCR

HunyuanOCR

Tencent
Mercury Edit 2

Mercury Edit 2

Inception
DeepSeek-OCR

DeepSeek-OCR

DeepSeek
ByteDance Seed

ByteDance Seed

ByteDance
Uni-1

Uni-1

Luma AI

Categories

Categories

Integrations

Cerebras
GPT-4.1
Groq
Inception Labs
LiveKit
OpenAI
Pipecat
Retell AI
Vapi AI

Integrations

Cerebras
GPT-4.1
Groq
Inception Labs
LiveKit
OpenAI
Pipecat
Retell AI
Vapi AI
Claim GLM-OCR and update features and information
Claim GLM-OCR and update features and information
Claim Mercury 2 and update features and information
Claim Mercury 2 and update features and information