GPT-4o mini

GPT-4o mini

OpenAI
Mercury 2

Mercury 2

Inception
+
+

Related Products

  • Gemini Enterprise Agent Platform
    984 Ratings
    Visit Website
  • Google AI Studio
    30 Ratings
    Visit Website
  • LM-Kit.NET
    29 Ratings
    Visit Website
  • Google Cloud Speech-to-Text
    366 Ratings
    Visit Website
  • Qloo
    23 Ratings
    Visit Website
  • MEXC
    188,765 Ratings
    Visit Website
  • CallHub
    426 Ratings
    Visit Website
  • iPlum
    9,148 Ratings
    Visit Website
  • Caller ID Reputation
    42 Ratings
    Visit Website
  • Juspay
    17 Ratings
    Visit Website

About

A small model with superior textual intelligence and multimodal reasoning. GPT-4o mini enables a broad range of tasks with its low cost and latency, such as applications that chain or parallelize multiple model calls (e.g., calling multiple APIs), pass a large volume of context to the model (e.g., full code base or conversation history), or interact with customers through fast, real-time text responses (e.g., customer support chatbots). Today, GPT-4o mini supports text and vision in the API, with support for text, image, video and audio inputs and outputs coming in the future. The model has a context window of 128K tokens, supports up to 16K output tokens per request, and has knowledge up to October 2023. Thanks to the improved tokenizer shared with GPT-4o, handling non-English text is now even more cost effective.

About

Mercury 2 is the first reasoning model fast enough to pick up the phone, a reasoning diffusion language model built for real-time voice agents. Instead of making callers wait through seconds of dead air while an autoregressive model generates thinking tokens one by one, Mercury 2 uses a diffusion large language model architecture to generate tokens in parallel, decoding 1000+ tokens per second on standard NVIDIA GPUs. That speed is fast enough to run a full reasoning pass and start speaking within the latency budget of a natural conversation, reducing the cost of reasoning from seconds of silence to roughly 300 milliseconds. Mercury models work by corrupting clean text into noise, then training a standard Transformer to reverse the process and predict clean text across all positions simultaneously. Because each denoising pass touches many tokens, generation uses the GPU more efficiently than one-token-at-a-time decoding, making custom-silicon-like speed possible on NVIDIA H100s.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

Users interested in a powerful and low cost AI model

Audience

Voice AI infrastructure teams that need low-latency reasoning models for phone agents, tool-calling workflows, and natural customer conversations

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version
Free Trial

Pricing

No information available.
Free Version
Free Trial

Reviews/Ratings

Overall 5.0 / 5
ease 5.0 / 5
features 5.0 / 5
design 5.0 / 5
support 5.0 / 5

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

OpenAI
Founded: 2015
United States
openai.com

Company Information

Inception
United States
www.inceptionlabs.ai/blog/mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone

Alternatives

Alternatives

Mercury Coder

Mercury Coder

Inception Labs
Mercury Edit 2

Mercury Edit 2

Inception
Inkling

Inkling

Thinking Machines Lab
ByteDance Seed

ByteDance Seed

ByteDance
GPT-5 mini

GPT-5 mini

OpenAI
Uni-1

Uni-1

Luma AI

Categories

Categories

Integrations

OpenAI
Bind AI
Cerebras
ChatGPT Plus
ChatPerk
Duck.ai
Elixir
Fynix
GPT4Sales
Kotlin
MacWhisper
MindMac
Moemate
Progress Agentic RAG
Rust
Tessl
Tune AI
Tune Studio
TypeScript
Vapi AI

Integrations

OpenAI
Bind AI
Cerebras
ChatGPT Plus
ChatPerk
Duck.ai
Elixir
Fynix
GPT4Sales
Kotlin
MacWhisper
MindMac
Moemate
Progress Agentic RAG
Rust
Tessl
Tune AI
Tune Studio
TypeScript
Vapi AI
Claim GPT-4o mini and update features and information
Claim GPT-4o mini and update features and information
Claim Mercury 2 and update features and information
Claim Mercury 2 and update features and information