DiffusionGemmaGoogle
|
Mercury VoiceInception
|
|||||
Related Products
|
||||||
About
DiffusionGemma is an experimental open model that explores text diffusion, an exceptionally fast approach to text generation. Released under an Apache 2.0 license, this 26B Mixture of Experts (MoE) model moves beyond the sequential token-by-token processing of typical autoregressive Large Language Models (LLMs). Instead, it generates entire blocks of text simultaneously, delivering up to 4x faster text generation on GPUs. Built on the intelligence-per-parameter of the Gemma 4 family and Gemini Diffusion research, DiffusionGemma integrates a novel diffusion head designed to maximize generation speed. It is designed for researchers and developers exploring speed-critical, interactive local workflows such as in-line editing, rapid iteration, and non-linear text structures. By shifting the decode bottleneck from memory bandwidth to compute, it can generate more than 1,000 tokens per second on a single NVIDIA H100 and more than 700 tokens per second on an NVIDIA GeForce RTX 5090.
|
About
Mercury is a family of diffusion large language models built to deliver frontier LLM quality at significantly higher speeds, running at 1,000+ tokens per second on commercial NVIDIA GPUs for instant, in-the-flow AI applications. The models are OpenAI-compatible and designed as drop-in replacements for traditional LLMs, making them easier to integrate into existing AI stacks. Mercury 2.5 is the family’s most intelligent reasoning dLLM, built for complex applications where both quality and speed matter. It supports a 260K context window, reasoning, tool use, and structured output, with use cases including rapid coding iteration, agents and subagents, customer support, and enterprise search. Mercury Voice is optimized for voice agents and delivers time-to-first-token under 170 ms while supporting reasoning, tool use, structured output, and a 128K context window. It is suited to applications such as customer support, patient care, education, and gaming.
|
|||||
Platforms Supported
Windows
Supported
Mac
Supported
Linux
Supported
Cloud
Not Supported
On-Premises
Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Researchers and developers wanting to explore fast local text generation for interactive AI applications, rapid iteration, editing, and other latency-sensitive workflows
|
Audience
Developers and AI teams seeking to build fast, high-performance applications, agents, voice systems, and reasoning workflows with diffusion-based language models
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Not Supported
|
API
Offers API
Supported
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
Free
Free Version
Supported
Free Trial
Not Supported
|
Pricing
$0.04 per 1M tokens
Free Version
Supported
Free Trial
Not Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Supported
In Person
Not Supported
|
|||||
Company InformationGoogle
Founded: 1998
United States
blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/
|
Company InformationInception
United States
www.inceptionlabs.ai/models
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
Claude Haiku 4.5
Not Supported
GPT-5.6 Luna
Not Supported
Gemini 3.5 Flash-Lite
Not Supported
Gemini Enterprise Agent Platform
Supported
Gemma
Supported
NVIDIA NIM
Supported
OpenAI
Not Supported
|
Integrations
Claude Haiku 4.5
Supported
GPT-5.6 Luna
Supported
Gemini 3.5 Flash-Lite
Supported
Gemini Enterprise Agent Platform
Not Supported
Gemma
Not Supported
NVIDIA NIM
Not Supported
OpenAI
Supported
|
|||||
|
|
|