+
+

Related Products

  • LTX
    182 Ratings
    Visit Website
  • LM-Kit.NET
    29 Ratings
    Visit Website
  • Google AI Studio
    30 Ratings
    Visit Website
  • Checksum.ai
    1 Rating
    Visit Website
  • Adobe Firefly
    25,029 Ratings
    Visit Website
  • Google Cloud Speech-to-Text
    366 Ratings
    Visit Website
  • Runpod
    230 Ratings
    Visit Website
  • Imorgon
    5 Ratings
    Visit Website
  • MEXC
    188,765 Ratings
    Visit Website
  • EBizCharge
    207 Ratings
    Visit Website

About

DiffusionGemma is an experimental open model that explores text diffusion, an exceptionally fast approach to text generation. Released under an Apache 2.0 license, this 26B Mixture of Experts (MoE) model moves beyond the sequential token-by-token processing of typical autoregressive Large Language Models (LLMs). Instead, it generates entire blocks of text simultaneously, delivering up to 4x faster text generation on GPUs. Built on the intelligence-per-parameter of the Gemma 4 family and Gemini Diffusion research, DiffusionGemma integrates a novel diffusion head designed to maximize generation speed. It is designed for researchers and developers exploring speed-critical, interactive local workflows such as in-line editing, rapid iteration, and non-linear text structures. By shifting the decode bottleneck from memory bandwidth to compute, it can generate more than 1,000 tokens per second on a single NVIDIA H100 and more than 700 tokens per second on an NVIDIA GeForce RTX 5090.

About

NVIDIA Nemotron 3.5 Lightning is an open 30B-parameter mixture-of-experts model with 3B active parameters, designed for high-volume, low-latency execution in long-running and always-on AI agents. Built for the execution layer of agentic systems, it handles frequent tasks such as tool calls, output validation, routine commands, and subagent delegation while larger reasoning models focus on planning and orchestration. Its MoE architecture activates only a fraction of parameters for each token, combining the capacity of a larger model with lower compute requirements. The model is trained for popular agent harnesses and supports speculative decoding through multi-token prediction, DFlash, and DSpark to improve inference speed across different serving scenarios. It is available with BF16 and NVFP4 checkpoints and can run from local systems such as DGX Spark and GeForce RTX hardware to data center environments.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

Researchers and developers wanting to explore fast local text generation for interactive AI applications, rapid iteration, editing, and other latency-sensitive workflows

Audience

AI developers, engineering teams, and organizations building autonomous or long-running agents seeking to execute high-volume agentic tasks

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

Free
Free Version
Free Trial

Pricing

No information available.
Free Version
Free Trial

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

Google
Founded: 1998
United States
blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/

Company Information

NVIDIA
Founded: 1993
United States
nvidia.com

Alternatives

Mercury 2

Mercury 2

Inception

Alternatives

Gemini Diffusion

Gemini Diffusion

Google DeepMind
Mercury Coder

Mercury Coder

Inception Labs
Gemma 4

Gemma 4

Google
ByteDance Seed

ByteDance Seed

ByteDance
Nemotron 3

Nemotron 3

NVIDIA

Categories

Categories

Integrations

Gemini Enterprise Agent Platform
Gemma
Hermes Agent
NVIDIA NIM
NVIDIA NemoClaw
OpenClaw

Integrations

Gemini Enterprise Agent Platform
Gemma
Hermes Agent
NVIDIA NIM
NVIDIA NemoClaw
OpenClaw
Claim DiffusionGemma and update features and information
Claim DiffusionGemma and update features and information
Claim Nemotron 3.5 Lightning and update features and information
Claim Nemotron 3.5 Lightning and update features and information