ModelScope

ModelScope

Alibaba Cloud
+
+

Related Products

  • LTX
    182 Ratings
    Visit Website
  • LM-Kit.NET
    29 Ratings
    Visit Website
  • IONOS Cloud GPU Servers
    45,199 Ratings
    Visit Website
  • JetBrains Junie
    12 Ratings
    Visit Website
  • Google AI Studio
    41 Ratings
    Visit Website
  • Checksum.ai
    1 Rating
    Visit Website
  • Adobe Firefly
    25,030 Ratings
    Visit Website
  • Google Cloud Speech-to-Text
    366 Ratings
    Visit Website
  • FinOpsly
    3 Ratings
    Visit Website
  • Runpod
    230 Ratings
    Visit Website

About

DiffusionGemma is an experimental open model that explores text diffusion, an exceptionally fast approach to text generation. Released under an Apache 2.0 license, this 26B Mixture of Experts (MoE) model moves beyond the sequential token-by-token processing of typical autoregressive Large Language Models (LLMs). Instead, it generates entire blocks of text simultaneously, delivering up to 4x faster text generation on GPUs. Built on the intelligence-per-parameter of the Gemma 4 family and Gemini Diffusion research, DiffusionGemma integrates a novel diffusion head designed to maximize generation speed. It is designed for researchers and developers exploring speed-critical, interactive local workflows such as in-line editing, rapid iteration, and non-linear text structures. By shifting the decode bottleneck from memory bandwidth to compute, it can generate more than 1,000 tokens per second on a single NVIDIA H100 and more than 700 tokens per second on an NVIDIA GeForce RTX 5090.

About

This model is based on a multi-stage text-to-video generation diffusion model, which inputs a description text and returns a video that matches the text description. Only English input is supported. This model is based on a multi-stage text-to-video generation diffusion model, which inputs a description text and returns a video that matches the text description. Only English input is supported. The text-to-video generation diffusion model consists of three sub-networks: text feature extraction, text feature-to-video latent space diffusion model, and video latent space to video visual space. The overall model parameters are about 1.7 billion. Support English input. The diffusion model adopts the Unet3D structure, and realizes the function of video generation through the iterative denoising process from the pure Gaussian noise video.

Platforms Supported

Windows Supported
Mac Supported
Linux Supported
Cloud Not Supported
On-Premises Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Platforms Supported

Windows Not Supported
Mac Not Supported
Linux Not Supported
Cloud Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Audience

Researchers and developers wanting to explore fast local text generation for interactive AI applications, rapid iteration, editing, and other latency-sensitive workflows

Audience

Users interested in an open source text-to-video AI video generation model

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Supported

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Not Supported

API

Offers API Not Supported

API

Offers API Not Supported

Screenshots and Videos

Screenshots and Videos

Pricing

Free
Free Version Supported
Free Trial Not Supported

Pricing

Free
Free Version Supported
Free Trial Not Supported

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Company Information

Google
Founded: 1998
United States
blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/

Company Information

Alibaba Cloud
China
modelscope.cn/

Alternatives

Mercury 2

Mercury 2

Inception

Alternatives

Gemini Diffusion

Gemini Diffusion

Google DeepMind
Kaggle

Kaggle

Google
Mercury Coder

Mercury Coder

Inception Labs
ByteDance Seed

ByteDance Seed

ByteDance

Categories

AI Models Supported

Categories

AI Gateways Supported
AI Inference Supported
AI Tools Supported

Integrations

GLM-4.5 Not Supported
Gemma Supported
NVIDIA NIM Supported
Qwen 4 Not Supported
Qwen-Image Not Supported
Qwen2 Not Supported
Qwen2-VL Not Supported
Qwen2.5 Not Supported
Qwen2.5-1M Not Supported
Qwen2.5-Coder Not Supported
Qwen2.5-Max Not Supported
Qwen3.6-35B-A3B Not Supported
Qwen3.6-Max-Preview Not Supported
Qwen3.7-Max Not Supported
Qwen3.8-2.4T-A95B Not Supported
Qwen3.8-27B Not Supported
Qwen3.8-Max Not Supported
Step 3.5 Flash Not Supported
Step 5 Preview Not Supported
Yi-Large Not Supported

Integrations

GLM-4.5 Supported
Gemma Not Supported
NVIDIA NIM Not Supported
Qwen 4 Supported
Qwen-Image Supported
Qwen2 Supported
Qwen2-VL Supported
Qwen2.5 Supported
Qwen2.5-1M Supported
Qwen2.5-Coder Supported
Qwen2.5-Max Supported
Qwen3.6-35B-A3B Supported
Qwen3.6-Max-Preview Supported
Qwen3.7-Max Supported
Qwen3.8-2.4T-A95B Supported
Qwen3.8-27B Supported
Qwen3.8-Max Supported
Step 3.5 Flash Supported
Step 5 Preview Supported
Yi-Large Supported
Claim DiffusionGemma and update features and information
Claim DiffusionGemma and update features and information
Claim ModelScope and update features and information
Claim ModelScope and update features and information