GLM-4.6V

GLM-4.6V

Z.ai
+
+

Related Products

  • LM-Kit.NET
    29 Ratings
    Visit Website
  • Gemini Enterprise Agent Platform
    1,161 Ratings
    Visit Website
  • LTX
    182 Ratings
    Visit Website
  • Google AI Studio
    41 Ratings
    Visit Website
  • RaimaDB
    12 Ratings
    Visit Website
  • Google Cloud Speech-to-Text
    443 Ratings
    Visit Website
  • Haast
    4 Ratings
    Visit Website
  • 4K Video Downloader
    13,127 Ratings
    Visit Website
  • LALAL.AI
    5,355 Ratings
    Visit Website
  • Planview AdaptiveWork
    718 Ratings
    Visit Website

About

EmbeddingGemma 2 is an open, lightweight multimodal embedding model designed to map text, code, images, video, and audio into a shared embedding space for search, retrieval, classification, routing, and RAG applications. Built on the Gemma 4 architecture and released under the Apache 2.0 license, it has 740 million parameters and is optimized for on-device inference. Its modular design can use as little as 270M parameters for text-only workloads, with optional vision and audio encoders for full multimodal support. Matryoshka Representation Learning lets developers reduce output vectors from 768 dimensions to 512, 256, or 128, lowering storage and memory requirements for local vector databases. The model supports an 8K-token context window and can process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations on local hardware.

About

GLM-4.6V is a state-of-the-art open source multimodal vision-language model from the Z.ai (GLM-V) family designed for reasoning, perception, and action. It ships in two variants: a full-scale version (106B parameters) for cloud or high-performance clusters, and a lightweight “Flash” variant (9B) optimized for local deployment or low-latency use. GLM-4.6V supports a native context window of up to 128K tokens during training, enabling it to process very long documents or multimodal inputs. Crucially, it integrates native Function Calling, meaning the model can take images, screenshots, documents, or other visual media as input directly (without manual text conversion), reason about them, and trigger tool calls, bridging “visual perception” with “executable action.” This enables a wide spectrum of capabilities; interleaved image-and-text content generation (for example, combining document understanding with text summarization or generation of image-annotated responses).

Platforms Supported

Windows Not Supported
Mac Not Supported
Linux Not Supported
Cloud Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Platforms Supported

Windows Supported
Mac Supported
Linux Supported
Cloud Supported
On-Premises Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Audience

Developers and AI teams wanting to build private, efficient, on-device multimodal search, retrieval, RAG, and semantic indexing systems

Audience

Developers, researchers, and AI engineers wanting a solution to build agents that understand images and text, manipulate documents or UIs, and generate complex image-text outputs

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Supported

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Supported

API

Offers API Supported

API

Offers API Supported

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version Not Supported
Free Trial Not Supported

Pricing

Free
Free Version Supported
Free Trial Not Supported

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Company Information

Google
Founded: 1998
United States
blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

Company Information

Z.ai
Founded: 2023
China
chat.z.ai/

Alternatives

Alternatives

GPT-5.2

GPT-5.2

OpenAI
GLM-4.1V

GLM-4.1V

Z.ai
txtai

txtai

NeuML
MiMo-V2.5

MiMo-V2.5

Xiaomi Technology

Categories

Embedding Models Supported

Categories

AI Coding Models Supported
AI Models Supported
Foundation Models Supported

Integrations

Claude Code Not Supported
Cline Not Supported
Kilo Code Not Supported
OpenRouter Not Supported
Roo Code Not Supported
Sup AI Not Supported

Integrations

Claude Code Supported
Cline Supported
Kilo Code Supported
OpenRouter Supported
Roo Code Supported
Sup AI Supported
Claim EmbeddingGemma 2 and update features and information
Claim EmbeddingGemma 2 and update features and information
Claim GLM-4.6V and update features and information
Claim GLM-4.6V and update features and information