FLUX 3 Action

FLUX 3 Action

Black Forest Labs
GLM-OCR

GLM-OCR

Z.ai
+
+

Related Products

  • Adobe Firefly
    25,030 Ratings
    Visit Website
  • SAP S/4HANA Cloud Public Edition
    4,559 Ratings
    Visit Website
  • LTX
    182 Ratings
    Visit Website
  • Screencapt
    140 Ratings
    Visit Website
  • Grafana Cloud
    860 Ratings
    Visit Website
  • Shoplogix Smart Factory Platform
    19 Ratings
    Visit Website
  • Jama Connect
    387 Ratings
    Visit Website
  • Haast
    4 Ratings
    Visit Website
  • JetBrains Junie
    12 Ratings
    Visit Website
  • UptimeRobot
    852 Ratings
    Visit Website

About

FLUX 3 Action is an open-weight 7B world-action model designed for action prediction in robotics and other latency-sensitive visual environments. Derived from the multimodal FLUX 3 backbone, it was pretrained on large-scale image, video, and audio data with a strong emphasis on video, then adapted through joint video-action training and fine-tuning for specific embodiments and action spaces. Given an instruction, camera observations, and robot joint positions, the model predicts motor commands together with their expected visual outcomes, allowing a robot to execute actions, observe the environment again, and continuously replan. Unlike approaches that separate visual prediction from control, FLUX 3 Action jointly models future video and actions, transferring world understanding learned from broad video pretraining into robot control. Its single-step 7B checkpoint reaches a 38.3% success rate on RoboLab-120.

About

GLM-OCR is a multimodal optical character recognition model and open source repository that provides accurate, efficient, and comprehensive document understanding by combining text and visual modalities into a unified encoder–decoder architecture derived from the GLM-V family. Built with a visual encoder pre-trained on large-scale image–text data and a lightweight cross-modal connector feeding into a GLM-0.5B language decoder, the model supports layout detection, parallel region recognition, and structured output for text, tables, formulas, and complicated real-world document formats. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization, achieving state-of-the-art benchmarks on major document understanding tasks.

Platforms Supported

Windows Not Supported
Mac Not Supported
Linux Not Supported
Cloud Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Platforms Supported

Windows Not Supported
Mac Not Supported
Linux Not Supported
Cloud Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Audience

Robotics researchers, AI developers, and embodied-agent teams seeking to predict and execute actions from multimodal observations with a fast, open-weight world-action model

Audience

Developers, researchers, and engineers wanting a tool to accurately parse and understand complex documents, layouts, and visual-text content at scale

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Supported

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Supported

API

Offers API Supported

API

Offers API Supported

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version Not Supported
Free Trial Not Supported

Pricing

Free
Free Version Supported
Free Trial Not Supported

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Company Information

Black Forest Labs
Founded: 2024
Germany
bfl.ai/models/flux-3-action

Company Information

Z.ai
Founded: 2019
China
github.com/zai-org/GLM-OCR

Alternatives

Gemini Robotics-ER 1.6

Gemini Robotics-ER 1.6

Google DeepMind

Alternatives

CodeT5

CodeT5

Salesforce
Gemini Robotics 2

Gemini Robotics 2

Google DeepMind
HunyuanOCR

HunyuanOCR

Tencent
FLUX 3

FLUX 3

Black Forest Labs
DeepSeek-OCR

DeepSeek-OCR

DeepSeek
Gemini Robotics

Gemini Robotics

Google DeepMind
Starchild-1

Starchild-1

Odyssey

Categories

AI Models Supported

Categories

AI Models Supported
OCR Supported

Integrations

No info available.

Integrations

No info available.
Claim FLUX 3 Action and update features and information
Claim FLUX 3 Action and update features and information
Claim GLM-OCR and update features and information
Claim GLM-OCR and update features and information