FLUX 3 ActionBlack Forest Labs
|
GLM-OCRZ.ai
|
|||||
Related Products
|
||||||
About
FLUX 3 Action is an open-weight 7B world-action model designed for action prediction in robotics and other latency-sensitive visual environments. Derived from the multimodal FLUX 3 backbone, it was pretrained on large-scale image, video, and audio data with a strong emphasis on video, then adapted through joint video-action training and fine-tuning for specific embodiments and action spaces. Given an instruction, camera observations, and robot joint positions, the model predicts motor commands together with their expected visual outcomes, allowing a robot to execute actions, observe the environment again, and continuously replan. Unlike approaches that separate visual prediction from control, FLUX 3 Action jointly models future video and actions, transferring world understanding learned from broad video pretraining into robot control. Its single-step 7B checkpoint reaches a 38.3% success rate on RoboLab-120.
|
About
GLM-OCR is a multimodal optical character recognition model and open source repository that provides accurate, efficient, and comprehensive document understanding by combining text and visual modalities into a unified encoder–decoder architecture derived from the GLM-V family. Built with a visual encoder pre-trained on large-scale image–text data and a lightweight cross-modal connector feeding into a GLM-0.5B language decoder, the model supports layout detection, parallel region recognition, and structured output for text, tables, formulas, and complicated real-world document formats. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization, achieving state-of-the-art benchmarks on major document understanding tasks.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Robotics researchers, AI developers, and embodied-agent teams seeking to predict and execute actions from multimodal observations with a fast, open-weight world-action model
|
Audience
Developers, researchers, and engineers wanting a tool to accurately parse and understand complex documents, layouts, and visual-text content at scale
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Supported
|
API
Offers API
Supported
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
No information available.
Free Version
Not Supported
Free Trial
Not Supported
|
Pricing
Free
Free Version
Supported
Free Trial
Not Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
|||||
Company InformationBlack Forest Labs
Founded: 2024
Germany
bfl.ai/models/flux-3-action
|
Company InformationZ.ai
Founded: 2019
China
github.com/zai-org/GLM-OCR
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
No info available.
|
Integrations
No info available.
|
|||||
|
|
|