FLUX 3 ActionBlack Forest Labs
|
Gemini Robotics 2Google DeepMind
|
|||||
Related Products
|
||||||
About
FLUX 3 Action is an open-weight 7B world-action model designed for action prediction in robotics and other latency-sensitive visual environments. Derived from the multimodal FLUX 3 backbone, it was pretrained on large-scale image, video, and audio data with a strong emphasis on video, then adapted through joint video-action training and fine-tuning for specific embodiments and action spaces. Given an instruction, camera observations, and robot joint positions, the model predicts motor commands together with their expected visual outcomes, allowing a robot to execute actions, observe the environment again, and continuously replan. Unlike approaches that separate visual prediction from control, FLUX 3 Action jointly models future video and actions, transferring world understanding learned from broad video pretraining into robot control. Its single-step 7B checkpoint reaches a 38.3% success rate on RoboLab-120.
|
About
Gemini Robotics 2 is Google DeepMind’s intelligence layer for adaptable robots, bringing whole-body control, advanced dexterity, embodied reasoning, and multi-robot collaboration to physical AI. It includes three models. Gemini Robotics 2 is a vision-language-action model that converts visual and language input into motor control, enabling humanoids and bi-arm robots to act from feet to fingertips. It can coordinate walking, crouching, reaching, balancing, and object manipulation, while controlling five-fingered hands or standard grippers for delicate and precise tasks. Gemini Robotics ER 2 serves as the high-level brain, communicating with people, understanding its surroundings, planning multi-step tasks that last several minutes, coordinating actions with the VLA, tracking progress, self-correcting failures, and allowing different robots to work together.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Robotics researchers, AI developers, and embodied-agent teams seeking to predict and execute actions from multimodal observations with a fast, open-weight world-action model
|
Audience
Industrial robotics developers building adaptable humanoid and multi-robot systems for complex physical environments
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Supported
|
API
Offers API
Not Supported
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
No information available.
Free Version
Not Supported
Free Trial
Not Supported
|
Pricing
No information available.
Free Version
Not Supported
Free Trial
Not Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
|||||
Company InformationBlack Forest Labs
Founded: 2024
Germany
bfl.ai/models/flux-3-action
|
Company InformationGoogle DeepMind
Founded: 2010
United Kingdom
deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
Gemini
Not Supported
|
||||||
|
|
|