Agora-1
Agora-1 is a multi-agent world model that enables multiple participants, human or AI, to share and interact within the same world simulation in real time. It is the first in a series of multi-agent world models exploring how world models can enable new shared experiences across gaming, robotics, defense, education, foundation models, and more. World models generate high-fidelity simulations of arbitrary environments, but until now, they have largely been limited to a single active participant inside those simulated worlds. Agora-1 introduces multi-agent world simulations by allowing up to four players to interact in the same generated world at once. Players are matched into a shared deathmatch simulation, where every participant interacts with the same world simultaneously while the model simulates player actions, maintains shared world state, and streams generated pixels to each player.
Learn more
MiniMax H3
MiniMax H3 is a general-purpose omni-modal generation model that jointly understands multimodal contexts spanning text, images, video, and audio. It generates videos with native stereo sound at up to 2K resolution and 15 seconds in length, delivering content for advertising, branding, ecommerce, product design, UI/UX, gaming, and creative workflows. Users can combine reference types in one instruction, for example, transferring camera movement from a video, placing a character from an image into the scene, and matching vocals from an audio clip, while describing the relationships in natural language. H3 supports text-to-image, text-to-video with jointly generated audio, multi-shot modeling, text-to-audio, and generalized reference and editing across images, videos, and audio. Voice, sound effects, and music are modeled together. The model excels at instruction following, accurate text and brand presentation, and video-to-video motion transfer.
Learn more
FLUX 3
FLUX 3 is a multimodal foundation model that jointly learns from images, video, and audio within one unified architecture, building a representation of how objects hold together, how things move, and how events sound. Built on the Self-Flow approach, it aligns multimodal generation and understanding in the same backbone so each modality constrains the others, sound matches impact, motion follows physical properties, and future events follow from the past. FLUX 3 can mix modalities and jointly generate images, video, and native audio from text prompts or references such as images, video, and audio. Its video capabilities include text-to-video, image-to-video animation, video-to-video transformation, generative video-and-audio continuation, keyframe-controlled transitions, multilingual dialogue, animated typography, diverse styles and aspect ratios, and agentic chaining into longer multi-shot sequences.
Learn more
Atlas by World Labs
Atlas is a next-generation omni world model for spatial intelligence that natively operates on text, images, video, and 3D. Built as a multimodal autoregressive diffusion transformer, it combines inputs into a shared spatial context and generates what comes next while staying consistent in 3D with everything it has seen and imagining what lies beyond it. Atlas supports world generation, reconstruction, and simulation across a broad range of tasks. It can generate images and videos from one or more reference images with pixel-perfect camera control, producing long, coherent videos with manually designed camera paths. For spatial reconstruction, Atlas can recreate real-world scenes from sparse input images, generate novel views, and produce explicit 3D outputs such as point clouds and 3D Gaussian splats. More input views provide additional context, allowing the model to reduce imagination and create increasingly faithful reconstructions. Atlas also models how worlds evolve over time.
Learn more