MiniMax H3
MiniMax H3 is a general-purpose omni-modal generation model that jointly understands multimodal contexts spanning text, images, video, and audio. It generates videos with native stereo sound at up to 2K resolution and 15 seconds in length, delivering content for advertising, branding, ecommerce, product design, UI/UX, gaming, and creative workflows. Users can combine reference types in one instruction, for example, transferring camera movement from a video, placing a character from an image into the scene, and matching vocals from an audio clip, while describing the relationships in natural language. H3 supports text-to-image, text-to-video with jointly generated audio, multi-shot modeling, text-to-audio, and generalized reference and editing across images, videos, and audio. Voice, sound effects, and music are modeled together. The model excels at instruction following, accurate text and brand presentation, and video-to-video motion transfer.
Learn more
FLUX 3
FLUX 3 is a multimodal foundation model that jointly learns from images, video, and audio within one unified architecture, building a representation of how objects hold together, how things move, and how events sound. Built on the Self-Flow approach, it aligns multimodal generation and understanding in the same backbone so each modality constrains the others, sound matches impact, motion follows physical properties, and future events follow from the past. FLUX 3 can mix modalities and jointly generate images, video, and native audio from text prompts or references such as images, video, and audio. Its video capabilities include text-to-video, image-to-video animation, video-to-video transformation, generative video-and-audio continuation, keyframe-controlled transitions, multilingual dialogue, animated typography, diverse styles and aspect ratios, and agentic chaining into longer multi-shot sequences.
Learn more
Grok Imagine Image 2.0
Grok Imagine Image 2.0 is an image generation and editing model from SpaceXAI built for precise creative work across photography, design, illustration, and multi-part visuals. The model is available as Quality Mode in Grok Imagine on grok.com, iOS, and Android. Grok Imagine Image 2.0 follows detailed instructions, preserves elements across generations and edits, and handles typography, layout, and sharp small text for practical creative assets. Its editing tools include magic wand region editing, segmentation, background removal, smart resize, and multi-reference editing with up to five input images. The platform also includes templates for photo editing, product shots, e-commerce photos, headshots, icons, game assets, emojis, merchandise, and more. Built for real creative workflows, Grok Imagine Image 2.0 helps users generate, edit, resize, and adapt images for professional and consumer use.
Learn more
Bonsai Image
Bonsai Image Ternary 4B MLX 2-bit is a ternary-weight text-to-image diffusion transformer deployment for Apple Silicon. It is built as a quality-oriented Bonsai Image variant, using ternary {−1, 0, +1} transformer weights with FP16 group-wise scaling in the matrix-heavy transformer layers, including Q/K/V projections, output projections, and MLP weights. The model reduces the FLUX.2 Klein 4B transformer from 7.75 GB FP16 to a 1.21 GB Bonsai Image transformer, a 6.4× smaller footprint, while keeping visual quality and prompt fidelity close to the original model. The Apple Silicon deployment payload is 3.88 GB, including the MLX 2-bit diffusion transformer, a 4-bit Qwen3-4B text encoder, and an FP16 Flux2 VAE. After prompt encoding, the text encoder is offloaded, so the denoising loop only keeps the compact transformer and VAE resident. The model uses a 4-step FlowMatchEuler sampler with guidance 1.0 and shift 3.0, with no CFG and no negative prompts required.
Learn more