MiniMax H3
MiniMax H3 is a general-purpose omni-modal generation model that jointly understands multimodal contexts spanning text, images, video, and audio. It generates videos with native stereo sound at up to 2K resolution and 15 seconds in length, delivering content for advertising, branding, ecommerce, product design, UI/UX, gaming, and creative workflows. Users can combine reference types in one instruction, for example, transferring camera movement from a video, placing a character from an image into the scene, and matching vocals from an audio clip, while describing the relationships in natural language. H3 supports text-to-image, text-to-video with jointly generated audio, multi-shot modeling, text-to-audio, and generalized reference and editing across images, videos, and audio. Voice, sound effects, and music are modeled together. The model excels at instruction following, accurate text and brand presentation, and video-to-video motion transfer.
Learn more
Qwen-Image-3.0-Pro
Qwen-Image-3.0-Pro is an image generation model designed to turn text and image inputs into detailed, information-rich visuals that are useful beyond simple aesthetics. It supports prompts of up to 4.5K tokens and dense information layouts with images within images, allowing complex compositions such as newspapers, storyboards, menus, and exam papers to be generated in a single pass. The model emphasizes authentic detail, with precise rendering of text as small as 10 pixels and fine visual features including micro-expressions, skin pores, and individual strands of hair, approaching the quality of real photography. Qwen-Image-3.0-Pro also brings deeper knowledge into generation, supporting native text rendering in 12 languages and more than 20 fonts. It can realistically simulate mainstream digital interfaces, including web pages, games, and live-stream environments, while incorporating external knowledge into the resulting image.
Learn more
Qwen-Image-2.0
Qwen-Image 2.0 is the latest AI image generation and editing model in the Qwen family that combines both generation and editing in a single unified architecture, delivering high-quality visuals with professional-grade typography and layout capabilities directly from natural-language prompts. It supports text-to-image and image editing workflows with a lightweight 7 billion-parameter model that runs quickly while producing native 2048x2048 resolution outputs and handling long, detailed instructions up to about 1,000 tokens so creators can generate complex infographics, posters, slides, comics, and photorealistic scenes with accurate, well-rendered English and other language text embedded in the visuals. The unified model design means users donāt need separate tools for creating and modifying images, making it easier to iterate on ideas and refine compositions.
Learn more
Qwen-Image
Qwen-Image is a multimodal diffusion transformer (MMDiT) foundation model offering state-of-the-art image generation, text rendering, editing, and understanding. It excels at complex text integration, seamlessly embedding alphabetic and logographic scripts into visuals with typographic fidelity, and supports diverse artistic styles from photorealism to impressionism, anime, and minimalist design. Beyond creation, it enables advanced image editing operations such as style transfer, object insertion or removal, detail enhancement, in-image text editing, and human pose manipulation through intuitive prompts. Its built-in vision understanding tasks, including object detection, semantic segmentation, depth and edge estimation, novel view synthesis, and super-resolution, extend its capabilities into intelligent visual comprehension. Qwen-Image is accessible via popular libraries like Hugging Face Diffusers and integrates prompt-enhancement tools for multilingual support.
Learn more