MagicQuill
MagicQuill is an intelligent and interactive system that achieves precise image editing. As a highly practical application, image editing encounters a variety of user demands and thus prioritizes excellent ease of use. In this paper, we unveil MagicQuill, an integrated image editing system designed to help users swiftly actualize their creativity. Our system starts with a streamlined yet functionally robust interface, enabling users to articulate their ideas (e.g., inserting elements, erasing objects, altering color, etc.) with just a few strokes. These interactions are then monitored by a multimodal large language model (MLLM) to anticipate user intentions in real-time, bypassing the need for prompt entry. Finally, we apply the powerful diffusion prior, enhanced by a carefully learned two-branch plug-in module, to process the editing request with precise control. It facilitates accurate local edits, enhancing the overall editing experience.
Learn more
Reve 2.1
Reve 2.1 is a new foundation image model that makes a rapid leap in visual intelligence and world knowledge, just one month after Reve 2.0. It extends the same foundation of controllability, but sharpens it at every stage with intuitive prompt understanding, stronger foreign-text rendering, and more precise native 4K output. Reve 2.1 plans in finer detail, reasons more accurately about how elements relate, and renders results with greater precision at full 16-megapixel resolution. Built around the belief that images should be structured like code, with hierarchical layouts and controllable regions, the model brings layout planning directly into visual intelligence. It reasons about structure, hierarchy, and spatial relationships before rendering, making it stronger for dense scenes, intricate compositions, complicated visual instructions, and fine text. Reve 2.1 also supports precision editing, where every element is addressable and editable.
Learn more
Gemini 2.5 Flash Image
Gemini 2.5 Flash Image is Google’s latest state-of-the-art image generation and editing model, now accessible via the Gemini API, Google AI Studio’s build mode, and Gemini Enterprise Agent Platform. It enables powerful creative control by allowing users to blend multiple input images into a single visual, maintain consistent characters or products across edits for rich storytelling, and apply precise, natural-language-based–based transformations, such as removing objects, changing poses, adjusting colors, or altering backgrounds. The model is backed by Gemini’s deep world knowledge, enabling it to understand and reinterpret scenes or diagrams in context, which unlocks dynamic use cases like educational tutors or scene-aware editing assistants. Demonstrated through customizable template apps in AI Studio (including photo editors, multi-image fusers, and interactive tools), the model supports rapid prototyping and remixing via prompts or UI.
Learn more
Editpal
Editpal is an AI-powered image editor that lets you transform pictures simply by typing what you want. Whether you’re replacing backgrounds, changing colors, adjusting poses, or combining multiple images into one, Editpal makes it easy — no professional editing skills needed. It keeps characters and objects consistent across different edits, so your subject always looks the same in every scene. Perfect for creating marketing images, fine-tuning photos, designing educational visuals, or merging several pictures into a seamless composition.
What You Can Do With Editpal:
Quickly generate multiple versions of a product in different settings for e-commerce or ads.
Create realistic group photos or fix portraits with precision, all through text commands.
Produce educational or conceptual visuals from rough drawings or ideas.
Learn more