Qwen-Image-2.1 is an open-source 7B-parameter image generation and editing model from the Qwen family. It unifies text-to-image creation and image editing within the same architecture. The model can natively generate transparent RGBA images, edit transparent layers, and extract subjects from photographs. Image editing supports up to 10 reference images for multi-subject composition and identity preservation. Local edits can be guided with circles, painted annotations, or separate masks. It supports native 2K output across multiple aspect ratios while improving typography, portrait lighting, textures, and fine details. Efficient inference features include mixed-granularity attention, prefix KV caching, model offloading, and integrations with Diffusers, ComfyUI, vLLM, SGLang, and LightX2V.
Features
- Unified text-to-image generation and image editing
- Native transparent RGBA image generation
- Up to 10 reference images for composition and editing
- Mask, circle, and painted-annotation guided local edits
- Native 2K output across multiple aspect ratios
- Diffusers, ComfyUI, vLLM, SGLang, and LightX2V support