The Qwen-Image-2.1 model integrates image creation and modification into a single checkpoint, facilitating seamless compositing via transparent image output. Released on Sept. 20, the 7B architecture supports up to 10 reference images for portrait and product fidelity and outperforms most closed-source models in a lightweight package. Integration is active via Day-0 support in ComfyUI, Diffusers, vLLM-Omni, and SGLang-Diffusion.

Hardware benchmarks show a single RTX 4090 24GB GPU generates 1024x1024 images in 18.7 seconds with 22.7 GiB peak memory. The tool is distributed under a strict non-commercial license with no revenue cap, prohibiting its use in monetized tutorials or sponsored social media posts.

Sign in to suggest edits

Key sources

  1. SOURCE@alibaba_qwen“lightweight 7B architecture that outperforms most closed-source models”x.com
  2. SUPPORT@alibaba_qwen“supports transparent image generation natively”x.com
  3. SUPPORT@alibaba_qwen“supports up to 10 reference images and can produce high-quality results”x.com
  4. SUPPORT@alibaba_qwen“Qwen-Image-2.1 is now supported in ComfyUI”x.com
  5. SUPPORT@risingsayak“Available Day-0 in Diffusers”x.com
  6. SUPPORT@sgl_project“1024×1024 generation in 18.7s and image editing in 21.7s with 22.7 GiB peak GPU memory during requests”x.com
  7. SUPPORT@ostrisai“Strict non-commercial license. No revenue cap.”x.com
  8. SOURCE@alibaba_qwen“performing removals and modifications in all three regions at once”x.com
Markdown