Alibaba's newest image model utilizes a 7B parameter architecture to combine generation and editing within a single checkpoint. The open-weight system enables native RGBA transparency for seamless compositing and accepts 10 reference images to maintain fidelity in portraits and products. In the Image Edit Arena, the model earned 1,367 points, placing it #16 overall and just behind GPT-Image-1.5-high-fidelity.

Optimization by Unsloth AI allows the model to operate on 12GB of VRAM locally, with a Dynamic FP8 version that reduces requirements to 6GB via offloading. The release includes day-0 support for vLLM, ComfyUI, and Diffusers, though a strict non-commercial license prohibits any revenue-generating activities.

Sign in to suggest edits

Key sources

  1. SOURCE@arena“took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15”x.com
  2. SOURCE@unslothai“can now run locally on 12GB VRAM with Unsloth GGUFs!”x.com
  3. SUPPORT@azmaeensami“tested Alibaba's brand-new Qwen-Image-2.1 (Full BF16) on an RTX 4080 via ComfyUI”x.com
  4. SOURCEmarketbrief.now
Markdown