Qwen-Image-2.1_clear: Cleaner Anime Results from Alibaba’s Latest Model (Research License Only)
- Qwen-Image-2.1 is a cutting-edge image generation AI
- Qwen-Image-2.1_clear produces clean, vivid results
- It comes with a strict non-commercial license
Introduction
Hello, this is Easygoing.
Today, I’d like to introduce Alibaba’s latest image generation AI model, Qwen-Image-2.1.
What is Qwen-Image-2.1?
Qwen-Image-2.1 is a new image generation AI model released by China’s Alibaba on September 20, 2026.
gantt
title Image Generation AI Models
dateFormat YYYY-MM-DD
axisFormat %Y
tickInterval 12month
section 1st Generation<br>(SD1_vae)
Stable Diffusion 1 : 2022-08-22, 2026-10-02
section 2nd Generation<br>(SDXL_0.9_vae)
Stable Diffusion XL : 2023-07-26, 2026-10-02
section 3rd Generation<br>(Flux.1_vae)<br>(Qwen-Image_vae)
Flux.1 : 2024-08-01, 2026-10-02
Qwen-Image : 2025-08-04, 2026-10-02
Z-Image : 2025-11-25, 2026-10-02
Krea 2 : 2026-06-22, 2026-10-02
section 4th Generation<br>(Flux.2_vae)<br>(Qwen-Image-2.1_vae)
Flux.2 : 2025-11-25, 2026-10-02
ERNIE-Image : 2026-04-07, 2026-10-02
Qwen-Image-2.1 : crit, 2026-09-20, 2026-10-02Alibaba has previously released two image generation AI series—Qwen-Image and Z-Image. While Qwen-Image-2.1 is the latest addition to the Qwen-Image lineup, it is in fact a completely new image generation AI model built from the ground up, with a newly developed VAE.
Architecture of Qwen-Image-2.1
Let’s take a look at the architecture of the Qwen-Image-2.1 model.
Architecture of Qwen-Image-2.1
| Qwen-Image-2.1 | Developer | Model | License |
|---|---|---|---|
| Text Encoder | Alibaba | Qwen3-VL (8B) | Apache-2.0 |
| Transformer | Alibaba | Qwen-Image-2.1_transformer | Qwen Research License |
| VAE | Alibaba | Qwen-Image-2.1_vae | Qwen Research License |
Qwen-Image-2.1 uses Alibaba’s own Qwen3-VL as the text encoder, Qwen-Image-2.1_transformer as the Transformer, and Qwen-Image-2.1_vae as the VAE. In other words, it is a fully in-house developed model by Alibaba.
Comparison of Flux.2_vae and Qwen-Image-2.1_vae
The characteristics of an image generation AI model are largely determined by the VAE it uses.
About VAEs in Image Generation AI Models
Comparison of 3rd-Generation and 4th-Generation VAEs
| Flux.1_vae | Qwen-Image_vae | Flux.2_vae | Qwen-Image-2.1_vae | |
|---|---|---|---|---|
| Generation | 3rd Generation | 4th Generation | ||
| Original Resolution | 1024 x 1024 x 3 | 1024 x 1024 x 3 | 1024 x 1024 x 3 | 1024 x 1024 x 3 |
| Latent Resolution | 128 x 128 x 16 | 128 x 128 x 16 | 128 x 128 x 32 | 64 x 64 x 64 |
| Compression Ratio | 1/12 | 1/12 | 1/6 | 1/12 |
| Representative Models | Flux.1 Z-Image |
Qwen-Image Krea 2 |
Flux.2 ERNIE-Image |
Qwen-Image-2.1 |
The Qwen-Image-2.1_vae installed in Qwen-Image-2.1 belongs to the same 4th generation as Flux.2_vae. However, unlike the 3rd-generation VAEs that had very similar structures, the 4th-generation VAEs from Black Forest Labs and Alibaba differ significantly in architecture.
Qwen-Image-2.1_vae has twice the compression rate of Flux.2_vae. In theory, this means latent-space processing can be twice as fast as Flux.2. In practice, however, actual generation time also depends on model size and the number of steps, so a simple comparison is not possible.
Illustrations Generated with Qwen-Image-2.1_clear
For this article, I created Qwen-Image-2.1_clear, a fine-tuned version of Qwen-Image-2.1 optimized primarily for anime-style illustrations.
Below, I compare outputs from Qwen-Image-2.1_clear and the original Qwen-Image-2.1 side by side.
Autumn Leaves Path
Denim Jacket
Qwen-Image-2.1_clear has been adjusted to suppress residual noise and to increase saturation and contrast compared with the original.
When the two are compared, the left-side Qwen-Image-2.1_clear produces a cleaner, clearer illustration with noticeably less residual noise.
More Sample Images
Here are additional examples generated with Qwen-Image-2.1_clear.
Silver-Haired Mage
anime, dutch angle, turn around, 1girl, solo, long hair, looking at viewer, smile, blue eyes, dress, closed mouth, jewelry, medium white hair, earrings, puffy sleeves, necklace, black dress, dark skin, eyeshadow, blouse, gem, dark-skinned female, purple dress, glitter, empressIced Tea
1girl, long hair, black hair, dress, closed mouth, cleavage, medium upper body, multicolored hair, closed eyes, outdoors, parted lips, drink, black dress, blurred, colored hair, lacy clothes, drinking glass, drink stirrer, cafe, drinking straw, bar counter, anime, dutch angle, colorful, vivid, turn aroundYellow Outfit
realistic, photorealistic, dutch angle, turn around, young woman, looking at viewer, smile, short hair, black hair, dress, closed mouth, jewelry, brown eyes, earrings, teeth, smiling face, turtleneck, freckles, yellow dress, turtleneck sweater, ear piercing, yellow dressAlthough the text encoder of Qwen-Image-2.1 (Qwen3-VL) is officially recommended for use with natural-language prompts, in my environment tag-style prompts consistently produced higher-quality illustrations.
Qwen-Image-2.1 Comes with a Strict Non-Commercial License
Previous Alibaba image generation AI models were released under the open Apache-2.0 license. In contrast, Qwen-Image-2.1 is distributed under Alibaba’s proprietary Qwen Research License.
About the Qwen Research License
- Non-commercial research license only
- Derivative models must clearly state that they are derived from Qwen
- Users of the model assume an indemnification obligation toward Alibaba
In short, Qwen-Image-2.1 is a research-only license that explicitly prohibits commercial use.
Any derivative model of Qwen-Image-2.1 must, upon release, clearly indicate that it is a derivative of Qwen.
Furthermore, if Alibaba suffers any damage (for example, being sued) as a result of the use of Qwen-Image-2.1, the user is obligated to cover the litigation costs and compensate Alibaba for the damage.
Like the Anima model I previously introduced, the Qwen Research License is designed to minimize risk for the rights holder and is therefore quite strict for end users.
About the Dual License of NVIDIA and the Anima Model
To reiterate: Qwen-Image-2.1 is not only non-commercial; in the event of any trouble, users also bear an indemnification obligation toward Alibaba. Please keep this firmly in mind.
Qwen-Image-2.1_clear Workflow
Because of the license restrictions above, I will not release the Qwen-Image-2.1_clear model on Civitai or Hugging Face.
Instead, I am sharing the ComfyUI workflow I used when creating Qwen-Image-2.1_clear.
Workflow
Models
-
Text Encoder (ComfyUI/models/text_encoders)
- qwen3vl_8b_bf16.safetensors
- qwen3vl_8b_int8_convrot.safetensors (lightweight version)
-
Diffusion Model (ComfyUI/models/diffusion_models)
-
VAE (ComfyUI/models/vae)
Custom Nodes
ComfyUI-easygoing-nodes
If you are interested in Qwen-Image-2.1_clear, please try it yourself while strictly complying with the Qwen Research License.
Summary: Qwen-Image-2.1 Is a State-of-the-Art Image Generation AI Model
- Qwen-Image-2.1 is a cutting-edge image generation AI
- Qwen-Image-2.1_clear produces clean, vivid results
- It comes with a strict non-commercial license
In this article, I introduced the Qwen-Image-2.1_clear model.
Qwen-Image-2.1 is a fascinating model that lets you experience the latest image-generation technology. At the same time, compared with previous Alibaba models, the more restrictive license undeniably makes it harder to use.
Major Open-Weight Image & Video Generation AI Models
| Developer | UNet / Transformer | Text Encoder | VAE |
|---|---|---|---|
| Stability AI | Stable Diffusion 1.x | CLIP-L | SD1.5_vae |
| Stable Diffusion XL | CLIP-L OpenCLIP-G |
SDXL0.9_vae | |
| Stable Diffusion 3 | CLIP-L OpenCLIP-G T5-XXL-v1.1 |
SD3_vae | |
| Fal.ai | AuraFlow | pile-T5-XL | Auraflow_vae |
| Black Forest Labs | Flux.1 [schnell / dev] | CLIP-L T5-XXL-v1.1 |
Flux.1_vae |
| Flux.2 [dev] | Mistral Small 3.2 / Pixtral | Flux.2_vae | |
| Flux.2 [klein] 9B | Qwen3 | Flux.2_vae | |
| Flux.2 [klein] 4B | Qwen3 | Flux.2_vae | |
| DeepSeek | Janus-Pro | SigLIP-L DeepSeek-LLM |
LlamaGen_vq |
| Zhipu AI | CogVideoX | T5-XXL | CogVideoX_vae |
| GLM-Image | GLM-4-9B Glyph Encoder |
GLM-Image_vae | |
| Genmo | Mochi | T5-XXL-v1.1 | Mochi_vae |
| Rhymes AI | Allegro | T5-XXL | Allegro_vae |
| Lightricks | LTX-Video | T5-XXL-v1.1 | LTX-Video_vae |
| LTX-2 | Gemma 3 | LTX-2_vae | |
| LTX-2.3 | Gemma 3 | LTX-2.3_vae | |
| LTX-2.5 | Gemma 4 (Custom) | LTX-2.5_vae | |
| NVIDIA | Cosmos-Predict2 | T5-XXL | Wan2.1_vae |
| Cosmos-Predict2.5 |
Cosmos-Reason1
(Custom of Qwen2.5-VL) |
Wan2.1_vae | |
|
Cosmos3 Reasoner & Generator
(Custom of Qwen3-VL) |
Wan2.2_vae | ||
| HiDream-ai | HiDream-I1 | CLIP-L OpenCLIP-G T5-XXL-v1.1 Llama-3.1-Instruct |
Flux.1_vae |
| HiDream-O1-Image (Custom of Qwen3-VL) | |||
| Tencent | Hunyuan Video | LLaVA-LLaMA-3 CLIP-L |
HunyuanVideo_vae |
| Hunyuan Video 1.5 |
Qwen2.5-VL
byT5 SigLIP (for I2V) |
HunyuanVideo-1.5_vae | |
| Hunyuan Image 3.0 (Proprietary MLLM) | HunyuanImage-3.0_vae | ||
| StepFun | Step-Video-T2V | Hunyuan-CLIP Step-LLM |
Step-Video_vae |
| Alibaba | Wan2.1 | UMT5-XXL | Wan2.1_vae |
| Wan2.2 | UMT5-XXL | Wan2.2_vae | |
| Qwen-Image | Qwen2.5-VL | Qwen-Image_vae | |
| Qwen-Image-2.1 | Qwen3-VL | Qwen-Image-2.1_vae | |
| Z-Image | Qwen3 | Flux.1_vae | |
| CircleStone Labs | Anima | Qwen3-Base | Qwen-Image_vae |
| Baidu | ERNIE-Image | Mistral3 Pixtral |
Flux.2_vae |
| Ideogram | Ideogram 4.0 | Qwen3-VL | Flux.2_vae |
| MiniMax | MiniMax H3 (Hailuo 3.0) | Qwen3-VL | MiniMax-H3_vae |
- Bold text indicates Alibaba AI models & components
As of October 2026, many of the core components of image and video generation AI are now dominated by Chinese models. The only companies capable of developing fully open models entirely in-house are Alibaba and Tencent.
The shift of Qwen-Image-2.1 to a proprietary license may turn out to be a symbolic moment—Chinese AI models that have taken the technical lead are now beginning to move toward closed strategies.
Thank you for reading to the end!