Qwen-Image-2.1 vs FLUX.2 Klein

Pranjal SrivastavaSep 22, 2026
5 min
Qwen-Image-2.1 vs FLUX.2 Klein

Two models, identical prompts, ten capability categories. One is a 7B research-licensed model running 40 bf16 steps on a DGX Spark; the other a 9B non-commercial model running 4 MLX steps on a Mac. This is what happened when we asked both to do the same jobs.

Qwen-Image-2.1 7B
4.49
out of 5 · wins text & identity
FLUX.2 Klein 9B
4.19
out of 5 · wins realism & speed
Gap
0.30
almost entirely text rendering
ModelsQwen-Image-2.1 7B  vs  FLUX.2 Klein 9B
Slots10 capability categories, 26 prompt/seed jobs
HardwareDGX Spark GB10 (CUDA 13, aarch64)  vs  Mac Studio M3 Ultra (mflux/MLX)
PromptsByte-identical — both runners import the same matrix.py
Scoring0–5 on prompt adherence, text correctness, anatomy, realism, artifacts
OverallQwen 4.49 / 5  ·  FLUX 4.19 / 5

1. High-level comparison

The bottom line first

Qwen wins the headline number by 0.30 points — but that gap is almost entirely text. Strip out the two text slots and FLUX is ahead. Qwen is best understood not as “the better image model” but as a text-rendering specialist that also happens to do images. It also ships under a research-only licence, which matters if you intend to use the model itself commercially (see section 5).

Category scorecard

#CategoryQwenFLUXWinner
1Text rendering — English4.784.06Qwen
2Text rendering — Hindi (Devanagari)4.833.33Qwen
3Humans — portrait4.384.62FLUX
4Humans — full body & pose4.254.12tie → FLUX
5Humans — groups & diversity4.384.12Qwen
6Humans — action / expression4.884.38Qwen
7Reference-image editing4.384.06Qwen
8RGBA / transparent output4.294.00Qwen
9Aspect ratios & speed3.565.00FLUX
10Prompt adherence stress5.005.00tie → FLUX
ALL SLOTS4.494.19

Ties are decided in FLUX’s favour when the margin is under 0.15 — it is faster and its outputs carry fewer restrictions. Without that rule the raw means are Qwen 4.49 / FLUX 4.19.

What each model actually buys you

Qwen-Image-2.1FLUX.2 Klein
English text in pixels4/4 strings letter-perfect0/4 — every one has a character slip
Devanagari55/55 characters correct55/55 wrong (confident gibberish)
Identity preservation (edits)5/53/5 — re-makes the face
Instruction-following (edits)3/5 — ignores garment detail5/5
Multi-reference (2 people)YesImpossible — single-ref API
Native alpha channelYes, one passNo — needs rembg, leaves green fringe
Landscapes / aspect ratiosWeak (2/5)Excellent (5/5)
Skin / portrait realismGood, smoothedBest in test
Speed (median)51.8 s/image28.2 s/image
Speed at 9:16 tall104 s231 s
LicenceResearch only — no commercial useNon-commercial 9B weights; outputs cleared for commercial use

The three things Qwen wins

  1. English text you can trust. Including a 38-character headline rendered exactly at two different seeds. FLUX’s errors (3 SIkills, Skills Th, Technlogies, latancy) are all single-character slips in otherwise crisp typography — which is worse, not better, because the image looks right until a human reads it.
  2. Devanagari at all. Not a quality gap — a capability gap. FLUX emits Devanagari-shaped noise, crisply and confidently. Correct nuqta, correct matras, correct shirorekha.
  3. Multi-reference identity. Two named people in one frame. FLUX’s edit endpoint accepts exactly one reference image, so this task cannot be expressed through it at all.

Don't Forget to share this post...!

Schedule a consultation with our AI safety experts today.

Book Your Free Consultation