Two models, identical prompts, ten capability categories. One is a 7B research-licensed model running 40 bf16 steps on a DGX Spark; the other a 9B non-commercial model running 4 MLX steps on a Mac. This is what happened when we asked both to do the same jobs.
| Models | Qwen-Image-2.1 7B  vs FLUX.2 Klein 9B |
| Slots | 10 capability categories, 26 prompt/seed jobs |
| Hardware | DGX Spark GB10 (CUDA 13, aarch64)  vs Mac Studio M3 Ultra (mflux/MLX) |
| Prompts | Byte-identical — both runners import the same
matrix.py
|
| Scoring | 0–5 on prompt adherence, text correctness, anatomy, realism, artifacts |
| Overall | Qwen 4.49 / 5  · FLUX 4.19 / 5 |
Qwen wins the headline number by 0.30 points — but that gap is almost entirely text. Strip out the two text slots and FLUX is ahead. Qwen is best understood not as “the better image model” but as a text-rendering specialist that also happens to do images. It also ships under a research-only licence, which matters if you intend to use the model itself commercially (see section 5).
| # | Category | Qwen | FLUX | Winner |
|---|---|---|---|---|
| 1 | Text rendering — English | 4.78 | 4.06 | Qwen |
| 2 | Text rendering — Hindi (Devanagari) | 4.83 | 3.33 | Qwen |
| 3 | Humans — portrait | 4.38 | 4.62 | FLUX |
| 4 | Humans — full body & pose | 4.25 | 4.12 | tie → FLUX |
| 5 | Humans — groups & diversity | 4.38 | 4.12 | Qwen |
| 6 | Humans — action / expression | 4.88 | 4.38 | Qwen |
| 7 | Reference-image editing | 4.38 | 4.06 | Qwen |
| 8 | RGBA / transparent output | 4.29 | 4.00 | Qwen |
| 9 | Aspect ratios & speed | 3.56 | 5.00 | FLUX |
| 10 | Prompt adherence stress | 5.00 | 5.00 | tie → FLUX |
| ALL SLOTS | 4.49 | 4.19 |
Ties are decided in FLUX’s favour when the margin is under 0.15 — it is faster and its outputs carry fewer restrictions. Without that rule the raw means are Qwen 4.49 / FLUX 4.19.
| Qwen-Image-2.1 | FLUX.2 Klein | |
|---|---|---|
| English text in pixels | 4/4 strings letter-perfect | 0/4 — every one has a character slip |
| Devanagari | 55/55 characters correct | 55/55 wrong (confident gibberish) |
| Identity preservation (edits) | 5/5 | 3/5 — re-makes the face |
| Instruction-following (edits) | 3/5 — ignores garment detail | 5/5 |
| Multi-reference (2 people) | Yes | Impossible — single-ref API |
| Native alpha channel | Yes, one pass | No — needs rembg, leaves green fringe |
| Landscapes / aspect ratios | Weak (2/5) | Excellent (5/5) |
| Skin / portrait realism | Good, smoothed | Best in test |
| Speed (median) | 51.8 s/image | 28.2 s/image |
| Speed at 9:16 tall | 104 s | 231 s |
| Licence | Research only — no commercial use | Non-commercial 9B weights; outputs cleared for commercial use |
3 SIkills, Skills Th,
Technlogies, latancy) are all single-character
slips in otherwise crisp typography — which is worse, not
better, because the image looks right until a human reads it.