
A sequel to Qwen-Image-2.1 vs FLUX.2 Klein — full capability matrix. Same prompts, same seeds, same scoring rubric. This post covers only what changed.
Last time Qwen-Image-2.1 beat FLUX.2 Klein 4.45 to 4.19 (0–5 scale) across 26 prompts. That score is after the correction for the three-legged jumper.
Qwen’s lead came from three things: English text, Devanagari text and edits that combine two reference images. It paid for them with time: 52 s per image against FLUX’s 28 s, and a research-only licence.
Five days later Viggle released Qwen-Image-2.1-viggle-turbo v0.2.1:
Those are exactly the two things we keep Qwen for, so we tested both.
true_cfg_scale=1.0.
Hardware: DGX Spark, GB10.| Qwen base, 40 steps | Qwen turbo, 6 steps | Qwen turbo, 8 steps | FLUX.2 Klein (first post) | |
|---|---|---|---|---|
| Text-to-image, 1024² / 1280×720 | 47–54 s | 11.3 s | 14.7 s | 28.2 s median |
| Edit, 1–2 references, 1024×1536 | 60–89 s | 20.9 s | 26.5 s | 87 s (1 ref max) |
| Peak memory | ~43 GB | ~43 GB | ~43 GB | ~46 GB |
The turbo is about 4.5× faster than base for text-to-image and 3–4× for edits.
This flips the speed argument from the first post: Qwen was the slow, careful option, and now it’s the fastest image model we run. FLUX ran on the Mac Studio, so this isn’t the same hardware, but it’s the same job.
Same rubric as last time: adherence, text, anatomy, realism, artifacts, each 0–5. I re-scored only the cells where the turbo image differs from base; everything else carries over.
| Slot | FLUX | Qwen base | Turbo 6 | Turbo 8 |
|---|---|---|---|---|
| 1 · English text | 4.06 | 4.78 | 4.67 | 4.78 |
| 2 · Devanagari | 3.33 | 4.83 | 4.83 | 4.83 |
| 3 · Portraits | 4.62 | 4.38 | 4.62 | 4.62 |
| 4 · Full body & pose | 4.12 | 4.25 | 3.88 | 4.25 |
| 5 · Groups | 4.12 | 4.38 | 4.25 | 4.38 |
| 6 · Action / expression | 4.38 | 4.38 | 5.00 | 5.00 |
| Slots 7, 8, 9, 10 | — | unchanged | unchanged | unchanged |
| All slots | 4.19 | 4.45 | 4.46 | 4.52 |
At 8 steps the turbo beats the model it was distilled from. At 6 steps it only draws level: its gains on the jump and the portraits are offset by two slips, a street-sign glitch and a hand with a finger in place of the thumb. Here’s what drives each number.
This was the error caught on manual review after the first post was published. Base Qwen’s jumper had a proper ground shadow but three legs, and FLUX’s had the right limbs but no shadow.
Both turbo versions draw a clean star jump: two legs, both fully extended as the prompt asked, a ground shadow and the whole body in frame. It’s the first image in the series that gets the physics and the anatomy right together. Slot 6 goes from a tie to a clear Qwen win.
The target was “दिल्ली में आज भारी बारिश की संभावना / मौसम विभाग ने येलो अलर्ट जारी किया”, plus “ताज़ा खबर” in a corner box.
| Errors | What went wrong | |
|---|---|---|
| Base, 40 steps | 3 | दिली (dropped the ल्ल conjunct), वारिश (ब→व), खवर (ब→व) |
| Turbo, 6 steps | 2 | वारिश: ब→व, and the final श is a malformed blend of स and श |
| Turbo, 8 steps | 1 | malformed ल्ल in दिल्ली |
It’s one seed, so I wouldn’t claim the turbo is better at Devanagari. But the specific weakness Viggle warned about didn’t show up in Hindi.
The three single-line Devanagari cases from the first post were letter-perfect in all three Qwen versions: the poster, the tea-stall board and the Hinglish lower-third. FLUX still garbles every Devanagari character.
Across all the English text, only one image had a mistake: the CODEFIRE street sign at 6 steps, where a stray stroke cuts through “COD”. At 8 steps it’s clean.
That’s the whole English text story. Everything else was letter-perfect at 6 and 8 steps, including the menu’s six prices, the slide’s four bullets, the infographic rows and the nutrition label.
The newspaper shows the pattern: every line we specified is exact in all three versions. The body columns are filler text in all three, which is fine because we never asked for body text.
The strongest complaint about base Qwen in round one was that it smooths skin and takes about ten years off. The “60-year-old with deep wrinkles” came out looking around 50.
The turbo keeps noticeably more texture: pores, deep forehead lines and crow’s feet. The same happens on the 32-year-old presenter and the laughing woman. That closes most of the portrait gap. Slot 3 now ties FLUX at 4.62, where base trailed at 4.38.
My guess at the cause: distillation strips out some of the teacher’s “prettifying” push. Whatever the reason, for presenter work it’s a clear improvement.
Viggle claims 0% layout drift, and it held for 45 of our 46 cases: same camera, same composition, same subjects as base. The exception was the street cricket shot at 6 steps. The batter changes into whites and actually raises the bat, which is closer to the prompt, and stumps appear, but a stray blob floats beside him. At 8 steps the image goes back to base’s layout, with batting pads added.
The second slip is in the typing shot. At 6 steps the right hand has a fifth finger-shaped digit where the thumb should be, a digit-count failure that base and the 8-step version both avoid. I scored its anatomy 2 out of 5, which drops the full-body slot for the 6-step turbo to 3.88, below FLUX.
Every slip in this test happened at 6 steps and was gone at 8: the sign glitch, the cricket blob and the extra finger, plus one of the Hindi errors. The two extra steps cost about 3 seconds.
On the transparent presenter sticker, base Qwen put a black top under the amber blazer. The turbo drops the top, leaving a plunging neckline, and at 8 steps adds a hand across the chest.
The transparent Sarah edit shows the same drift in a way that matters more. The turbo lowers the neckline, so the clip-on mic ends up on bare skin, attached to nothing, while base clips it to her top.
Transparency itself works fine in the turbo. The lesson is to spell out the layers (“a black crew-neck top under the blazer”) when you use it for presenter wardrobe.
Identity-preserving edits were the other weakness Viggle warned about, and I couldn’t find the problem in any of our 12 edit cases:
In each one the turbo image is almost indistinguishable from base. Face, hair, hands and pose all hold, at about a quarter of the time.
Also unchanged: groups, walking and full body; the spatial-reasoning test (cube on sphere, left of the pyramid); the mug, the courtroom sketch, and the 9:16 and 16:9 landscapes.
Two base failures the turbo did not fix:
# DGX, ~/qwenimage (same venv as round one + `peft`)
hf download Viggle/Qwen-Image-2.1-viggle-turbo \
Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors scheduler/scheduler_config.json
python turbo_ab/turbo_run.py # 2 base re-renders (bit-exact check) + 12 cases × turbo 6/8
python turbo_ab/turbo_run2.py # remaining 28 cases × turbo 6/8 + 6 dense-text × base/6/8
pipe.load_lora_weights("Viggle/Qwen-Image-2.1-viggle-turbo",
weight_name="Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors")
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained(
"Viggle/Qwen-Image-2.1-viggle-turbo", subfolder="scheduler")
SIGMAS = {6: [1.0, 0.9375, 0.875, 0.75, 0.5, 0.25],
8: [1.0, 0.96875, 0.9375, 0.90625, 0.875, 0.75, 0.5, 0.25]} # extra steps split only 1→0.875
img = pipe(prompt=..., num_inference_steps=6, sigmas=SIGMAS[6], true_cfg_scale=1.0).images[0]
Qwen-Image-2.1 and the Viggle turbo are released under the Qwen Research licence (non-commercial). Built with Qwen.