Qwen-Image-2.1 at 6 steps: does the turbo keep what beat FLUX?

A sequel to Qwen-Image-2.1 vs FLUX.2 Klein — full capability matrix. Same prompts, same seeds, same scoring rubric. This post covers only what changed.

14.7 s
per image at 8 steps (11.3 s at 6), vs 47–54 s for base Qwen
4.52
overall score at 8 steps (4.46 at 6), vs 4.45 base and 4.19 FLUX
12 / 12
identity edits held, despite the model card’s warning

Where we left off

Last time Qwen-Image-2.1 beat FLUX.2 Klein 4.45 to 4.19 (0–5 scale) across 26 prompts. That score is after the correction for the three-legged jumper.

Qwen’s lead came from three things: English text, Devanagari text and edits that combine two reference images. It paid for them with time: 52 s per image against FLUX’s 28 s, and a research-only licence.

Five days later Viggle released Qwen-Image-2.1-viggle-turbo v0.2.1:

  • It’s a distilled LoRA (1.3 GB) that runs on top of the same base weights.
  • It draws an image in 6 steps instead of 40 and drops classifier-free guidance.
  • Viggle’s model card claims about 5× the speed, with “very competitive” quality.
  • It names two weak spots: small dense text and identity-preserving edits.

Those are exactly the two things we keep Qwen for, so we tested both.

What we ran

  • Every prompt from the first round, same seeds: all 26 capability-matrix prompts plus the 14 bulletin-thumbnail and outfit-edit prompts from the experiment before it.
  • 6 new dense-text prompts: a chalkboard menu with six prices, a slide with four bullets, a newspaper front page, a five-row scores infographic, a two-line Hindi news ticker and a nutrition label.
  • Settings: the turbo at 6 steps (the recommended schedule) and at 8 steps (Viggle’s suggested fix for dense text), shipped scheduler, true_cfg_scale=1.0. Hardware: DGX Spark, GB10.
  • Baseline check: fresh 40-step renders of the base model reproduced the first round’s images bit-exactly, so the base images below are the same ones the first post scored.

Speed: the headline

Qwen base, 40 steps Qwen turbo, 6 steps Qwen turbo, 8 steps FLUX.2 Klein (first post)
Text-to-image, 1024² / 1280×720 47–54 s 11.3 s 14.7 s 28.2 s median
Edit, 1–2 references, 1024×1536 60–89 s 20.9 s 26.5 s 87 s (1 ref max)
Peak memory ~43 GB ~43 GB ~43 GB ~46 GB

The turbo is about 4.5× faster than base for text-to-image and 3–4× for edits.

This flips the speed argument from the first post: Qwen was the slow, careful option, and now it’s the fastest image model we run. FLUX ran on the Mac Studio, so this isn’t the same hardware, but it’s the same job.

Scoreboard: three slots move at 6 steps, two at 8

Same rubric as last time: adherence, text, anatomy, realism, artifacts, each 0–5. I re-scored only the cells where the turbo image differs from base; everything else carries over.

Slot FLUX Qwen base Turbo 6 Turbo 8
1 · English text 4.06 4.78 4.67 4.78
2 · Devanagari 3.33 4.83 4.83 4.83
3 · Portraits 4.62 4.38 4.62 4.62
4 · Full body & pose 4.12 4.25 3.88 4.25
5 · Groups 4.12 4.38 4.25 4.38
6 · Action / expression 4.38 4.38 5.00 5.00
Slots 7, 8, 9, 10 — unchanged unchanged unchanged
All slots 4.19 4.45 4.46 4.52

At 8 steps the turbo beats the model it was distilled from. At 6 steps it only draws level: its gains on the jump and the portraits are offset by two slips, a street-sign glitch and a hand with a finger in place of the thumb. Here’s what drives each number.

1.The three-legged jumper is fixed

jumper-pics
FLUX · Qwen base 40 · turbo 6 · turbo 8. Base has three legs; both turbo versions draw a clean two-legged star jump with a shadow.

This was the error caught on manual review after the first post was published. Base Qwen’s jumper had a proper ground shadow but three legs, and FLUX’s had the right limbs but no shadow.

Both turbo versions draw a clean star jump: two legs, both fully extended as the prompt asked, a ground shadow and the whole body in frame. It’s the first image in the series that gets the physics and the anatomy right together. Slot 6 goes from a tie to a clear Qwen win.

2. Dense Hindi: the turbo made fewer mistakes than its teacher

hindi_ticker_zoom
The ticker band at full resolution: base 40 · turbo 6 · turbo 8.

The target was “दिल्ली में आज भारी बारिश की संभावना / मौसम विभाग ने येलो अलर्ट जारी किया”, plus “ताज़ा खबर” in a corner box.

Errors What went wrong
Base, 40 steps 3 दिली (dropped the ल्ल conjunct), वारिश (ब→व), खवर (ब→व)
Turbo, 6 steps 2 वारिश: ब→व, and the final श is a malformed blend of स and श
Turbo, 8 steps 1 malformed ल्ल in दिल्ली

It’s one seed, so I wouldn’t claim the turbo is better at Devanagari. But the specific weakness Viggle warned about didn’t show up in Hindi.

The three single-line Devanagari cases from the first post were letter-perfect in all three Qwen versions: the poster, the tea-stall board and the Hinglish lower-third. FLUX still garbles every Devanagari character.

3. The only English text miss: a 6-step glitch

streetsign_codefire
A stray stroke cuts through “COD” at 6 steps; 8 steps is clean.

Across all the English text, only one image had a mistake: the CODEFIRE street sign at 6 steps, where a stray stroke cuts through “COD”. At 8 steps it’s clean.

That’s the whole English text story. Everything else was letter-perfect at 6 and 8 steps, including the menu’s six prices, the slide’s four bullets, the infographic rows and the nutrition label.

dense_newspaper_parity
Every specified line is exact in all three: masthead, date line, headline and subhead.

The newspaper shows the pattern: every line we specified is exact in all three versions. The body columns are filler text in all three, which is fine because we never asked for body text.

4. Portraits: the beautification goes away

elder_portrait
FLUX · Qwen base 40 · turbo 6 · turbo 8. The turbo keeps pores, forehead lines and crow’s feet.

The strongest complaint about base Qwen in round one was that it smooths skin and takes about ten years off. The “60-year-old with deep wrinkles” came out looking around 50.

The turbo keeps noticeably more texture: pores, deep forehead lines and crow’s feet. The same happens on the 32-year-old presenter and the laughing woman. That closes most of the portrait gap. Slot 3 now ties FLUX at 4.62, where base trailed at 4.38.

My guess at the cause: distillation strips out some of the teacher’s “prettifying” push. Whatever the reason, for presenter work it’s a clear improvement.

5. The 6-step slips: a stray blob and a missing thumb

boys_playing_cricket
At 6 steps the batter raises the bat and stumps appear, but a stray blob floats beside him. 8 steps returns to base’s layout.

Viggle claims 0% layout drift, and it held for 45 of our 46 cases: same camera, same composition, same subjects as base. The exception was the street cricket shot at 6 steps. The batter changes into whites and actually raises the bat, which is closer to the prompt, and stumps appear, but a stray blob floats beside him. At 8 steps the image goes back to base’s layout, with batting pads added.

typing_thumb
At 6 steps the right hand has a fifth finger where its thumb should be. Base and 8 steps both draw a correct hand.

The second slip is in the typing shot. At 6 steps the right hand has a fifth finger-shaped digit where the thumb should be, a digit-count failure that base and the 8-step version both avoid. I scored its anatomy 2 out of 5, which drops the full-body slot for the 6-step turbo to 3.88, below FLUX.

Every slip in this test happened at 6 steps and was gone at 8: the sign glitch, the cricket blob and the extra finger, plus one of the Hindi errors. The two extra steps cost about 3 seconds.

6. Wardrobe drift on the transparent-background images

sticker_wardrobe
Base puts a black top under the blazer; the turbo drops it.

On the transparent presenter sticker, base Qwen put a black top under the amber blazer. The turbo drops the top, leaving a plunging neckline, and at 8 steps adds a hand across the chest.

mic_zoom
With the neckline lowered, the turbo’s clip-on mic sits on bare skin, attached to nothing.

The transparent Sarah edit shows the same drift in a way that matters more. The turbo lowers the neckline, so the clip-on mic ends up on bare skin, attached to nothing, while base clips it to her top.

Transparency itself works fine in the turbo. The lesson is to spell out the layers (“a black crew-neck top under the blazer”) when you use it for presenter wardrobe.

What didn’t change, and that’s the point

two_presenters_parity
The two-reference presenter shot: turbo is almost indistinguishable from base, at 16 s instead of 60 s.

Identity-preserving edits were the other weakness Viggle warned about, and I couldn’t find the problem in any of our 12 edit cases:

  • 10 Sarah outfit edits: amber, cream, lavender and wine with a text instruction; the same four with a second reference image for the garment; navy at two seeds;
  • the newsroom background swap;
  • the two-presenter shot built from two reference images.

In each one the turbo image is almost indistinguishable from base. Face, hair, hands and pose all hold, at about a quarter of the time.

Also unchanged: groups, walking and full body; the spatial-reasoning test (cube on sphere, left of the pyramid); the mug, the courtroom sketch, and the 9:16 and 16:9 landscapes.

Two base failures the turbo did not fix:

  • The typing shot still crops out the man. All three Qwen versions show only hands and a sleeve.
  • The kite is still a bird of prey, not the red kite FLUX drew.

Verdict

  1. Turbo at 8 steps replaces base as the default for everything we used Qwen for: English and Devanagari text, two-reference presenter shots, outfit edits and native transparency. At about 15 s per image it’s 3.5× faster than base, and it scored higher (4.52 vs 4.45).
  2. Use 6 steps for quick drafts only. It’s another 3 seconds faster, but every slip in this test happened at 6 steps: a text glitch, a stray blob, an extra finger and one more Hindi error.
  3. Keep base 40 around as a fallback for multi-constraint edits, which Viggle still flags, even though our cases didn’t expose the problem.
  4. The licence hasn’t changed. The turbo inherits the Qwen Research licence, so it’s non-commercial only. That limits where any of this can ship, exactly as it did for base.

Caveats

  • One seed per prompt. This is a strong impression, not a statistical result. Viggle’s own 96-prompt evaluation is the better source for averages.
  • Scored by the same reviewer against the same rubric, re-scoring only images that visibly differ from base. Small differences in scoring judgement can move a slot by ±0.1.
  • Different hardware for FLUX: its timings are from the Mac Studio, so treat the cross-model speed comparison as practical, not controlled.

Reproduce

# DGX, ~/qwenimage (same venv as round one + `peft`)
hf download Viggle/Qwen-Image-2.1-viggle-turbo \
   Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors scheduler/scheduler_config.json
python turbo_ab/turbo_run.py     # 2 base re-renders (bit-exact check) + 12 cases × turbo 6/8
python turbo_ab/turbo_run2.py    # remaining 28 cases × turbo 6/8 + 6 dense-text × base/6/8
pipe.load_lora_weights("Viggle/Qwen-Image-2.1-viggle-turbo",
    weight_name="Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors")
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained(
    "Viggle/Qwen-Image-2.1-viggle-turbo", subfolder="scheduler")
SIGMAS = {6: [1.0, 0.9375, 0.875, 0.75, 0.5, 0.25],
          8: [1.0, 0.96875, 0.9375, 0.90625, 0.875, 0.75, 0.5, 0.25]}  # extra steps split only 1→0.875
img = pipe(prompt=..., num_inference_steps=6, sigmas=SIGMAS[6], true_cfg_scale=1.0).images[0]

Qwen-Image-2.1 and the Viggle turbo are released under the Qwen Research licence (non-commercial). Built with Qwen.