Repository navigation
[Diffusion] Optimize Qwen-Image TP collectives and attention - #36680
Conversation
d46d461 to
2a37870
Compare
Preserve the PR's Qwen-Image NVFP4 fallback mapping while adopting main's transformer override quantization detection.
|
Updated the H100 consistency-data pin in This revision keeps the Qwen-Image fused-QKV GT from ci-data-diffusion#5, while restoring the three Sana H100 frames to the main-compatible blobs used before The Sana mismatch was unrelated to this PR's Qwen-Image runtime changes. I also merged the latest Validation performed locally:
The updated head is |
|
Applied the quality-gating follow-up in |
|
/rerun-failed-ci |
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…ject#36680) Co-authored-by: Mick <mickjagger19@icloud.com> Co-authored-by: Cursor <cursoragent@cursor.com>
…-wrapped layers sgl-project#36680 added _split_unquantized_merged_linear, which reads output_partition_sizes and weight straight off to_added_qkv. Once LoRA wraps the layer (MergedColumnParallelLinearWithLoRA) the attribute is missing, and splitting the base weight would drop the LoRA delta anyway. That broke miles_diffusion test_qwenimage_pickscore_grpo_5xGPU on sglang-miles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one MergedColumnParallelLinear (to_added_qkv). The parity patch routed every ColumnParallelLinearWithLoRA through the 2D PEFT delta, which fails on the merged layer's 3D stacked LoRA (size 8 vs 3072). Split the packed weight into three reference GEMMs, each with its own PEFT-ordered delta, which matches the separate linears PEFT trains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ity patch sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one MergedColumnParallelLinear (to_added_qkv). The parity patch forced every LoRA layer through a 2D PEFT-ordered delta, which crashes on the merged layer's 3D stacked LoRA (size 8 vs 3072). Let sgl-d's native LoRA forwards run; the norm/RoPE/split_seqs parity patches stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ity patch sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one MergedColumnParallelLinear (to_added_qkv). The parity patch forced every LoRA layer through a 2D PEFT-ordered delta, which crashes on the merged layer's 3D stacked LoRA (size 8 vs 3072). Let sgl-d's native LoRA forwards run; the norm/RoPE/split_seqs parity patches stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Attribution, split by source: - Step 0 (first rollout and the first train step): only this PR's 7d83ef7, which drops the LoRA forward overrides from the qwen_image parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks) with this branch already gives this standard's step-0 values bit for bit. - Everything after the first weight update: sgl-project/sglang#36680 (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the only contributing commit in a CI-runner bisect from 5375babbac to 5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into one to_added_qkv, so once LoRA is live the added-text Q/K/V run as one packed GEMM plus a stacked LoRA delta instead of three GEMMs. Effect: the train<->rollout residual grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series barely moves. Recorded on the h200/5gpu runner against sgl-project/sglang#40786 (run 35803865496). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ity patch sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one MergedColumnParallelLinear (to_added_qkv). The parity patch forced every LoRA layer through a 2D PEFT-ordered delta, which crashes on the merged layer's 3D stacked LoRA (size 8 vs 3072). Let sgl-d's native LoRA forwards run; the norm/RoPE/split_seqs parity patches stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Attribution, split by source: - Step 0 (first rollout and the first train step): only this PR's 7d83ef7, which drops the LoRA forward overrides from the qwen_image parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks) with this branch already gives this standard's step-0 values bit for bit. - Everything after the first weight update: sgl-project/sglang#36680 (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the only contributing commit in a CI-runner bisect from 5375babbac to 5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into one to_added_qkv, so once LoRA is live the added-text Q/K/V run as one packed GEMM plus a stacked LoRA delta instead of three GEMMs. Effect: the train<->rollout residual grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series barely moves. Recorded on the h200/5gpu runner against sgl-project/sglang#40786 (run 35803865496). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-wrapped layers sgl-project#36680 added _split_unquantized_merged_linear, which reads output_partition_sizes and weight straight off to_added_qkv. Once LoRA wraps the layer (MergedColumnParallelLinearWithLoRA) the attribute is missing, and splitting the base weight would drop the LoRA delta anyway. That broke miles_diffusion test_qwenimage_pickscore_grpo_5xGPU on sglang-miles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Attribution, split by source: - Step 0 (first rollout and the first train step): only this PR's 7d83ef7, which drops the LoRA forward overrides from the qwen_image parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks) with this branch already gives this standard's step-0 values bit for bit. - Everything after the first weight update: sgl-project/sglang#36680 (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the only contributing commit in a CI-runner bisect from 5375babbac to 5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into one to_added_qkv, so once LoRA is live the added-text Q/K/V run as one packed GEMM plus a stacked LoRA delta instead of three GEMMs. Effect: the train<->rollout residual grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series barely moves. Recorded on the h200/5gpu runner against sgl-project/sglang#40786 (run 35803865496). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Summary
This PR consolidates the Qwen-Image TP kernel work previously split across
#36680, #36693, and #36694:
dispatcher and admit 24 MiB Qwen-Image row-parallel outputs to
CustomAllReduceV2three reference GEMMs for default
quality=lossless, and mount the singlecolumn-parallel GEMM only for request-scoped
quality=high; Nunchaku keepsits existing quantized packed path
kernel instead of materializing three concatenated tensors
and scatter back to the padded/BCG bucket when needed
quantization backends, and non-CUDA platforms
skill/kernel documentation
Request quality semantics
Only the unquantized BF16 added-text QKV GEMM packing is non-bit-exact. It is
now a request-scoped quality site:
--quality=lossless(default): applies the packed Q/K/V weight slices as theoriginal three independent projection GEMMs, preserving the reference
reduction order and output;
--quality=high: mounts the single packed GEMM used for the performancenumbers below;
losslessunmounts the site at the next batch boundary.The checkpoint remains packed in memory in both modes, so this does not add a
second copy of the Q/K/V weights. The segmented pack, attention scheduler, and
collective changes keep their existing independently validated fallbacks.
Actual blast radius
This is broader than the original Qwen-only benchmark, but it does not
speed up every diffusion model or every Qwen checkpoint automatically.
quality=highon unquantized Qwen-Image architecture checkpoints: Qwen-Image, 2512, Edit, Edit-2509, Edit-2511, Layered, and FireRed Image Edit 1.0/1.1. Defaultlosslessuses three projection GEMMs; segmented packing remains independently guarded.LTX-2/2.3 and single-GPU FLUX.2 Klein do not execute these changed branches in
their tested presets; their measurements below are controls, not PR speedup
claims. Qwen/Ideogram NVFP4 are Blackwell-only and the quantized projection path
does not inherit the unquantized fused-QKV change.
Motivation and profiler evidence
The original Qwen-Image-2512 TP2 path left three independent gaps relative to
vLLM-Omni:
implementation. The 24 MiB row-parallel outputs therefore fell back to NCCL
instead of using the CUDA V2 dispatcher.
before packing. The trace contained 1,212
CatArraylaunches taking10.768 ms in the sampled window.
VarlenDynamicat 159.403 us/call, while vLLM selectedStaticPersistentat 147.402 us/call for this batch-one dense sequence.The combined path changes those hot spots as follows:
81.754 us (21.5% faster)
production batch-one shape
VarlenDynamic159.403 us ->StaticPersistent147.063 us,matching the vLLM kernel timing
quality=high, fused added-text QKV removes 1,212 small split-KGEMMs/reductions and saves about 0.43 ms per denoise step
Cross-model H200 audit
Method
a7e3f590cabb676a02181fb3328026587822f9348e056cfbfe68ddc46c1f869b7853ec9fc7e94cfemeasurements were collected before the request gate was added, when the packed
GEMM ran unconditionally; they therefore describe the current
quality=highpath. The new default
quality=losslesspath retains the reference three-GEMMadded-QKV computation. Eager with compile/BCG off unless explicitly shown
immediately after that model, including failed 403 attempts.
Full JSON summaries, output metrics, all contact sheets, and both cleanup
ledgers are in the
evidence branch.
Successfully executed affected checkpoints
Positive percentages mean lower PR latency. Bold rows cross the audit's 1.5%
signal threshold.
Qwen conclusion: no, the entire Qwen family does not get a measurable
speedup. The two T2I BCG lanes are clear wins; Edit, Edit-2509, Edit-2511,
Layered, and FireRed 1.0/1.1 are neutral or slightly slower at these fixed
shapes. The Qwen-specific code is shared, but its saved launches are too small
relative to the full denoiser on several variants, and only TP runs benefit
from the custom-all-reduce change.
ERNIE's internal prompt-enhancement generation is nondeterministic across fresh
processes on both base and PR. Its pictures are included for disclosure, but
their cross-side SSIM is not evidence of a PR correctness change. Denoise timing
is measured after the independently generated prompt.
Affected but blocked on this assignment
Negative controls (not attributed to this PR)
These controls show why an isolated ABBA shift is not generalized unless the
changed branch is proven to execute.
Generated output: Base vs PR
Every successfully executed affected checkpoint has a Base/PR sheet below.
Video sheets show the first and middle frames; Layered shows all four layers.
Qwen family (8 checkpoints)
Qwen-Image-2512
Qwen-Image
Qwen-Image-Edit
Qwen-Image-Edit-2509
Qwen-Image-Edit-2511
Qwen-Image-Layered
FireRed-Image-Edit-1.0
FireRed-Image-Edit-1.1
Z-Image, Cosmos3, ERNIE, and LingBot (7 checkpoints)
Z-Image-Turbo
Z-Image
Cosmos3-Super T2V
Cosmos3-Super I2V
ERNIE-Image (nondeterministic prompt enhancement on both sides)
ERNIE-Image-Turbo (nondeterministic prompt enhancement on both sides)
LingBot Video MoE 30B
The control sheets are retained in the evidence branch but intentionally not
presented as affected-model results.
Separate vLLM-Omni comparison lane
This is the original fixed no-CFG benchmark and is separate from the cross-model
ABBA sweep above. It used Qwen-Image-2512, 1024x1024, 50 steps, guidance scale
1, seed 42, TP2, BCG enabled, compile off, five measured requests, and the same
adjacent full-NVLink H200 pair.
The combined SGLang candidate is 7.7% faster than the fastest observed
vLLM-Omni result under that fixed workload. The fused added-QKV part of this
candidate now requires
quality=high; defaultlosslesskeeps the referenceprojection order.
Correctness and cleanup
bit-exact Base vs PR. Qwen-Image-2512 (SSIM 0.9713 / PSNR 30.15 dB) and
Qwen-Image-Layered (minimum SSIM 0.9891 across four layers) were not bit-exact,
but both were stable within each side.
quality=highBF16 Qwen added-text projection fusion changes reductionassociation. Default
quality=losslessnow runs the reference three GEMMs. Inthe original no-CFG high-path lane, the fixed-seed image measured PSNR 40.55 dB and
pixel MAE 1.393/255 against the unfused baseline.
prompt enhancement is regenerated in each fresh process; the output sheets
are disclosure rather than a correctness attribution.
and LTX-2.3's extra materialized checkpoint). Every record reports zero
remaining bytes/weight files. The task cache itself is now empty except for
its marker file.
Test plan
pre-commit run --from-ref origin/main --to-ref HEADThe BF16 reduction-order tradeoff is now isolated behind
quality=highanddocumented explicitly. Supersedes #36693 and #36694.
CI States
Latest PR Test (Base): ⏳ Run #33459254730
Latest PR Test (Extra): ⏳ Run #33459254596
Latest PR Test (AMD ROCm 7.2): ⏳ Run #33459254715