Repository navigation
[Diffusion] Fuse Joy Image Edit QKV concatenation and avoid QK copies - #40494
Conversation
mickqian
left a comment
There was a problem hiding this comment.
Exact-head CI follow-up for 4811f4d:
The root failure is qwen_image_t2i_cache_dit_enabled, Denoise Step 49, in single-GPU partition 2. Initial execution plus all six testcase retries exceeded the unchanged 65.1100 ms limit (65.7389–66.9358 ms; final 66.4692 ms). This is persistent within this job, not a startup failure. The modified Joy Image paths do not establish a cause for this Qwen failure, but the log alone also does not prove a runner defect.
The downstream Base-B failures inspected are health-gate cascades from this job, not additional model-test failures. I have not changed thresholds or requested another blind retry. Please investigate the exact Qwen configuration on a comparable runner before changing the assertion. I cannot push to the head repository.
Summary
Reduce data movement around Joy Image Edit's joint attention: reuse the existing out-of-place QK-Norm/RoPE operation to avoid copying packed image Q/K, then concatenate image/text Q/K/V with one pure-copy kernel.
On 2× H200, native eager 1024×1024 / 40 steps / BF16 / lossless, two fixed ABBA groups of persistent native generators using native client-side PNG saving (
return_file_paths_only=False) reduce steady-state worker E2E by 3.16–3.27% and client E2E including pixel transport and PNG saving by 3.13–3.25%. All valid lossless images are byte-identical. The default worker-save path has allocator-release long tails and its repeated saved-client results did not qualify; no reliable default-path client speedup is claimed. All earlier measurements are retained below.Modifications
SGLANG_ENABLE_FUSED_QKNORM_ROPE=0; the joint-copy path is eligible for CUDA FP16/BF16 tensors of at least 32 MiB.The existing BBuf #34616/#34617 lossless guards and packed-input work informed this approach; they are already in the baseline. The existing out-of-place operation comes from #37903 (KevinMi) and is reused without changing its CUDA math. The actual Joy request's text branch has Q/K normalization but no RoPE. Source/prior-art review.
End-to-end benchmark
Model
jdopensource/JoyAI-Image-Edit-Diffusers@4b41fb25d961f37668750178ccbb380da326201c. 2× H200, CFG parallel 2, TP1/SP1, resident transformer/text encoder/VAE, manual performance mode, native SGLang backend, BF16,quality=lossless, no torch.compile. 1024×1024, one image, 40 steps, CFG4, seed42, promptMake the cat wear a red hat; input image.Baseline
80da4432d085ed4d6166ef643d9fd2b829dbb0c5; combined final candidate4811f4d52aa2586412f699b3bb84ed184d76250a. Torch 2.13.0+cu130, Triton 3.7.1, CUDA 13.0, driver 595.71.05. Environment.Same assigned GPUs and fixed A–B–B–A order. Every process uses 40-step same-shape native request warmup on both arms. The headline client-save protocol uses the existing native
return_file_paths_only=FalseAPI in both revisions and adds one complete generated-and-saved warmup, then five measured saved requests per process. Two groups contain 40 measured requests and 8 saved warmups, all retained. Native generation, returned pixel transport, PNG saving and reporting remain inside the client measurement. Earlier worker-save persistent groups use the same repeat count; earlier oneshot groups each measure one saved request per fresh process. Seconds; percentages are latency reductions from arithmetic means:copy-only v2is0cff694f701987ba694954a4cf54257f7c3881e6;copy-only finalis3b00154ddb33ac436906a493c9706ed3855ec918. The latter adds backend-import guards and documentation, with the same copy kernel. Copy-only final group 1 reaches only about 1.43% worker / 1.35% client and does not qualify. This prompted the profile-driven removal of the remaining image Q/K copies. Copy-only v2 group 1 contains a 15.36 s baseline client outlier; its larger client percentage is not the headline. The earlier default one-step-warmup single-request screen improved worker time by about 1.1% and also remains non-qualifying.Worker E2E is the native stage/perf-dump time before PNG saving. Client E2E is the native request-and-save timer, logged to 0.01s. Both exclude loading/warmup. No profiler request supplies an E2E row. Recorded lossless peak reserved memory: 45.83594, 45.87305 GiB per worker.
Combined oneshot group 2 has a 16.08 s candidate client outlier despite 13.824 s worker time, so that client's group does not qualify. Adding a complete saved-request warmup did not remove the issue: the worker-save persistent groups also fail saved-client qualification, including a 16.99 s candidate observation. Separate instrumented diagnostic worktrees localize a 1.76 s output-rank
empty_cachedelay in baseline, with PNG saving around 0.16 s and sub-millisecond peak collective/report writing. Candidate's twelve diagnostic requests did not reproduce its earlier long tail; the same cause for those particular earlier observations remains an inference. Startup/shutdown GC is recorded separately. The diagnostic runs are excluded from benchmark qualification.The final protocol explicitly selects native client-side PNG saving in both arms, retaining pixel data through the worker response. The existing worker condition then does not call request-tail
empty_cache; the native client saves the same PNG inside its timer. No allocator/GC function is patched, no output work is removed from the timed request, and no generation math changes. This is a supported different output configuration, with the default path's failed comparisons disclosed. Predeclared client-save protocol · Earlier persistent protocol · Default-path audit · Tail diagnosis.The persistent driver also measures an unrounded outer wall timer around native
generateplus the CLI performance dump, including input validation and the saved response. Output hashing and campaign JSON writing occur after this timer:All 24 oneshot, 40 worker-save persistent and 40 client-save persistent observations
All 8 explicitly excluded client-save warmups (earlier warmups also retained in raw data)
Raw native logs, presets, source SHAs and perf dumps · Complete audit.
Profile and kernel benchmark
Three-step native profiler requests select the third complete Joy model forward using CUDA launch correlation. These traces use the original worker-save configuration; the generation parameters and measured source revisions match the client-save benchmark, and the forward slice precedes either output transport/save path. Copy scopes are identified by their actual image/text input shapes. The optimized-chain subtotal includes image Q/K copies, image QK-Norm/RoPE and Q/K/V concatenations; unchanged position arange is excluded.
The 80 text Q/K copies and two unrelated FP32 concatenations remain. Attention, GEMMs, MLP and modulation math are unchanged. FA3 is attention; the generic analyzer's GEMM/MoE classifications do not apply to every named kernel in this dense BF16 model. Full/sliced traces and analysis.
Production helper versus the original image QK copy/norm/RoPE + three-cat chain, BF16, 32 heads ×128. Median of 40 CUDA-event samples after 10 warmups, normal output allocations, first-signature verification completed before timing:
Smaller shapes retain native operations; their eager rows expose the wrapper/guard overhead and are not claimed faster. Standalone graph is a kernel diagnostic, not native model BCG. Samples · Harness.
Focused NCU on the production shape:
The out-of-place norm itself is slightly slower; avoiding the two input copies (169.120 µs combined) compensates. Joint-copy L2 throughput is 88.94%, versus 12.18% for the original generic V cat. No measured local-memory spills. Source stalls map to the new Triton loads/stores; the reused CUDA JIT norm lacks usable source-line mapping in this report. PM raw instances are preserved, with no unsupported tail-balance claim. NCU reports, six-part analysis and limitations · Earlier copy-only NCU.
Output comparison
All 127 valid lossless PNGs are byte- and pixel-identical, including profiles and all measured revisions: SSIM 1, PSNR ∞. SHA256
76cfd40eb83feacb1b0292e171b9beb2a8cd7eb7fadc606cb1632e886d960ce6.All 5 valid high-mode PNGs are also byte-identical to lossless and to each other for this request; they are correctness checks, excluded from the lossless timing table.
Original baseline PNG · Original candidate PNG · Baseline high · Candidate high · Absolute pixel difference.
All 9 native BCG probes report
[Diffusion BCG] disabledand no capture. Their saved eager fallback images are independently hashed and excluded from performance claims. This PR improves eager execution and does not add Joy model BCG support.Validation and reproduction
Validation scope: The test counts, logs and GPU measurements below are historical results from the recorded benchmark revisions, before test cleanup. The current diff retains two joint-copy correctness/replay tests. This cleanup does not change runtime code or claim a new GPU measurement.
sys.argvbookkeeping and failed before model initialization. Retry1 supplies the same explicit CLI arguments and preserves the original failed log.The following native CLI reproduces the workload and the earlier worker-save oneshot measurement. For the headline client-save protocol, use run-joy-client-save.py and validate-joy-client-save.py; they construct the same native CLI arguments and retain a local
DiffGenerator, callinggenerate(sampling_params_kwargs={..., "return_file_paths_only": False, "save_output": True})for the saved warmup and five measured requests. Both revisions use the same pinned checkpoint and input. To select this output mode in a one-request CLI, pass a JSON config containing{"return_file_paths_only": false}with--config:Exact driver scripts · SHA256 manifest.
CI States
Latest PR Test (Base): ⏳ Run #35871593583
Latest PR Test (Extra): ❌ Run #35871592959
Latest PR Test (AMD ROCm 10): ⏳ Run #35871593623