Repository navigation
[Diffusion] Fuse LongCat Image normalization and modulation - #38530
Merged
BBuf merged 2 commits intoSep 9, 2026
Merged
Conversation
BBuf
requested review from
AgainstEntropy,
HaiShaw,
mickqian,
ping1jing2 and
yichiche
as code owners
September 8, 2026 15:39
Collaborator
Author
This was referenced Sep 19, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reuse the existing bit-exact
fused_layernorm_modulatein LongCat's dual/single AdaLN and norm2 sites. The initial profiles showed LayerNorm followed by scale+1, multiply and shift still taking5.4%–6.7% of cumulative kernel time after the previous QKNorm/RoPE and residual-gate optimizations.The subclasses preserve Diffusers parameters/state-dict names. A per-shape/stride comparison enables fusion only after matching the original chain. Mismatch/JIT failure disables fusion; unsupported inputs, gradients, compilation and unverified capture keep the original formula. Final AdaLayerNormContinuous is unchanged. No new kernel or checkpoint-specific implementation is added.
H200 setup
1x H200, PyTorch2.11.0+cu130, native pipelines, seed42, prompt rewrite off. All three checkpoints were profiled and validated separately.
meituan-longcat/LongCat-ImageA red panda reading a book beside a sunlit window.meituan-longcat/LongCat-Image-EditMake the cat wear a red hat.meituan-longcat/LongCat-Image-Edit-TurboMeasured baseline:
295e64d7000cfbb4d3bf9f4e46efcdacbc4d84a8; candidate:7f5848461cb0da7e11a3485806fb283da97a132d(same model/test content as local77fb7a7288). PR base:482e9f257bb9d9ac57c2245c93768a96ec70edbe, whose baseline LongCat model code matches the measured baseline. Final review adds explicit gradient/compile fallback, lazy resolution of new kernel symbols, and current CI registration. Fast-kernel arithmetic is unchanged. Full-model timings below are from the recorded benchmark commits, not a new-base E2E rerun.Fresh-process A-B-B-A requests use the same GPU, request warmup and saved outputs. E2E is SGLang perf-dump
total_duration_ms; loading, warmup and profile overhead are excluded. Image lossless eager A1 reuses the immediately preceding same-configurationbase-eager-a2sample; other rows use their own recorded four-cell matrix.Eager end-to-end
Denoise stage
Breakable CUDA Graph
All four accepted lossless BCG cells captured and replayed without invalid-signature/fallback signals. Source/candidate diagnostic traces each contain186
cudaGraphLaunchcalls. BCG itself showed no extra speedup over eager here; this table compares the fusion within the same mode. Invalid or disabled BCG fallback timings are excluded.Peak reserved memory is unchanged within each mode: Image lossless eager30.990GiB / BCG32.336GiB / high eager28.613GiB; Edit/Turbo lossless31.328GiB / high29.791GiB. One Image BCG candidate had a transient14MiB extra context during loading, absent in later measurement snapshots; the second candidate was close in latency. Other final rows did not sample an extra process. Ten-second snapshots do not prove continuous exclusivity.
Kernel / profile evidence
The microbenchmark script calls the actual existing fused function vs
F.layer_norm(..., eps=1e-6) * (1+scale[:,None]) + shift[:,None], BF16[1,S,3072]:Complete GPU
ProfilerStepranges avoid asynchronous trace-edge counting. Kernel time is the mean cumulative device time of three complete profiled steps:Image shapes use S512/4096/4608; Edit/Turbo use S859/8374/9233, D3072. Turbo has one DiT call per step without CFG, so its counts are half Edit's. Scale+1, multiply and shift-add each fall by120 calls/step for Image/Edit and60 for Turbo. Device measurements explain the change; they do not replace E2E.
Output comparison
All28 formal PNGs were rehashed. Within each model/quality, baseline and PR are byte-identical (SSIM1, zero pixel error). Image lossless eager/BCG outputs also match. High vs lossless is not byte-identical:
Image
Edit
Edit-Turbo
All28 measurements and SHA256 hashes.
Reproduce
Select one idle H200 with
CUDA_VISIBLE_DEVICES. Run both commits in A-B-B-A order and repeat for both quality values.For Turbo use
meituan-longcat/LongCat-Image-Edit-Turbo, retaining its native8-step/guidance1 preset. For accepted Image BCG add--enable-breakable-cuda-graph --warmup-resolutions 1024x1024. Run separate--profile --num-profiled-timesteps 3diagnostics, excluding profiled timings.Validation
Final rebased head
cafc9addfd05ccb17a99011bc49d7b9154250200: 5 H200 tests passed, including output/gradient equality after prior signature verification.Full pre-commit on both changed files and
git diff --check: passed.Measured candidate:4H200 integration tests passed: strict Diffusers checkpoint/output compatibility, changed-input graph replay, unverified capture and mismatch fallback.
Final-head tests additionally check gradients after the signature was already verified.
Each model's task-owned retained cache was removed after completion, leaving zero files/weights/bytes: Image29,318,787,592bytes; Edit29,319,288,298bytes; Turbo29,319,187,912bytes. Per-cell overlays were removed in
finally; shared read-only seeds were preserved.CI States
Latest PR Test (Base): 🚫 Run #34246132100
Latest PR Test (Extra): ❌ Run #34246582630
Latest PR Test (AMD ROCm 7.2): ❌ Run #34246131975