Skip to content

[Diffusion] Fuse LongCat Image normalization and modulation - #38530

Merged
BBuf merged 2 commits into
sgl-project:mainfrom
BBuf:codex/h200-longcat-norm-modulate-20260908
Sep 9, 2026
Merged

BBuf merged 2 commits into
sgl-project:mainfrom
BBuf:codex/h200-longcat-norm-modulate-20260908

Conversation

@BBuf

@BBuf BBuf commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Reuse the existing bit-exact fused_layernorm_modulate in LongCat's dual/single AdaLN and norm2 sites. The initial profiles showed LayerNorm followed by scale+1, multiply and shift still taking5.4%–6.7% of cumulative kernel time after the previous QKNorm/RoPE and residual-gate optimizations.

The subclasses preserve Diffusers parameters/state-dict names. A per-shape/stride comparison enables fusion only after matching the original chain. Mismatch/JIT failure disables fusion; unsupported inputs, gradients, compilation and unverified capture keep the original formula. Final AdaLayerNormContinuous is unchanged. No new kernel or checkpoint-specific implementation is added.

H200 setup

1x H200, PyTorch2.11.0+cu130, native pipelines, seed42, prompt rewrite off. All three checkpoints were profiled and validated separately.

Checkpoint Request Steps / guidance
meituan-longcat/LongCat-Image 1024x1024; A red panda reading a book beside a sunlit window. 50 / 4.5
meituan-longcat/LongCat-Image-Edit 1264x848; cat input; Make the cat wear a red hat. 50 / 4.5
meituan-longcat/LongCat-Image-Edit-Turbo Same editing input and prompt 8 / 1

Measured baseline: 295e64d7000cfbb4d3bf9f4e46efcdacbc4d84a8; candidate: 7f5848461cb0da7e11a3485806fb283da97a132d (same model/test content as local77fb7a7288). PR base: 482e9f257bb9d9ac57c2245c93768a96ec70edbe, whose baseline LongCat model code matches the measured baseline. Final review adds explicit gradient/compile fallback, lazy resolution of new kernel symbols, and current CI registration. Fast-kernel arithmetic is unchanged. Full-model timings below are from the recorded benchmark commits, not a new-base E2E rerun.

Fresh-process A-B-B-A requests use the same GPU, request warmup and saved outputs. E2E is SGLang perf-dump total_duration_ms; loading, warmup and profile overhead are excluded. Image lossless eager A1 reuses the immediately preceding same-configuration base-eager-a2 sample; other rows use their own recorded four-cell matrix.

Eager end-to-end

Model Quality Baseline runs PR runs Mean reduction
Image lossless 7.446897 / 7.449725 s 7.165762 / 7.175107 s 3.731%
Image high 7.206739 / 7.220042 s 6.953937 / 6.967295 s 3.504%
Edit lossless 17.175946 / 17.208413 s 16.705402 / 16.732543 s 2.752%
Edit high 16.781902 / 16.826185 s 16.356323 / 16.331516 s 2.738%
Edit-Turbo lossless 1.662449 / 1.668380 s 1.638115 / 1.631985 s 1.823%
Edit-Turbo high 1.631460 / 1.595840 s 1.587681 / 1.552630 s 2.695%

Denoise stage

Model Quality Baseline mean PR mean Reduction
Image lossless 7.323091 s 7.045838 s 3.786%
Image high 7.144771 s 6.890551 s 3.558%
Edit lossless 16.774165 s 16.292357 s 2.872%
Edit high 16.424011 s 15.956956 s 2.844%
Edit-Turbo lossless 1.324374 s 1.287778 s 2.763%
Edit-Turbo high 1.298517 s 1.260509 s 2.927%

Breakable CUDA Graph

Model / quality Baseline E2E runs PR E2E runs Reduction / status
Image / lossless 7.449374 / 7.467250 s 7.181330 / 7.194634 s 3.625%
Image / high — — Rejected: request-time linear+GELU mount differs from warmup graph
Edit and Edit-Turbo — — Native support gate disables BCG; no capture

All four accepted lossless BCG cells captured and replayed without invalid-signature/fallback signals. Source/candidate diagnostic traces each contain186 cudaGraphLaunch calls. BCG itself showed no extra speedup over eager here; this table compares the fusion within the same mode. Invalid or disabled BCG fallback timings are excluded.

Peak reserved memory is unchanged within each mode: Image lossless eager30.990GiB / BCG32.336GiB / high eager28.613GiB; Edit/Turbo lossless31.328GiB / high29.791GiB. One Image BCG candidate had a transient14MiB extra context during loading, absent in later measurement snapshots; the second candidate was close in latency. Other final rows did not sample an extra process. Ten-second snapshots do not prove continuous exclusivity.

Kernel / profile evidence

The microbenchmark script calls the actual existing fused function vs F.layer_norm(..., eps=1e-6) * (1+scale[:,None]) + shift[:,None], BF16 [1,S,3072]:

S Original Fused
512 19.0896 us 6.3901 us
4096 83.2337 us 25.9894 us
4608 96.4625 us 28.3634 us

Complete GPU ProfilerStep ranges avoid asynchronous trace-edge counting. Kernel time is the mean cumulative device time of three complete profiled steps:

Model / mode Kernels per step before -> after LayerNorm before -> after New fused calls per step Kernel time before -> after
Image / eager 1760 -> 1400 122 -> 2 120 138.15 -> 132.25 ms
Image / BCG 1760 -> 1400 122 -> 2 120 139.08 -> 133.96 ms
Edit / eager 1755 -> 1395 122 -> 2 120 328.389 -> 319.164 ms
Edit-Turbo / eager 879 -> 699 61 -> 1 60 160.640 -> 154.279 ms

Image shapes use S512/4096/4608; Edit/Turbo use S859/8374/9233, D3072. Turbo has one DiT call per step without CFG, so its counts are half Edit's. Scale+1, multiply and shift-add each fall by120 calls/step for Image/Edit and60 for Turbo. Device measurements explain the change; they do not replace E2E.

Output comparison

All28 formal PNGs were rehashed. Within each model/quality, baseline and PR are byte-identical (SSIM1, zero pixel error). Image lossless eager/BCG outputs also match. High vs lossless is not byte-identical:

Model High vs lossless SSIM PSNR
Image 0.99469157 43.53648 dB
Edit 0.99398641 40.15614 dB
Edit-Turbo 0.99290020 35.83691 dB

Image

Quality Baseline PR
lossless Baseline PR
high Baseline PR

Edit

Quality Baseline PR
lossless Baseline PR
high Baseline PR

Edit-Turbo

Quality Baseline PR
lossless Baseline PR
high Baseline PR

All28 measurements and SHA256 hashes.

Reproduce

Select one idle H200 with CUDA_VISIBLE_DEVICES. Run both commits in A-B-B-A order and repeat for both quality values.

SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 sglang generate \
  --backend=sglang --model-path meituan-longcat/LongCat-Image \
  --prompt 'A red panda reading a book beside a sunlit window.' \
  --width=1024 --height=1024 --num-inference-steps=50 --guidance-scale=4.5 \
  --seed=42 --enable-prompt-rewrite=false --performance-mode=manual \
  --quality=lossless --warmup-mode=request --save-output --perf-dump-path image.json

SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 sglang generate \
  --backend=sglang --model-path meituan-longcat/LongCat-Image-Edit \
  --image-path https://github.com/lm-sys/lm-sys.github.io/releases/download/test/TI2I_Qwen_Image_Edit_Input.jpg \
  --prompt 'Make the cat wear a red hat.' --seed=42 \
  --enable-prompt-rewrite=false --performance-mode=manual \
  --quality=lossless --warmup-mode=request --save-output --perf-dump-path edit.json

For Turbo use meituan-longcat/LongCat-Image-Edit-Turbo, retaining its native8-step/guidance1 preset. For accepted Image BCG add --enable-breakable-cuda-graph --warmup-resolutions 1024x1024. Run separate --profile --num-profiled-timesteps 3 diagnostics, excluding profiled timings.

Validation

  • Final rebased head cafc9addfd05ccb17a99011bc49d7b9154250200: 5 H200 tests passed, including output/gradient equality after prior signature verification.

  • Full pre-commit on both changed files and git diff --check: passed.

  • Measured candidate:4H200 integration tests passed: strict Diffusers checkpoint/output compatibility, changed-input graph replay, unverified capture and mismatch fallback.

  • Final-head tests additionally check gradients after the signature was already verified.

  • Each model's task-owned retained cache was removed after completion, leaving zero files/weights/bytes: Image29,318,787,592bytes; Edit29,319,288,298bytes; Turbo29,319,187,912bytes. Per-cell overlays were removed in finally; shared read-only seeds were preserved.


CI States

Latest PR Test (Base): 🚫 Run #34246132100
Latest PR Test (Extra): ❌ Run #34246582630
Latest PR Test (AMD ROCm 7.2): ❌ Run #34246131975

@github-actions github-actions Bot added the diffusion SGLang Diffusion label Sep 8, 2026
@BBuf BBuf added the run-ci CI: run the baseline test suite on this PR label Sep 8, 2026
@BBuf

BBuf commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator Author

@BBuf
BBuf merged commit 2952c8d into sgl-project:main Sep 9, 2026
228 of 283 checks passed
@BBuf
BBuf deleted the codex/h200-longcat-norm-modulate-20260908 branch September 9, 2026 03:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant