Skip to content

[Diffusion] Accelerate Cosmos3 Edge on Hopper with lossless fusions - #40386

Merged
BBuf merged 5 commits into
sgl-project:mainfrom
BBuf:diffusion/cosmos3-edge-hopper-fusions
Sep 23, 2026
Merged

BBuf merged 5 commits into
sgl-project:mainfrom
BBuf:diffusion/cosmos3-edge-hopper-fusions

Conversation

@BBuf

@BBuf BBuf commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Cosmos3 Edge's single-image path still uses separate Q/K normalization, RoPE and KV concatenation on Hopper. Its dense MLP also materializes ReLU before a separate square. Native profiles identify both chains in all 28 generation blocks, with separate positive and negative CFG passes.

This enables the existing rounded QK norm/RoPE/KV packing for the validated Edge shape and reuses the unary ReLU² kernel with fast math disabled. On one H200, two native A–B–B–A groups reduce image worker E2E by 29.61–32.23% and saved-image client E2E by 26.92–30.59%. Video worker E2E improves about 2.2%. All generated images and all 81 video frames remain exact.

Modifications

  • Admit the 2048-wide dense Edge T=1 shape to the existing Hopper TP1/SP1 QK norm/RoPE/KV-pack path. Preserve the rounded BF16 cache and Edge's separately normalized UND keys. Larger dense, parallel and compiled configurations retain the previous gate.
  • Expose the existing activation compiler's fast_math option through the unary wrapper, defaulting to the existing behavior. Cosmos contiguous BF16 CUDA inference uses fast_math=False for ReLU²; other model paths retain Torch.
  • Add finite-domain, production-shape, storage-ownership, changed-input CUDA-graph replay and fallback tests, plus a registered marker benchmark. No CUDA arithmetic implementation or checkpoint change.

The first fast-math experiment flushed 511 BF16 subnormal products to zero and was rejected before native combined benchmarks. Disabling fast math preserves all 65,280 finite BF16 encodings, including signed zero and subnormal results. Rejected experiment, source review.

This builds on the actual rounded Cosmos T1 and Hopper Nano implementations in #34932 and #36571, following the lossless fusion approach in #38530. Edge's architecture comes from #31590. Closed #34618 does not provide a usable BCG path in the current custom denoising stage.

Benchmark

nvidia/Cosmos3-Edge@344d602b128d1bbdacb43b08d0a3626f46343e29, 1× H200, TP1/SP1, BF16, native eager, quality=lossless, manual performance mode, GPU-resident components, no torch.compile. Native same-shape request warmup, two warmup steps; fixed sampling in both arms:

Workload Output Steps / CFG / seed Prompt
T2I 640×640, 1 frame 35 / 7 / 0 A warehouse robot folds a blue cloth on a clean workbench.
T2V 832×480, 81 frames, 24 fps (3.375 s) 35 / 5 / 42 A warehouse robot carefully places a blue box on a shelf.

Baseline 993d1fccbaafe3e79d91567d2fc1d665cc94fa50; candidate 18f1417a8b189fcb80b38d6ca3b487ecca89f39d. Torch2.13.0+cu130, Triton3.7.1, CUDA13.0, driver595.71.05. Environment. Both arms use the existing native helper's Cosmos guardrail setting; the checkpoint already disables its safety checker.

Fresh native CLI processes, same idle GPU, two A–B–B–A groups per workload. Profile and microbenchmark runs are excluded. Seconds, with reduction from each group's arithmetic means:

Workload / repeat Worker E2E Denoise Client E2E including saved output
T2I-R1 0.9986 → 0.6768 (32.23%) 0.9653 → 0.6435 (33.33%) 1.095 → 0.760 (30.59%)
T2I-R2 0.9588 → 0.6749 (29.61%) 0.9253 → 0.6414 (30.68%) 1.040 → 0.760 (26.92%)
T2V-R1 6.2802 → 6.1386 (2.25%) 4.5535 → 4.4176 (2.98%) 6.625 → 6.475 (2.26%)
T2V-R2 6.2865 → 6.1464 (2.23%) 4.5580 → 4.4190 (3.05%) 7.550 → 6.490 (14.04%)

T2V-R2 contains a baseline client-only outlier (A1 8.47 s, worker 6.2879 s). All rows are retained; its inflated client percentage is not treated as a repeatable benefit. T2V-R1 saved-client reduction is about 2.3%; both video groups have consistent worker reductions. The image result independently exceeds 1.5% in both worker and saved-client E2E in both groups.

All 16 unprofiled native measurements
Group Arm Worker E2E Denoise Client E2E incl. save
T2I-R1 A1 1.0380 1.0044 1.15
T2I-R1 B1 0.6858 0.6522 0.77
T2I-R1 B2 0.6677 0.6349 0.75
T2I-R1 A2 0.9591 0.9261 1.04
T2I-R2 A1 0.9567 0.9235 1.04
T2I-R2 B1 0.6775 0.6439 0.76
T2I-R2 B2 0.6722 0.6389 0.76
T2I-R2 A2 0.9609 0.9270 1.04
T2V-R1 A1 6.2796 4.5538 6.62
T2V-R1 B1 6.1399 4.4225 6.48
T2V-R1 B2 6.1374 4.4128 6.47
T2V-R1 A2 6.2807 4.5532 6.63
T2V-R2 A1 6.2879 4.5580 8.47
T2V-R2 B1 6.1471 4.4184 6.49
T2V-R2 B2 6.1458 4.4196 6.49
T2V-R2 A2 6.2851 4.5580 6.63

Worker E2E is perf-dump stage duration before saving, excluding loading and warmup. Client E2E is the native scheduler-request/output-save timer, logged to 0.01 s. Peak reserved memory is unchanged: 8.9258 GiB T2I / 13.2598 GiB T2V. Raw perf dumps, source SHAs, configs and native logs, complete timing audit.

The current PR head 1430010b7b20f9d56d6d6d4caeebee1980e28c73 adds only a test coverage correction after the measured head: the newly enabled Hopper split-path parity checks are registered on the Hopper CI lane and skipped on other architectures. B200 continues to run all four ReLU² tests, which passed in the first CI run. That run exposed three failures from applying the Hopper split-path bitwise contract to the unchanged Blackwell QK path. Runtime files and all benchmark measurements are unchanged.

The corrected seven-test file passed on H200 (14.27 s). In the next CI run, the H100, B200 and BCG jobs were skipped by fast-fail after the unrelated CPU fixture failure: test_auxiliary_output.py supplies a mock batch without split_index, which the existing scheduler reads. Those skipped GPU jobs are not recorded as passing validation.

The T1-only ablation also has exact images: worker mean 0.9589 → 0.6666 s, client 1.040 → 0.765 s. This is a separate exploratory group at 8d329ce689; the headline uses the final combined head and its repeated groups.

Profile and operator benchmark

Each slice is one complete second CFG iteration, covering both positive and negative passes. CPU model-call boundaries are joined to GPU execution through CUDA launch correlations; all five retained slices have zero boundary-crossing GPU kernels. The baseline video path already uses the QK fusion; its gain comes from ReLU².

Workload All kernels, before → after Eager ReLU + square calls / ms Fused ReLU² calls / ms Candidate fused QK/RoPE calls / ms
T2I 1996 → 708 112 / 0.4350 56 / 0.2138 56 / 0.2485
T2V 708 → 652 112 / 8.2427 56 / 3.8970 56 / 3.7125

T1-only reduces complete-iteration kernels from 1996 → 764. Final ReLU² fusion removes another 56 launches. T2I is sensitive to launch overhead; T2V is largely GPU busy. Cumulative profile device times explain the changed operators and are not E2E measurements. Original and complete-iteration traces, attribution and three-table triage.

Committed-wrapper marker benchmark, rotating inputs, median microseconds, including dispatch and normal output allocation:

Operation / shape Eager Fused
QK norm + RoPE + KV pack, S400, Q16/KV8, D128 64.9402 4.7117
QK norm + RoPE + KV pack, S1024, Q16/KV8, D128 85.8332 8.9444
ReLU², [1, 400, 9216] 8.6710 4.0407
ReLU², [1, 8190, 9216] 146.1616 71.6109

Raw marker output. Mutable Q/K views belong to the current forward; V and cached UND K/V remain unchanged, as tested.

Output comparison

The audit covers 39 output records: 33 valid eager requests, plus six explicitly unsupported BCG requests. Every valid image/video, including separate lossless/high checks and profiled requests, is byte-identical within its workload. All videos independently decode to 81/81 identical frames, SSIM1, PSNR∞, no audio stream. Before/after image and video contact sheets were visually inspected. The T2I native result contains a blue cloth/workbench but no visible robot in either arm; this comparison establishes output preservation.

Baseline and optimized image

Original baseline PNG · Original optimized PNG

Baseline and optimized video: frames 0, 26, 53, 80

Original baseline MP4 · Original optimized MP4 · Per-frame hashes

PNG SHA256: 5a3849ada21e0799fefa4c571cdf9cf005eab33b013eeb074c2f6c0782e67b23. MP4 SHA256: 6d7403a334c8a4f7487ee13e13e7c707898aecd5eb2529f34408a76b26f64350.

BCG was actually attempted for T2I/T2V and in the lossless/high T2I matrix. Every attempt reports [diffusion bcg] disabled with no capture, so fallback outputs are retained only as applicability evidence. No Cosmos BCG speedup is claimed. High eager T2I/T2V checks are exact, with no separate high-mode performance claim.

Validation and reproduction

  • 145 Cosmos/kernel tests and 46 subtests passed, including the new CUDA tests and existing model tests. Log.
  • 85 existing unary activation compatibility tests passed. Log.
  • All five changed files passed pre-commit and CI-registration validation.
  • Model weights were deleted after all native/profile/output checks: zero residual files, weight files or bytes. Shared caches were not modified. Cleanup audit.

Run each clean revision with the same pinned checkpoint:

CUDA_VISIBLE_DEVICES=0 SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 \
SGLANG_DISABLE_COSMOS3_GUARDRAILS=1 \
sglang generate --backend sglang --model-path nvidia/Cosmos3-Edge \
  --revision 344d602b128d1bbdacb43b08d0a3626f46343e29 \
  --num-gpus 1 --tp-size 1 --enable-torch-compile=false \
  --performance-mode manual --quality lossless \
  --width 640 --height 640 --num-frames 1 --num-inference-steps 35 \
  --guidance-scale 7 --seed 0 \
  --prompt 'A warehouse robot folds a blue cloth on a clean workbench.' \
  --warmup-mode request --warmup-steps 2 \
  --warmup-resolutions 640x640 --warmup-num-frames 1 \
  --save-output --perf-dump-path perf.json --output-path outputs --output-file-name sample

For T2V, use its prompt, --width 832 --height 480 --num-frames 81 --fps 24 --guidance-scale 5 --seed 42 --warmup-resolutions 832x480 --warmup-num-frames 81. Add --profile --num-profiled-timesteps 3 only for separate diagnostic requests. Exact native runner, pinned preparation, audit and extraction scripts, SHA256 manifest.

CI follow-up (2026-09-21)

Merged main 80da4432d085ed4d6166ef643d9fd2b829dbb0c5, including the upstream scheduler fixture repair #40411 and offline DFlash fixture repair #40427. Current head: 25ff8a99305c14fce39b73f4b5219eb6ef07bc7a. Merged-head validation and all logs, source delta in this PR's files, pre-commit. The original E2E and output comparisons above remain measurements of their explicitly pinned benchmark heads; this follow-up does not claim new model-request timings. The common repaired CPU fixture suite passes 49 tests + 57 subtests.


CI States

Latest PR Test (Base): ❌ Run #35525995597
Latest PR Test (Extra): ❌ Run #35525995362
Latest PR Test (AMD ROCm 10): ❌ Run #35525995657

@BBuf BBuf added the run-ci CI: run the baseline test suite on this PR label Sep 19, 2026
@BBuf
BBuf requested a review from mickqian as a code owner September 19, 2026 22:51
@BBuf BBuf added the diffusion SGLang Diffusion label Sep 19, 2026
@BBuf BBuf added the run-ci CI: run the baseline test suite on this PR label Sep 19, 2026
@BBuf BBuf added the diffusion SGLang Diffusion label Sep 19, 2026
@BBuf
BBuf requested a review from yuan-luo as a code owner September 19, 2026 22:51

@mickqian mickqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exact-head CI follow-up (1430010): test_disagg_idle_gap_0_prefill_normal fails because the Scheduler fixture lacks _prev_prefill_end_ts. Log: https://github.com/sgl-project/sglang/actions/runs/35475607241/job/105984387117 . This CPU test-fixture issue is fixed on main by #40411. Please merge current main and run CI on the resulting head; rerunning this unchanged head will not repair the fixture. I cannot push to the head repository. This diagnosis concerns this CPU failure only and does not establish that every GPU failure is unrelated.

@mickqian mickqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exact-head CI follow-up for 25ff8a9:

The current blocker is Base-C B200 partition 0: test_glm53_flash_b200.py / test_gsm8k, score 0.928 and then 0.924 on the built-in retry, below 0.930. This supersedes the earlier CPU-fixture diagnosis on the old head. Other inspected Base-C failures are health-gate cascades.

The same GLM test also fails on #40486 (0.926/0.920), whose changes are Cosmos-specific. That is evidence of a shared CI/model issue, not proof of a runner defect or a regression from this PR. This PR also changes a shared unary activation API (preserving the default fast_math=True), so its scope alone is not enough to declare every serving failure unrelated. I have not changed accuracy thresholds or blindly rerun it. I cannot push to the head repository.

@BBuf

BBuf commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator Author

@BBuf
BBuf merged commit 172b1b4 into sgl-project:main Sep 23, 2026
177 of 207 checks passed
@BBuf
BBuf deleted the diffusion/cosmos3-edge-hopper-fusions branch September 23, 2026 13:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion jit-kernel run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants