Repository navigation
[Diffusion] Accelerate Cosmos3 Edge on Hopper with lossless fusions - #40386
Conversation
mickqian
left a comment
There was a problem hiding this comment.
Exact-head CI follow-up (1430010): test_disagg_idle_gap_0_prefill_normal fails because the Scheduler fixture lacks _prev_prefill_end_ts. Log: https://github.com/sgl-project/sglang/actions/runs/35475607241/job/105984387117 . This CPU test-fixture issue is fixed on main by #40411. Please merge current main and run CI on the resulting head; rerunning this unchanged head will not repair the fixture. I cannot push to the head repository. This diagnosis concerns this CPU failure only and does not establish that every GPU failure is unrelated.
mickqian
left a comment
There was a problem hiding this comment.
Exact-head CI follow-up for 25ff8a9:
The current blocker is Base-C B200 partition 0: test_glm53_flash_b200.py / test_gsm8k, score 0.928 and then 0.924 on the built-in retry, below 0.930. This supersedes the earlier CPU-fixture diagnosis on the old head. Other inspected Base-C failures are health-gate cascades.
The same GLM test also fails on #40486 (0.926/0.920), whose changes are Cosmos-specific. That is evidence of a shared CI/model issue, not proof of a runner defect or a regression from this PR. This PR also changes a shared unary activation API (preserving the default fast_math=True), so its scope alone is not enough to declare every serving failure unrelated. I have not changed accuracy thresholds or blindly rerun it. I cannot push to the head repository.
Motivation
Cosmos3 Edge's single-image path still uses separate Q/K normalization, RoPE and KV concatenation on Hopper. Its dense MLP also materializes ReLU before a separate square. Native profiles identify both chains in all 28 generation blocks, with separate positive and negative CFG passes.
This enables the existing rounded QK norm/RoPE/KV packing for the validated Edge shape and reuses the unary ReLU² kernel with fast math disabled. On one H200, two native A–B–B–A groups reduce image worker E2E by 29.61–32.23% and saved-image client E2E by 26.92–30.59%. Video worker E2E improves about 2.2%. All generated images and all 81 video frames remain exact.
Modifications
fast_mathoption through the unary wrapper, defaulting to the existing behavior. Cosmos contiguous BF16 CUDA inference usesfast_math=Falsefor ReLU²; other model paths retain Torch.The first fast-math experiment flushed 511 BF16 subnormal products to zero and was rejected before native combined benchmarks. Disabling fast math preserves all 65,280 finite BF16 encodings, including signed zero and subnormal results. Rejected experiment, source review.
This builds on the actual rounded Cosmos T1 and Hopper Nano implementations in #34932 and #36571, following the lossless fusion approach in #38530. Edge's architecture comes from #31590. Closed #34618 does not provide a usable BCG path in the current custom denoising stage.
Benchmark
nvidia/Cosmos3-Edge@344d602b128d1bbdacb43b08d0a3626f46343e29, 1× H200, TP1/SP1, BF16, native eager,quality=lossless, manual performance mode, GPU-resident components, no torch.compile. Native same-shape request warmup, two warmup steps; fixed sampling in both arms:Baseline
993d1fccbaafe3e79d91567d2fc1d665cc94fa50; candidate18f1417a8b189fcb80b38d6ca3b487ecca89f39d. Torch2.13.0+cu130, Triton3.7.1, CUDA13.0, driver595.71.05. Environment. Both arms use the existing native helper's Cosmos guardrail setting; the checkpoint already disables its safety checker.Fresh native CLI processes, same idle GPU, two A–B–B–A groups per workload. Profile and microbenchmark runs are excluded. Seconds, with reduction from each group's arithmetic means:
T2V-R2 contains a baseline client-only outlier (A1 8.47 s, worker 6.2879 s). All rows are retained; its inflated client percentage is not treated as a repeatable benefit. T2V-R1 saved-client reduction is about 2.3%; both video groups have consistent worker reductions. The image result independently exceeds 1.5% in both worker and saved-client E2E in both groups.
All 16 unprofiled native measurements
Worker E2E is perf-dump stage duration before saving, excluding loading and warmup. Client E2E is the native scheduler-request/output-save timer, logged to 0.01 s. Peak reserved memory is unchanged: 8.9258 GiB T2I / 13.2598 GiB T2V. Raw perf dumps, source SHAs, configs and native logs, complete timing audit.
The current PR head
1430010b7b20f9d56d6d6d4caeebee1980e28c73adds only a test coverage correction after the measured head: the newly enabled Hopper split-path parity checks are registered on the Hopper CI lane and skipped on other architectures. B200 continues to run all four ReLU² tests, which passed in the first CI run. That run exposed three failures from applying the Hopper split-path bitwise contract to the unchanged Blackwell QK path. Runtime files and all benchmark measurements are unchanged.The corrected seven-test file passed on H200 (14.27 s). In the next CI run, the H100, B200 and BCG jobs were skipped by fast-fail after the unrelated CPU fixture failure:
test_auxiliary_output.pysupplies a mock batch withoutsplit_index, which the existing scheduler reads. Those skipped GPU jobs are not recorded as passing validation.The T1-only ablation also has exact images: worker mean 0.9589 → 0.6666 s, client 1.040 → 0.765 s. This is a separate exploratory group at
8d329ce689; the headline uses the final combined head and its repeated groups.Profile and operator benchmark
Each slice is one complete second CFG iteration, covering both positive and negative passes. CPU model-call boundaries are joined to GPU execution through CUDA launch correlations; all five retained slices have zero boundary-crossing GPU kernels. The baseline video path already uses the QK fusion; its gain comes from ReLU².
T1-only reduces complete-iteration kernels from 1996 → 764. Final ReLU² fusion removes another 56 launches. T2I is sensitive to launch overhead; T2V is largely GPU busy. Cumulative profile device times explain the changed operators and are not E2E measurements. Original and complete-iteration traces, attribution and three-table triage.
Committed-wrapper marker benchmark, rotating inputs, median microseconds, including dispatch and normal output allocation:
Raw marker output. Mutable Q/K views belong to the current forward; V and cached UND K/V remain unchanged, as tested.
Output comparison
The audit covers 39 output records: 33 valid eager requests, plus six explicitly unsupported BCG requests. Every valid image/video, including separate lossless/high checks and profiled requests, is byte-identical within its workload. All videos independently decode to 81/81 identical frames, SSIM1, PSNR∞, no audio stream. Before/after image and video contact sheets were visually inspected. The T2I native result contains a blue cloth/workbench but no visible robot in either arm; this comparison establishes output preservation.
Original baseline PNG · Original optimized PNG
Original baseline MP4 · Original optimized MP4 · Per-frame hashes
PNG SHA256:
5a3849ada21e0799fefa4c571cdf9cf005eab33b013eeb074c2f6c0782e67b23. MP4 SHA256:6d7403a334c8a4f7487ee13e13e7c707898aecd5eb2529f34408a76b26f64350.BCG was actually attempted for T2I/T2V and in the lossless/high T2I matrix. Every attempt reports
[diffusion bcg] disabledwith no capture, so fallback outputs are retained only as applicability evidence. No Cosmos BCG speedup is claimed. High eager T2I/T2V checks are exact, with no separate high-mode performance claim.Validation and reproduction
Run each clean revision with the same pinned checkpoint:
CUDA_VISIBLE_DEVICES=0 SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 \ SGLANG_DISABLE_COSMOS3_GUARDRAILS=1 \ sglang generate --backend sglang --model-path nvidia/Cosmos3-Edge \ --revision 344d602b128d1bbdacb43b08d0a3626f46343e29 \ --num-gpus 1 --tp-size 1 --enable-torch-compile=false \ --performance-mode manual --quality lossless \ --width 640 --height 640 --num-frames 1 --num-inference-steps 35 \ --guidance-scale 7 --seed 0 \ --prompt 'A warehouse robot folds a blue cloth on a clean workbench.' \ --warmup-mode request --warmup-steps 2 \ --warmup-resolutions 640x640 --warmup-num-frames 1 \ --save-output --perf-dump-path perf.json --output-path outputs --output-file-name sampleFor T2V, use its prompt,
--width 832 --height 480 --num-frames 81 --fps 24 --guidance-scale 5 --seed 42 --warmup-resolutions 832x480 --warmup-num-frames 81. Add--profile --num-profiled-timesteps 3only for separate diagnostic requests. Exact native runner, pinned preparation, audit and extraction scripts, SHA256 manifest.CI follow-up (2026-09-21)
Merged main
80da4432d085ed4d6166ef643d9fd2b829dbb0c5, including the upstream scheduler fixture repair #40411 and offline DFlash fixture repair #40427. Current head:25ff8a99305c14fce39b73f4b5219eb6ef07bc7a. Merged-head validation and all logs, source delta in this PR's files, pre-commit. The original E2E and output comparisons above remain measurements of their explicitly pinned benchmark heads; this follow-up does not claim new model-request timings. The common repaired CPU fixture suite passes 49 tests + 57 subtests.CI States
Latest PR Test (Base): ❌ Run #35525995597
Latest PR Test (Extra): ❌ Run #35525995362
Latest PR Test (AMD ROCm 10): ❌ Run #35525995657