Repository navigation
[diffusion] Flatten Wan VAE RMSNorm row addressing - #35981
Merged
Merged
Conversation
BBuf
force-pushed
the
bbuf/wan-rmsnorm-silu-flat-rows
branch
from
August 22, 2026 19:17
f54f854 to
e0fe07e
Compare
BBuf
marked this pull request as ready for review
August 23, 2026 12:42
Collaborator
Author
|
/tag-and-rerun-ci extra |
BBuf
force-pushed
the
bbuf/wan-rmsnorm-silu-flat-rows
branch
from
August 23, 2026 14:00
e0fe07e to
54c752e
Compare
Collaborator
Author
|
/rerun-failed-ci |
Collaborator
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
[pixel, channel]rowsb/t/h/wdiv/mod chain and general stride reconstructionWhy
Nsight Compute on the representative FastWan2.1 shape (
[1, 96, 4, 480, 832], BF16 input / FP32 affine) shows 80.58% SM throughput but only 9.14% DRAM throughput and 93.93% achieved occupancy. Source counters attribute 120,771 / 196,793 samples to integer/index instructions. The hottest PCs are the reciprocal-based lowering of the row-to-b/t/h/wdiv/mod chain (I2F.RP,IABS, andMUFU.RCP).Dense channels-last-3d already stores every pixel as a contiguous channel row, so those coordinates are unnecessary.
stride(C) == 1is checked explicitly because size-one channel tensors can satisfy the memory-format predicate with channel-first strides.Performance
Single NVIDIA H200, cold-L2 registered benchmark:
[1,96,4,480,832][1,256,4,384,576]FastWan2.1 end-to-end ABBA, one H200, eager, 832x480x61, 3 steps, no compile, resident DiT:
The quality-high full-stage trace reduces all 415 Wan RMSNorm+SiLU launches from 215.234 ms to 127.128 ms (40.94%). The largest C96 shape falls from 171.438 ms to 101.697 ms over 105 launches. Peak reserved memory stays at 31.402 GiB.
TurboWan2.1 T2V 1.3B provides a second model-level check at 832x480x81, 4 steps, quality=high, and no compile. Its main trace spends 286.083 ms across 545 Wan RMSNorm+SiLU launches (11.84% of GPU time). A clean main/PR/PR/main ABBA measured:
Denoise is 0.57% slower, confirming that the request gain comes from the VAE path rather than timing drift in the DiT. Peak reserved memory remains 30.729 GiB.
The 14B TurboWan checkpoints provide resolution scaling checks with the same 545 launches:
Both ABBA brackets are byte-exact across main and PR. These longer DiT-heavy requests dilute the VAE gain, but independently confirm the kernel reduction at both production resolutions.
TurboWan2.2 I2V A14B adds a four-H200 CFG/Ulysses check at 1280x720x81. Main averaged 7.6857 / 12.1847 s and the PR averaged 7.6780 / 12.2241 s. Rank-0 recorded GPU time fell from 12,020.198 to 11,866.292 ms (-1.28%), but synchronized e2e regressed 0.32%, so this checkpoint is validation evidence rather than a claimed request speedup. All five outputs are byte-exact.
Output comparison
Prompt:
A curious raccoon walks through a sunlit forest.Seed:42.main video · PR video · side-by-side video
All four FastWan quality-high outputs have the same SHA256:
fcd22372b15c84dca9b7e18848a4642103bbe6f52dc7fbb3314e9d3e81bceb8a. All four TurboWan ABBA outputs have the same SHA256:4eaef46a23e5e8c1e5a7a16941ba4c8b9080e2620ec2551848a11e0060270e46.Validation
pytest -q test/registered/kernels/ops/diffusion/test_norm.py -k wan_rmsnorm_silu: 4 passedpytest -q test/registered/kernels/ops/diffusion/test_model_fast_paths.py -k wan_vae_gate: 1 passedCI States
Latest PR Test (Base): ✅ Run #32644106273
Latest PR Test (Extra): ✅ Run #32661657125
Latest PR Test (AMD ROCm 7.2): ⏳ Run #32644106161