Repository navigation
[Model] Add MiniMax H3 attention strategy PoC - #8644
Draft
rahul-steiger-nv wants to merge 16 commits into
Draft
rahul-steiger-nv wants to merge 16 commits into
rahul-steiger-nv wants to merge 16 commits into
Conversation
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Allow pure strict Ulysses through the shared Attention wrapper without requiring an attention strategy. Prepare sparse adapters with post-all-to-all head counts, retain full heads for replicated attention, and reject unsupported parallel combinations and inactive sharded execution. Let Cosmos3 remove synthetic suffix padding before sparse selection and restore zero rows before reverse communication. Preserve the protected understanding KV prefix and the ordinary dense mask path. Compose capability queries from actual local tensors without launching kernels or collectives. Validate static per-role sparse configuration against single-device outputs using real NCCL and FA4, padded and unpadded sequences, repeated requests, regional compilation, and model/layerwise CPU offloading. Sequence lengths exceed the selector's minimum retained budget to exercise sparse selection. Validation: - Two Hopper GPUs, FA4 4.0.0b33: 4 passed (156.11s). - Focused sparse attention, SP-hook, capability, adapter, parallel admission, and Cosmos3 regression tests passed after extraction fixes. - Changed-file pre-commit and git diff --check passed. Environment-specific validation launchers and logs remain local. Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
rahul-steiger-nv
force-pushed
the
experiment/minimax-h3-strategies-fa4
branch
from
October 8, 2026 16:31
4a3abde to
6819aa3
Compare
1 task done
Contributor
Author
|
Cosmos3-Nano generated outputs: dense vs. mixed sparse vs. all-sparse, using the same prompt and seed. Labels show generation speedup, not playback speed. comparison-three-way.mp4 |
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
rahul-steiger-nv
force-pushed
the
experiment/minimax-h3-strategies-fa4
branch
from
October 9, 2026 11:33
6819aa3 to
d975a3a
Compare
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
rahul-steiger-nv
force-pushed
the
experiment/minimax-h3-strategies-fa4
branch
from
October 9, 2026 12:55
d975a3a to
4aeb2dd
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Extend the attention-strategy PoC from #8583 to MiniMax H3, following RFC #8382 and the SubBlock sparse RFC #8396. The recipe keeps the first twenty denoising iterations dense, then uses FA4 SubBlock attention in all 50 DiT layers. Token refinement stays dense.
Dependency: #8583 is the base for this stacked PR and builds on #8335. This change adds H3 integration, a recipe and tests; the shared strategy implementation comes from the base. The diff against
mainincludes these prerequisites; review only the H3 changes.Configuration, supported execution modes and test commands.
The recipe is illustrative. We do not claim an optimal schedule, a minimal sparse-layer count or general quality parity. Conditioning tokens remain eligible for pruning; always-sparse execution is experimental.
Test Plan
vLLM Version:
0.31.0Validated commit:
d975a3abb160857ad6ccaf80e984384f181fb7cb(base: #8583 atb9f81e58c). The single H3 commit was rebased onto the updated Cosmos3 stack on 2026-10-09; its foundation merges upstreammainat4c5541cfc.GH200, Enroot ARM64 image, PyTorch
2.13.0+cu130, FA3 dense and FA44.0.0b33sparse attention. Tests cover dispatch, packed padding, schedule cleanup, compilation/replay, component/layerwise offload and legacy regressions.Benchmark: one prompt, seed 42, 1280×736, 192 frames at 24 fps, 50 denoising iterations, 75% target sparsity. Regional dynamic compilation, component CPU offload, VAE tiling and no Cache-DiT. One warmup and three measured requests per mode.
Test Result
After the 2026-10-09 stack update: 499 tests passed, 5 skipped across the H3 integration/shared-strategy suite (470 passed) and parallelism/VDN regressions (29 passed). Three skips require two Hopper GPUs; two reference removed upstream APIs. All applicable pre-commit hooks passed. The rebased H3 patch has only an upstream context adjustment. Earlier two-GPU and generation measurements below were not rerun for this update. These are local checks, not a claim that PR CI has passed.
Strict Ulysses is now implemented for H3 strategies in eager/regional execution,
reusing rank-local preparation and the shared attention wrapper. An earlier
focused regression run passed 303 distinct tests, with eight skips and all
applicable pre-commit checks passing. Real two-rank CPU/Gloo parity passed.
All three two-GPU NCCL/FA4 strategy cases passed on two NVIDIA H100 NVL GPUs:
resident, model-level CPU offload and ordinary layerwise CPU offload. Each case
checks eager and regional execution. The x86 v0.31.0rc1 run used vLLM
0.31.0,PyTorch
2.13.0+cu130and FA44.0.0b33, and completed in 191.57 seconds(3m11s). Tested snapshot before squashing:
69b8c17a6.Full-checkpoint distributed generation quality and latency have not been measured. Full-forward compiled Ulysses, Ring/hybrid
parallelism and distributed layerwise offload remain unsupported.
The following single-GPU validation and generation measurements were collected
at
a753b8854, before the Ulysses addition:458 distinct tests passed, six skipped. Four skips require two GPUs and two reference removed upstream APIs. All applicable pre-commit checks passed. Full CI has not run.
Three measured requests per mode after one full warmup: dense 447.03/447.58/447.71 s; mixed 349.89/350.70/350.03 s. Timing includes text encoding, denoising and video/audio decode; excludes loading, warmup and MP4 encoding. Default exact AdaLN caching is enabled in both modes. Per-step dispatch was verified for all requests; measured requests did not recompile. No thermal throttling was recorded.
Comparison exports are 1280x720, 189 frames at 24 fps, center-cropped and trimmed from H3-native 1280x736, 192-frame generation. Both native clips are retained. Four sampled frames show coherent scenes, with differences in fabric folds and composition; this one prompt/seed does not establish general quality parity. Benchmark artifacts and comparison videos are retained locally and have not been attached to this draft.
AI assistance: Codex assisted with implementation, review, validation and this description.