Skip to content

[Model] Add MiniMax H3 attention strategy PoC - #8644

Draft
rahul-steiger-nv wants to merge 16 commits into
vllm-project:mainfrom
rahul-steiger-nv:experiment/minimax-h3-strategies-fa4
Draft

rahul-steiger-nv wants to merge 16 commits into
vllm-project:mainfrom
rahul-steiger-nv:experiment/minimax-h3-strategies-fa4

Conversation

@rahul-steiger-nv

@rahul-steiger-nv rahul-steiger-nv commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Extend the attention-strategy PoC from #8583 to MiniMax H3, following RFC #8382 and the SubBlock sparse RFC #8396. The recipe keeps the first twenty denoising iterations dense, then uses FA4 SubBlock attention in all 50 DiT layers. Token refinement stays dense.

Dependency: #8583 is the base for this stacked PR and builds on #8335. This change adds H3 integration, a recipe and tests; the shared strategy implementation comes from the base. The diff against main includes these prerequisites; review only the H3 changes.

Configuration, supported execution modes and test commands.

The recipe is illustrative. We do not claim an optimal schedule, a minimal sparse-layer count or general quality parity. Conditioning tokens remain eligible for pruning; always-sparse execution is experimental.

Test Plan

vLLM Version: 0.31.0
Validated commit: d975a3abb160857ad6ccaf80e984384f181fb7cb (base: #8583 at b9f81e58c). The single H3 commit was rebased onto the updated Cosmos3 stack on 2026-10-09; its foundation merges upstream main at 4c5541cfc.

GH200, Enroot ARM64 image, PyTorch 2.13.0+cu130, FA3 dense and FA4 4.0.0b33 sparse attention. Tests cover dispatch, packed padding, schedule cleanup, compilation/replay, component/layerwise offload and legacy regressions.

Benchmark: one prompt, seed 42, 1280×736, 192 frames at 24 fps, 50 denoising iterations, 75% target sparsity. Regional dynamic compilation, component CPU offload, VAE tiling and no Cache-DiT. One warmup and three measured requests per mode.

Test Result

After the 2026-10-09 stack update: 499 tests passed, 5 skipped across the H3 integration/shared-strategy suite (470 passed) and parallelism/VDN regressions (29 passed). Three skips require two Hopper GPUs; two reference removed upstream APIs. All applicable pre-commit hooks passed. The rebased H3 patch has only an upstream context adjustment. Earlier two-GPU and generation measurements below were not rerun for this update. These are local checks, not a claim that PR CI has passed.

Strict Ulysses is now implemented for H3 strategies in eager/regional execution,
reusing rank-local preparation and the shared attention wrapper. An earlier
focused regression run passed 303 distinct tests, with eight skips and all
applicable pre-commit checks passing. Real two-rank CPU/Gloo parity passed.
All three two-GPU NCCL/FA4 strategy cases passed on two NVIDIA H100 NVL GPUs:
resident, model-level CPU offload and ordinary layerwise CPU offload. Each case
checks eager and regional execution. The x86 v0.31.0rc1 run used vLLM 0.31.0,
PyTorch 2.13.0+cu130 and FA4 4.0.0b33, and completed in 191.57 seconds
(3m11s). Tested snapshot before squashing: 69b8c17a6.

Full-checkpoint distributed generation quality and latency have not been measured. Full-forward compiled Ulysses, Ring/hybrid
parallelism and distributed layerwise offload remain unsupported.

The following single-GPU validation and generation measurements were collected
at a753b8854, before the Ulysses addition:

458 distinct tests passed, six skipped. Four skips require two GPUs and two reference removed upstream APIs. All applicable pre-commit checks passed. Full CI has not run.

Mode Median generation latency Speedup
Dense 447.58 s 1.00x
20 dense steps, then sparse DiT 350.03 s 1.28x

Three measured requests per mode after one full warmup: dense 447.03/447.58/447.71 s; mixed 349.89/350.70/350.03 s. Timing includes text encoding, denoising and video/audio decode; excludes loading, warmup and MP4 encoding. Default exact AdaLN caching is enabled in both modes. Per-step dispatch was verified for all requests; measured requests did not recompile. No thermal throttling was recorded.

Comparison exports are 1280x720, 189 frames at 24 fps, center-cropped and trimmed from H3-native 1280x736, 192-frame generation. Both native clips are retained. Four sampled frames show coherent scenes, with differences in fabric folds and composition; this one prompt/seed does not establish general quality parity. Benchmark artifacts and comparison videos are retained locally and have not been attached to this draft.

AI assistance: Codex assisted with implementation, review, validation and this description.

Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Allow pure strict Ulysses through the shared Attention wrapper without
requiring an attention strategy. Prepare sparse adapters with post-all-to-all
head counts, retain full heads for replicated attention, and reject unsupported
parallel combinations and inactive sharded execution.

Let Cosmos3 remove synthetic suffix padding before sparse selection and
restore zero rows before reverse communication. Preserve the protected
understanding KV prefix and the ordinary dense mask path. Compose capability
queries from actual local tensors without launching kernels or collectives.

Validate static per-role sparse configuration against single-device outputs
using real NCCL and FA4, padded and unpadded sequences, repeated requests,
regional compilation, and model/layerwise CPU offloading. Sequence lengths
exceed the selector's minimum retained budget to exercise sparse selection.

Validation:
- Two Hopper GPUs, FA4 4.0.0b33: 4 passed (156.11s).
- Focused sparse attention, SP-hook, capability, adapter, parallel admission,
  and Cosmos3 regression tests passed after extraction fixes.
- Changed-file pre-commit and git diff --check passed.

Environment-specific validation launchers and logs remain local.

Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
@rahul-steiger-nv

Copy link
Copy Markdown
Contributor Author

Cosmos3-Nano generated outputs: dense vs. mixed sparse vs. all-sparse, using the same prompt and seed. Labels show generation speedup, not playback speed.

comparison-three-way.mp4

@hsliuustc0106 hsliuustc0106 added the enhancement New feature or request label Oct 9, 2026
@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Oct 9, 2026 — with ChatGPT Codex Connector
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
@rahul-steiger-nv
rahul-steiger-nv force-pushed the experiment/minimax-h3-strategies-fa4 branch from 6819aa3 to d975a3a Compare October 9, 2026 11:33
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
@rahul-steiger-nv
rahul-steiger-nv force-pushed the experiment/minimax-h3-strategies-fa4 branch from d975a3a to 4aeb2dd Compare October 9, 2026 12:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request high priority high priority issue, needs to be done asap

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants