Skip to content

[Core] Add shared block-sparse attention adapters with FlashAttention-4 - #8335

Open
rahul-steiger-nv wants to merge 10 commits into
vllm-project:mainfrom
rahul-steiger-nv:rsteiger/block-sparse-foundation
Open

rahul-steiger-nv wants to merge 10 commits into
vllm-project:mainfrom
rahul-steiger-nv:rsteiger/block-sparse-foundation

Conversation

@rahul-steiger-nv

@rahul-steiger-nv rahul-steiger-nv commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Add shared selected-block attention and a FlashAttention-4 execution adapter. Model roles choose the method through existing attention configuration; the selector produces per-query-head block indices/counts, and backend-owned adapters execute that selection.

The shared layer owns protected KV prefixes, typed single-document suffix padding, actual-request preparation, and the opaque compilation boundary. FA4 preserves native MHA/MQA/GQA head mapping and block geometry. Provider failures propagate without dense fallback or K/V expansion. Provider selection remains platform-owned through OmniPlatform.resolve_diffusion_attn_backend(), with sparse compatibility checked independently of dense-kernel restrictions.

FA4 is pinned to 4.0.0b33. The example recipes use Hopper 64×64 blocks; local Blackwell validation uses explicit 256×128 blocks. The 75% target sparsity is experimental, not a quality-qualified default.

Merged upstream main at 4c5541cfc; current PR head: 5e2569678. Builds on #7379 and execution-contract RFC #7226, implementing the foundation for RFC #8396.

Design, configuration composition, support matrix, and limitations.

Supported scope

  • Cosmos3 and MiniMax H3: single-device execution and pure strict Ulysses, with Q and KV head counts divisible by the Ulysses degree.
  • Model and ordinary layerwise CPU offloading: single-device or strict Ulysses, using eager or regional compilation. MiniMax layerwise offload streams DiT blocks while keeping the token refiner resident.
  • Local sparse execution: eager and dynamic/fullgraph compilation. This does not establish fullgraph capture of Ulysses communication or full-model compilation with offload.
  • Outside scope: distributed layerwise offloading, Ring, AllGather-KV, advanced Ulysses, HSDP, hybrid parallelism, arbitrary/causal masks, multi-document packing, and paged/quantized KV. Cosmos3 TP evidence covers rank-local projections and attention, not full TP collectives.

The design proposes composition with #8382: one schedule chooses complete-backend or selected-block method specifications, reusing platform/plugin dispatch and preserving metadata capabilities for all reachable methods. AttentionSpec.sparse is not implemented here; the shared public schema still requires coordination. Layer/step scheduling is separate, demonstrated for Cosmos3 in #8583.

Provider follow-ups: FlashInfer #8336, TRTLLM #8337, and cuDNN #8338.

Validation

After the main merge (2026-10-09): one GH200, vLLM-Omni 0.31.0rc1 ARM64 image, FA4 4.0.0b33.

  • 761 tests passed; 10 skipped. Eight skips require two GPUs; two concern unavailable MiniMax APIs.
  • Covers selection/adapter correctness, native head mapping, configuration and dispatch, capability queries, sequence-parallel hooks, compilation, Cosmos3/MiniMax integration, and single-GPU model/layerwise offloading.
  • Applicable pre-commit checks passed, including mypy. This is local validation, not a claim that PR CI has passed.

Reproduce the post-merge checks:

VLLM_FLASH_ATTN_VERSION=4 python -m pytest -q -ra \
  tests/diffusion/attention/test_block_selection.py \
  tests/diffusion/attention/test_block_sparse*.py \
  tests/diffusion/attention/test_attention_config.py \
  tests/diffusion/attention/test_selector.py \
  tests/diffusion/attention/test_attention_capabilities.py \
  tests/diffusion/distributed/test_sp_plan_hooks.py \
  tests/diffusion/distributed/test_cosmos3_pre_sharded.py \
  tests/diffusion/models/cosmos3/test_cosmos3_transformer.py \
  tests/diffusion/models/cosmos3/test_cosmos3_sparse_ulysses.py \
  tests/diffusion/models/minimax_h3/test_minimax_h3_contract.py \
  tests/diffusion/models/minimax_h3/test_minimax_h3_sparse_ulysses.py

Earlier evidence, not rerun after this merge:

  • Hopper/FA4 Ulysses and offloading: all six Cosmos3 cases and five MiniMax cases passed, including two-rank execution, padding, actual sparse dispatch, and eager/regional compilation. These use small randomly initialized models.
  • Prior foundation adapter regressions: GH200 827 passed / 12 skipped; B200 764 passed / 75 skipped at be9ab0e84. These results used the 0.30.0 image and do not qualify Blackwell distributed/offload combinations.

Two-rank validation can be rerun with the two test_*_sparse_ulysses.py modules above and CUDA_VISIBLE_DEVICES=0,1; the tests spawn their own ranks.

Quality and performance limits

Full-checkpoint quality remains unqualified for the static recipes. #8583 reports preliminary Cosmos3 generation latency with the pinned wheel for one prompt/seed, not general quality preservation or distributed throughput.

Upstream MiniMax evaluation motivates ten initial dense steps followed by SubBlock. It is evidence for an upstream speed/similarity trade-off, not validation of this port's always-sparse recipes. Matched local MiniMax dense/mixed quality and latency evaluation remains follow-up work. Conditioning tokens are currently eligible for pruning; other models and block geometries need separate qualification.

AI assistance: Codex helped implement code, tests and documentation, run validation, and prepare this PR.

@rahul-steiger-nv
rahul-steiger-nv force-pushed the rsteiger/block-sparse-foundation branch 3 times, most recently from b076b99 to 56a7da3 Compare September 30, 2026 17:01
@hsliuustc0106 hsliuustc0106 added core related to core module: cache, scheduler, engine, worker, modelrunner diffusion codes related to diffusion models labels Sep 30, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

This PR touches vllm_omni/diffusion/, tests/diffusion/, docs/design/, recipes/attention/, .gitignore (40 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be:

@wtomin @david6666666 @yenuo26

Could one of you take a look when you get a chance? Thanks!

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Oct 1, 2026 — with ChatGPT Codex Connector

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two issues need fixing before merge; details are in the inline comments.

Reviewed commit 56a7da3. Static Python syntax, recipe JSON, and diff checks passed. Runtime/GPU tests were not run because no disposable sandbox was available for the fork.

Comment thread vllm_omni/diffusion/attention/backends/flash_attn.py
Comment thread vllm_omni/diffusion/attention/block_sparse.py
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/index.md.

Module owners: @david6666666 @wtomin @xuechendi

Routing: @david6666666 via module of the changed files, module named in the PR description, semantic router, model owner; @wtomin via module of the changed files, module named in the PR description, semantic router, CODEOWNERS; @xuechendi via module of the changed files, module named in the PR description, semantic router, CODEOWNERS

@rahul-steiger-nv, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

Allow pure strict Ulysses through the shared Attention wrapper without
requiring an attention strategy. Prepare sparse adapters with post-all-to-all
head counts, retain full heads for replicated attention, and reject unsupported
parallel combinations and inactive sharded execution.

Let Cosmos3 remove synthetic suffix padding before sparse selection and
restore zero rows before reverse communication. Preserve the protected
understanding KV prefix and the ordinary dense mask path. Compose capability
queries from actual local tensors without launching kernels or collectives.

Validate static per-role sparse configuration against single-device outputs
using real NCCL and FA4, padded and unpadded sequences, repeated requests,
regional compilation, and model/layerwise CPU offloading. Sequence lengths
exceed the selector's minimum retained budget to exercise sparse selection.

Validation:
- Two Hopper GPUs, FA4 4.0.0b33: 4 passed (156.11s).
- Focused sparse attention, SP-hook, capability, adapter, parallel admission,
  and Cosmos3 regression tests passed after extraction fixes.
- Changed-file pre-commit and git diff --check passed.

Environment-specific validation launchers and logs remain local.

Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
@vllm-omni-review-bot

vllm-omni-review-bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Resolved as of 5e2569678108: the high-priority or low-quality signal noted on an earlier commit no longer applies.

Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner diffusion codes related to diffusion models high priority high priority issue, needs to be done asap

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants