Repository navigation
[Core] Add shared block-sparse attention adapters with FlashAttention-4 - #8335
rahul-steiger-nv wants to merge 10 commits into
Conversation
b076b99 to
56a7da3
Compare
|
This PR touches vllm_omni/diffusion/, tests/diffusion/, docs/design/, recipes/attention/, .gitignore (40 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be: Could one of you take a look when you get a chance? Thanks! |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Two issues need fixing before merge; details are in the inline comments.
Reviewed commit 56a7da3. Static Python syntax, recipe JSON, and diff checks passed. Runtime/GPU tests were not run because no disposable sandbox was available for the fork.
a041037 to
10d130a
Compare
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
676558f to
be9ab0e
Compare
|
This PR appears to belong to: docs/design/module/diffusion/index.md. Module owners: @david6666666 @wtomin @xuechendi Routing: @david6666666 via module of the changed files, module named in the PR description, semantic router, model owner; @wtomin via module of the changed files, module named in the PR description, semantic router, CODEOWNERS; @xuechendi via module of the changed files, module named in the PR description, semantic router, CODEOWNERS @rahul-steiger-nv, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot routing recordAssigned Strict on zcode (GLM-5.3-Flash) under experiment |
Allow pure strict Ulysses through the shared Attention wrapper without requiring an attention strategy. Prepare sparse adapters with post-all-to-all head counts, retain full heads for replicated attention, and reject unsupported parallel combinations and inactive sharded execution. Let Cosmos3 remove synthetic suffix padding before sparse selection and restore zero rows before reverse communication. Preserve the protected understanding KV prefix and the ordinary dense mask path. Compose capability queries from actual local tensors without launching kernels or collectives. Validate static per-role sparse configuration against single-device outputs using real NCCL and FA4, padded and unpadded sequences, repeated requests, regional compilation, and model/layerwise CPU offloading. Sequence lengths exceed the selector's minimum retained budget to exercise sparse selection. Validation: - Two Hopper GPUs, FA4 4.0.0b33: 4 passed (156.11s). - Focused sparse attention, SP-hook, capability, adapter, parallel admission, and Cosmos3 regression tests passed after extraction fixes. - Changed-file pre-commit and git diff --check passed. Environment-specific validation launchers and logs remain local. Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
8c21c01 to
b95ff49
Compare
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Omni ReviewBot triage noteResolved as of |
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Purpose
Add shared selected-block attention and a FlashAttention-4 execution adapter. Model roles choose the method through existing attention configuration; the selector produces per-query-head block indices/counts, and backend-owned adapters execute that selection.
The shared layer owns protected KV prefixes, typed single-document suffix padding, actual-request preparation, and the opaque compilation boundary. FA4 preserves native MHA/MQA/GQA head mapping and block geometry. Provider failures propagate without dense fallback or K/V expansion. Provider selection remains platform-owned through
OmniPlatform.resolve_diffusion_attn_backend(), with sparse compatibility checked independently of dense-kernel restrictions.FA4 is pinned to
4.0.0b33. The example recipes use Hopper 64×64 blocks; local Blackwell validation uses explicit 256×128 blocks. The 75% target sparsity is experimental, not a quality-qualified default.Merged upstream
mainat4c5541cfc; current PR head:5e2569678. Builds on #7379 and execution-contract RFC #7226, implementing the foundation for RFC #8396.Design, configuration composition, support matrix, and limitations.
Supported scope
The design proposes composition with #8382: one schedule chooses complete-backend or selected-block method specifications, reusing platform/plugin dispatch and preserving metadata capabilities for all reachable methods.
AttentionSpec.sparseis not implemented here; the shared public schema still requires coordination. Layer/step scheduling is separate, demonstrated for Cosmos3 in #8583.Provider follow-ups: FlashInfer #8336, TRTLLM #8337, and cuDNN #8338.
Validation
After the main merge (2026-10-09): one GH200, vLLM-Omni
0.31.0rc1ARM64 image, FA44.0.0b33.Reproduce the post-merge checks:
VLLM_FLASH_ATTN_VERSION=4 python -m pytest -q -ra \ tests/diffusion/attention/test_block_selection.py \ tests/diffusion/attention/test_block_sparse*.py \ tests/diffusion/attention/test_attention_config.py \ tests/diffusion/attention/test_selector.py \ tests/diffusion/attention/test_attention_capabilities.py \ tests/diffusion/distributed/test_sp_plan_hooks.py \ tests/diffusion/distributed/test_cosmos3_pre_sharded.py \ tests/diffusion/models/cosmos3/test_cosmos3_transformer.py \ tests/diffusion/models/cosmos3/test_cosmos3_sparse_ulysses.py \ tests/diffusion/models/minimax_h3/test_minimax_h3_contract.py \ tests/diffusion/models/minimax_h3/test_minimax_h3_sparse_ulysses.pyEarlier evidence, not rerun after this merge:
be9ab0e84. These results used the 0.30.0 image and do not qualify Blackwell distributed/offload combinations.Two-rank validation can be rerun with the two
test_*_sparse_ulysses.pymodules above andCUDA_VISIBLE_DEVICES=0,1; the tests spawn their own ranks.Quality and performance limits
Full-checkpoint quality remains unqualified for the static recipes. #8583 reports preliminary Cosmos3 generation latency with the pinned wheel for one prompt/seed, not general quality preservation or distributed throughput.
Upstream MiniMax evaluation motivates ten initial dense steps followed by SubBlock. It is evidence for an upstream speed/similarity trade-off, not validation of this port's always-sparse recipes. Matched local MiniMax dense/mixed quality and latency evaluation remains follow-up work. Conditioning tokens are currently eligible for pruning; other models and block geometries need separate qualification.
AI assistance: Codex helped implement code, tests and documentation, run validation, and prepare this PR.