Repository navigation
[Core] Add Cosmos3 attention strategy PoC - #8583
Draft
rahul-steiger-nv wants to merge 14 commits into
Draft
rahul-steiger-nv wants to merge 14 commits into
rahul-steiger-nv wants to merge 14 commits into
Conversation
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Allow pure strict Ulysses through the shared Attention wrapper without requiring an attention strategy. Prepare sparse adapters with post-all-to-all head counts, retain full heads for replicated attention, and reject unsupported parallel combinations and inactive sharded execution. Let Cosmos3 remove synthetic suffix padding before sparse selection and restore zero rows before reverse communication. Preserve the protected understanding KV prefix and the ordinary dense mask path. Compose capability queries from actual local tensors without launching kernels or collectives. Validate static per-role sparse configuration against single-device outputs using real NCCL and FA4, padded and unpadded sequences, repeated requests, regional compilation, and model/layerwise CPU offloading. Sequence lengths exceed the selector's minimum retained budget to exercise sparse selection. Validation: - Two Hopper GPUs, FA4 4.0.0b33: 4 passed (156.11s). - Focused sparse attention, SP-hook, capability, adapter, parallel admission, and Cosmos3 regression tests passed after extraction fixes. - Changed-file pre-commit and git diff --check passed. Environment-specific validation launchers and logs remain local. Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
rahul-steiger-nv
force-pushed
the
experiment/cosmos3-strategies-fa4
branch
from
October 7, 2026 08:42
a852599 to
5363636
Compare
Contributor
Author
|
Cosmos3-Nano generated outputs: dense vs. mixed sparse vs. all-sparse, using the same prompt and seed. Labels show generation speedup, not playback speed. comparison-captioned.mp4 |
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
This was referenced Oct 8, 2026
rahul-steiger-nv
force-pushed
the
experiment/cosmos3-strategies-fa4
branch
from
October 8, 2026 10:37
5363636 to
a16c015
Compare
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
rahul-steiger-nv
force-pushed
the
experiment/cosmos3-strategies-fa4
branch
from
October 9, 2026 11:33
852f74f to
b9f81e5
Compare
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
This is the Cosmos3 PoC I offered in @nodeeeeee's RFC #8382. It switches between dense and sparse attention by denoising step and layer, with layout selection outside compiled execution and shared model weights.
It adds attention presets, layouts, schedules, checkpoint defaults with deployment overrides, and warmup across layouts. Recipes demonstrate a uniform dense-to-sparse switch and a mixed-layer policy. See design, usage, and scope.
Dependency: #8335 supplies the shared block-sparse foundation and FA4 adapter. The diff against
mainincludes prerequisite changes; review only the strategy PoC.Test Plan
vLLM Version:
0.31.0vLLM-Omni Commit:
b9f81e58cdbc5f05a6497ffd90960867f8b36acf(rebased onto the updated #8335 foundation at5e2569678, which merges upstreammainat4c5541cfc)Environment: GH200, Enroot ARM64 image, PyTorch
2.13.0+cu130, FA3 dense and FA44.0.0b33sparse attention.python3 -m pytest -q \ tests/diffusion/attention/test_attention_strategy*.py \ tests/diffusion/attention/test_attention_checkpoint_policy.py \ tests/diffusion/attention/test_cosmos3_action_strategy.py \ tests/diffusion/attention/test_attention_config.py \ tests/diffusion/models/cosmos3/test_cosmos3_attention_strategy.py \ tests/diffusion/models/cosmos3/test_cosmos3_transformer.py \ tests/diffusion/test_diffusion_engine_dummy_run.py \ tests/diffusion/test_diffusion_model_runner.py \ tests/diffusion/models/cosmos3/test_cosmos3_strategy_offload.py \ tests/diffusion/models/cosmos3/test_cosmos3_strategy_ulysses.py \ tests/diffusion/distributed/test_sp_plan_hooks.py \ tests/diffusion/distributed/test_cosmos3_pre_sharded.py \ tests/diffusion/diffusion_kv/test_paged_attention_adapter.py \ tests/diffusion/models/wan2_2/test_wan22_quant_config_propagation.pyTest Result
516 tests passed, 5 skipped after the 2026-10-09 rebase. Only the three Cosmos3 strategy commits were replayed onto the updated foundation; their patches are unchanged. Coverage includes strategy/configuration behavior, CUDA compilation and reuse, Cosmos3 regressions, single-GPU CPU offload, CPU/Gloo Ulysses, and upstream sharding/paged-KV integration. The rebase preserves upstream Cosmos3 sharding before the generation stack and rank-local embedding ownership.
The five skips require two GPUs. Earlier two-Hopper strategy validation covered strict Ulysses with no/model/layerwise CPU offload in eager and regional modes; those CUDA cases have not been rerun after this rebase. Applicable pre-commit hooks passed except
mypy-3.10: all 43 errors exactly match the foundation, with no new diagnostics. Full CI has not run for this PR.Historical generation measurements: taken before this rebase at
a8525998dwith vLLM0.30.0; not rerun on the rebased branch.Cosmos3-Nano: default prompt, seed 42, 720p, 189 frames, 35 steps, mixed precision disabled, and 75% target sparsity. One full warmup and one measured generation per mode.
Expected sparse dispatch was verified at every step, with no measured-run recompilation or thermal throttling. Latency includes decode; it excludes loading, warmup, and MP4 encoding. The benchmark replaces synthetic startup warmup with a full generation.
These settings are illustrative: no optimal strategy or minimal sparse-layer count is claimed. One prompt/seed and one measurement per mode demonstrate preliminary speedup, not general quality preservation or completion of the RFC's validation requirements.
AI assistance: Codex assisted with code review, validation, benchmarking, and preparation of this PR description.