Skip to content

[Core] Add Cosmos3 attention strategy PoC - #8583

Draft
rahul-steiger-nv wants to merge 14 commits into
vllm-project:mainfrom
rahul-steiger-nv:experiment/cosmos3-strategies-fa4
Draft

rahul-steiger-nv wants to merge 14 commits into
vllm-project:mainfrom
rahul-steiger-nv:experiment/cosmos3-strategies-fa4

Conversation

@rahul-steiger-nv

@rahul-steiger-nv rahul-steiger-nv commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

This is the Cosmos3 PoC I offered in @nodeeeeee's RFC #8382. It switches between dense and sparse attention by denoising step and layer, with layout selection outside compiled execution and shared model weights.

It adds attention presets, layouts, schedules, checkpoint defaults with deployment overrides, and warmup across layouts. Recipes demonstrate a uniform dense-to-sparse switch and a mixed-layer policy. See design, usage, and scope.

Dependency: #8335 supplies the shared block-sparse foundation and FA4 adapter. The diff against main includes prerequisite changes; review only the strategy PoC.

Test Plan

vLLM Version: 0.31.0
vLLM-Omni Commit: b9f81e58cdbc5f05a6497ffd90960867f8b36acf (rebased onto the updated #8335 foundation at 5e2569678, which merges upstream main at 4c5541cfc)

Environment: GH200, Enroot ARM64 image, PyTorch 2.13.0+cu130, FA3 dense and FA4 4.0.0b33 sparse attention.

python3 -m pytest -q \
  tests/diffusion/attention/test_attention_strategy*.py \
  tests/diffusion/attention/test_attention_checkpoint_policy.py \
  tests/diffusion/attention/test_cosmos3_action_strategy.py \
  tests/diffusion/attention/test_attention_config.py \
  tests/diffusion/models/cosmos3/test_cosmos3_attention_strategy.py \
  tests/diffusion/models/cosmos3/test_cosmos3_transformer.py \
  tests/diffusion/test_diffusion_engine_dummy_run.py \
  tests/diffusion/test_diffusion_model_runner.py \
  tests/diffusion/models/cosmos3/test_cosmos3_strategy_offload.py \
  tests/diffusion/models/cosmos3/test_cosmos3_strategy_ulysses.py \
  tests/diffusion/distributed/test_sp_plan_hooks.py \
  tests/diffusion/distributed/test_cosmos3_pre_sharded.py \
  tests/diffusion/diffusion_kv/test_paged_attention_adapter.py \
  tests/diffusion/models/wan2_2/test_wan22_quant_config_propagation.py

Test Result

516 tests passed, 5 skipped after the 2026-10-09 rebase. Only the three Cosmos3 strategy commits were replayed onto the updated foundation; their patches are unchanged. Coverage includes strategy/configuration behavior, CUDA compilation and reuse, Cosmos3 regressions, single-GPU CPU offload, CPU/Gloo Ulysses, and upstream sharding/paged-KV integration. The rebase preserves upstream Cosmos3 sharding before the generation stack and rank-local embedding ownership.

The five skips require two GPUs. Earlier two-Hopper strategy validation covered strict Ulysses with no/model/layerwise CPU offload in eager and regional modes; those CUDA cases have not been rerun after this rebase. Applicable pre-commit hooks passed except mypy-3.10: all 43 errors exactly match the foundation, with no new diagnostics. Full CI has not run for this PR.

Historical generation measurements: taken before this rebase at a8525998d with vLLM 0.30.0; not rerun on the rebased branch.

Cosmos3-Nano: default prompt, seed 42, 720p, 189 frames, 35 steps, mixed precision disabled, and 75% target sparsity. One full warmup and one measured generation per mode.

Generation attention Generation latency Speedup
Dense 194.66 s 1.00×
Mixed: 10 dense steps, then 28/36 sparse layers 159.42 s 1.22×
All 36 generation layers sparse throughout 131.21 s 1.48×

Expected sparse dispatch was verified at every step, with no measured-run recompilation or thermal throttling. Latency includes decode; it excludes loading, warmup, and MP4 encoding. The benchmark replaces synthetic startup warmup with a full generation.

These settings are illustrative: no optimal strategy or minimal sparse-layer count is claimed. One prompt/seed and one measurement per mode demonstrate preliminary speedup, not general quality preservation or completion of the RFC's validation requirements.

AI assistance: Codex assisted with code review, validation, benchmarking, and preparation of this PR description.

Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Allow pure strict Ulysses through the shared Attention wrapper without
requiring an attention strategy. Prepare sparse adapters with post-all-to-all
head counts, retain full heads for replicated attention, and reject unsupported
parallel combinations and inactive sharded execution.

Let Cosmos3 remove synthetic suffix padding before sparse selection and
restore zero rows before reverse communication. Preserve the protected
understanding KV prefix and the ordinary dense mask path. Compose capability
queries from actual local tensors without launching kernels or collectives.

Validate static per-role sparse configuration against single-device outputs
using real NCCL and FA4, padded and unpadded sequences, repeated requests,
regional compilation, and model/layerwise CPU offloading. Sequence lengths
exceed the selector's minimum retained budget to exercise sparse selection.

Validation:
- Two Hopper GPUs, FA4 4.0.0b33: 4 passed (156.11s).
- Focused sparse attention, SP-hook, capability, adapter, parallel admission,
  and Cosmos3 regression tests passed after extraction fixes.
- Changed-file pre-commit and git diff --check passed.

Environment-specific validation launchers and logs remain local.

Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
@rahul-steiger-nv
rahul-steiger-nv force-pushed the experiment/cosmos3-strategies-fa4 branch from a852599 to 5363636 Compare October 7, 2026 08:42
@rahul-steiger-nv

Copy link
Copy Markdown
Contributor Author

Cosmos3-Nano generated outputs: dense vs. mixed sparse vs. all-sparse, using the same prompt and seed. Labels show generation speedup, not playback speed.

comparison-captioned.mp4

@hsliuustc0106 hsliuustc0106 added the core related to core module: cache, scheduler, engine, worker, modelrunner label Oct 8, 2026
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
@rahul-steiger-nv
rahul-steiger-nv force-pushed the experiment/cosmos3-strategies-fa4 branch from 852f74f to b9f81e5 Compare October 9, 2026 11:33
Signed-off-by: Rahul Steiger <rsteiger@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants