Skip to content

[Attention][DSA] Enable W4A16 DSA - #51724

Merged
DarkLight1337 merged 1 commit into
vllm-project:mainfrom
sychen52:W4A8R8_DSA
Sep 1, 2026
Merged

DarkLight1337 merged 1 commit into
vllm-project:mainfrom
sychen52:W4A8R8_DSA

Conversation

@sychen52

@sychen52 sychen52 commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Corresponding FlashMLA PR: vllm-project/FlashMLA#18

  • concat_and_cache_nvfp4_ds_mla — quantize and store

    • Write path for --kv-cache-dtype nvfp4_fp8_ds_mla, reached through the existing concat_and_cache_mla op.
    • Quantizes the bf16 NoPE latent to e2m1 with one e4m3 scale per 16 elements, and the RoPE to unscaled e4m3.
    • Writes a 352 B entry per token into the paged cache.
  • cp_gather_and_upconvert_nvfp4_kv_cache — dequantize for prefill

    • Read path, mirroring cp_gather_and_upconvert_fp8_kv_cache.

    • Gathers a batch's scattered cache pages and upconverts them to a contiguous bf16 [total_tokens, 576] workspace for the bf16 prefill kernel.

    • Only runs at ≥32 query heads per rank, so at TP8 neither DSv3.2 nor GLM-5.2 reaches it — covered by unit tests, not their E2E runs.

Purpose

Enable W4A16 DSA

This PR is now in draft for testing purpose, the pin of FlashMLA needs to be changed before merging

Test Plan

Unittest
E2E with GLM 5.2 and DSV3.2

Test Result

passed.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify

mergify Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--51724.org.readthedocs.build/en/51724/

@mergify mergify Bot added documentation Improvements or additions to documentation ci/build labels Aug 10, 2026
@mergify

mergify Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @sychen52.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@MatthewBonanni

Copy link
Copy Markdown
Member

Marking as [Do Not Merge] until vllm-project/FlashMLA#18 lands, this is convention to prevent landing a change to the FlashMLA source repo

@MatthewBonanni MatthewBonanni changed the title Enable W4A16 DSA [Do Not Merge] Enable W4A16 DSA Aug 10, 2026
@mergify mergify Bot removed the needs-rebase label Aug 10, 2026
@sychen52
sychen52 force-pushed the W4A8R8_DSA branch 6 times, most recently from d3703dd to c249781 Compare August 11, 2026 01:32
@github-actions

Copy link
Copy Markdown

✅ Queued 2 failed job(s) for retry in Buildkite CI #86016.

@sychen52

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

Copy link
Copy Markdown

✅ Queued 2 failed job(s) for retry in Buildkite CI #86016.

@sychen52 sychen52 changed the title Enable W4A16 DSA [Attention][DSA] Enable W4A16 DSA Aug 29, 2026
@sychen52

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #86110 for commit 9ee374ec2e92.

@sychen52

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

Copy link
Copy Markdown

✅ Queued 4 failed job(s) for retry in Buildkite CI #86110.

@sychen52

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

Copy link
Copy Markdown

✅ Queued 2 failed job(s) for retry in Buildkite CI #86110.

@LucasWilkinson LucasWilkinson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM other then a couple nits

Comment thread vllm/v1/kv_cache_interface.py Outdated
Comment thread vllm/config/cache.py Outdated
@sychen52

sychen52 commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor Author

For the 2 models and 5 tasks, NVFP4 NoPE and FP8 RoPE is almost always better than NVFP4 NoPE and RoPE (only exception is DSV3.2 on Tau2). The accuracy difference between FP8 and NVFP4 NoPE and FP8 RoPE kv is < 1% except DSV3.2 on Tau2.

model kv_cache_dtype AIME25 GPQA LCB AALCR Tau2
GLM-5.2-NVFP4 bfloat16 0.945 0.901 76.9 0.698 0.927
GLM-5.2-NVFP4 fp8 0.927 0.899 77.836 0.696 0.918
GLM-5.2-NVFP4 nvfp4_ds_mla 0.938 0.896 77.148 0.697 0.932
GLM-5.2-NVFP4 nvfp4_nvfp4RoPE_ds_mla 0.932 0.892 76.294 0.695 0.927
DeepSeek-V3.2 bfloat16 0.935 0.842 0.797 0.622 0.705 / 0.721
DeepSeek-V3.2 fp8 0.928 0.844 0.787 0.621 0.718 / 0.755
DeepSeek-V3.2 nvfp4_ds_mla 0.935 0.843 0.791 0.634 0.702 / 0.738
DeepSeek-V3.2 nvfp4_nvfp4RoPE_ds_mla 0.920 0.827 0.787 0.602 0.753 / 0.735

Note that nvfp4_nvfp4RoPE_ds_mla is not added in this PR. It is only used during experimentation.

AIME25: 64 repeats
GPQA: 64 repeats
LCB: 8 repeats
AALCR: 16 repeats
Tau2: 8 repeats, thinking false.
GLM5.2: reasoning_effort = max, output_cap=100k, except LCB (reasoning_effort = high, output_cp=250k, otherwise the output reaches the cap with no score)
DSV3.2: thinking = true (no effort set), output_cap=64k

@sychen52

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #86391 for commit 35bc90b3f7fa.

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
@sychen52

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #86408 for commit 9969b3c09c18.

@DarkLight1337
DarkLight1337 merged commit 7c5dc57 into vllm-project:main Sep 1, 2026
277 checks passed
am-cohere pushed a commit to am-cohere/vllm that referenced this pull request Sep 1, 2026
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
mylibrar pushed a commit to tanyuqian/vllm that referenced this pull request Sep 3, 2026
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
sychen52 added a commit to sychen52/vllm that referenced this pull request Sep 11, 2026
 vllm-project#52861 routed the DSA models to a fused norm+rope Triton kernel that
writes the MLA KV cache itself and only supports fp8_ds_mla. vllm-project#51724
added nvfp4_ds_mla after that and the rebase missed it.

Teach the fused kernel the nvfp4_ds_mla layout.

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
SeungminHeo pushed a commit to SeungminHeo/vllm that referenced this pull request Oct 4, 2026
The A.X-K2 line lets a dense EAGLE3/DFlash/DSpark draft ride a sparse-MLA
target by defaulting the draft's KV cache dtype to "auto" whenever the
target cache is fp8_ds_mla, since the draft's dense attention backend
cannot hold the opaque DS-MLA record. Upstream vllm-project#51724/vllm-project#55538 added a
second such record, nvfp4_ds_mla (SM100 only), which the planned B200
deployment will use; match on the "_ds_mla" suffix, as upstream's
attention layer does, so the override covers both.

Validation: ruff check/format and py_compile on vllm/config/vllm.py
passed on macOS; no GPU runtime test was run on this host.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Seungmin <seungminheo@sk.com>
(cherry picked from commit 6961a1a)
@gaby gaby mentioned this pull request Oct 5, 2026
15 of 22 tasks
SeungminHeo pushed a commit to SeungminHeo/vllm that referenced this pull request Oct 6, 2026
The A.X-K2 line lets a dense EAGLE3/DFlash/DSpark draft ride a sparse-MLA
target by defaulting the draft's KV cache dtype to "auto" whenever the
target cache is fp8_ds_mla, since the draft's dense attention backend
cannot hold the opaque DS-MLA record. Upstream vllm-project#51724/vllm-project#55538 added a
second such record, nvfp4_ds_mla (SM100 only), which the planned B200
deployment will use; match on the "_ds_mla" suffix, as upstream's
attention layer does, so the override covers both.

Validation: ruff check/format and py_compile on vllm/config/vllm.py
passed on macOS; no GPU runtime test was run on this host.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Seungmin <seungminheo@sk.com>
(cherry picked from commit 6961a1a)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/build documentation Improvements or additions to documentation ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants