Skip to content

[Bugfix][Model] Keep GLM DSA's DCP defaults off a2a where FlashMLA sparse serves DCP - #59310

Open
drakosha wants to merge 1 commit into
vllm-project:mainfrom
drakosha:fix-glm-dcp-default-hopper
Open

drakosha wants to merge 1 commit into
vllm-project:mainfrom
drakosha:fix-glm-dcp-default-hopper

Conversation

@drakosha

Copy link
Copy Markdown
Contributor

Purpose

Fixes the second failure in #59306, the DCP a2a default on Hopper. The MTP head failure from the same issue is fixed by #58209.

GlmMoeDsaForCausalLM.verify_and_update_config defaults DCP to the a2a combine with query replication (#50382). FlashMLA sparse rejects anything but ag_rs under DCP; that guard is ours, from #46514, merged two days before #50382. On SM90 FlashMLA sparse is also the only sparse MLA backend that supports DCP (FLASH_ATTN_MLA_SPARSE and FLASHINFER_MLA_SPARSE_SM90 do not set supports_dcp, so the selector drops them), which is why the startup log lists a single candidate:

Using FLASHMLA_SPARSE attention backend out of potential backends: ['FLASHMLA_SPARSE']
NotImplementedError: DCP for FlashMLA sparse is only validated with the default 'ag_rs' DCP comm backend; got 'a2a'

GLM-5.x with DCP on Hopper therefore defaulted into a configuration its only backend refuses and needed --dcp-comm-backend ag_rs --no-dcp-q-replicate by hand.

The hook now leaves the stock defaults in place where FlashMLA sparse serves DCP: on SM90 (CUDA), and when the backend is forced to FLASHMLA_SPARSE on any platform. Unchanged:

  • Blackwell keeps a2a + qrep.
  • An explicit --dcp-comm-backend still wins. set_dcp_defaults only fills what the user left unset, so --dcp-comm-backend a2a on Hopper still reaches the guard and is refused as before.
  • ROCm, gated by is_cuda().

The other way to close this would be lifting the guard so FlashMLA sparse accepts the a2a combine. The DSA path already merges through MLADCPManager.combine, which implements a2a, so it may work as is, but it has not been validated on Hopper and this PR does not touch the guard.

The q_replicate=True half of the default is inert on the CUDA DSA path

Since #52861 GlmMoeDsaForCausalLM on CUDA is vllm/models/deepseek_v32/nvidia/model.py. Its DeepseekV32Attention builds q_b_proj as a plain ColumnParallelLinear, never passes dcp_q_replicate to MLAAttention, and always all-gathers the query through MLADCPManager.query_gather. Only the generic deepseek_v2.py path (and dots3_note, kimi_k3/amd) reads parallel_config.dcp_q_replicate. So on main the flag neither replicates the projection (no +4.57 GiB/rank) nor skips the all-gather, on Hopper and Blackwell alike; the qrep measurement in #50382 predates #52861 by two days. This PR leaves that default alone. dcp_q_replicate=True still shows up in the config log. @LucasWilkinson, should qrep be wired into DeepseekV32Attention, or dropped from the GLM default?

Not a duplicate

Searched open PRs for dcp_comm_backend, a2a DCP, set_dcp_defaults and 59306 in:body. #56906 adds an a2a default for Kimi-K3 on ROCm through the same hook, gated on the platform the same way, and does not touch GLM. #54472 is about direct A2A layouts. Nothing addresses the GLM default on Hopper.

Test Plan

CPU-only unit tests next to the existing set_dcp_defaults tests, with current_platform patched so both branches run on any runner:

pytest tests/distributed/test_dcp_a2a.py -k TestDCPCommBackendConfig

The parametrized test pins three cases: SM100 keeps a2a + qrep, SM90 gets the stock ag_rs without qrep, FLASHMLA_SPARSE forced on SM100 gets the stock defaults. A second test pins that an explicit --dcp-comm-backend a2a is not overridden.

Both runs used the unpatched vllm/vllm-openai:nightly image (af7f948) in a CPU container: first with only the test file swapped in, then with the patched config.py mounted over the installed one.

Test Result

Before, unpatched hook with the new tests:

FAILED tests/distributed/test_dcp_a2a.py::TestDCPCommBackendConfig::test_glm_moe_dsa_default_follows_backend[capability1-None-expected1]
FAILED tests/distributed/test_dcp_a2a.py::TestDCPCommBackendConfig::test_glm_moe_dsa_default_follows_backend[capability2-AttentionBackendEnum.FLASHMLA_SPARSE-expected2]
2 failed, 6 passed, 27 deselected in 2.46s

After, TestDCPCommBackendConfig plus the neighbouring CPU class TestLSEWeightedCombine:

18 passed, 17 deselected in 8.17s

ruff check, ruff format --check and typos are clean on both files.

Hardware: 4x H200 NVL, Inferact/GLM-5.3-NVFP4, TP4 + DCP4 + EP, MTP 3, --kv-cache-dtype fp8_ds_mla, nightly af7f948 with this change and #58209 applied, and no DCP flags on the command line. The engine config logs dcp_comm_backend=ag_rs, the model starts with FLASHMLA_SPARSE and serves normally. Without this change the same command stops at the NotImplementedError above. The change does not affect model output: it picks between configurations that already exist, and the one it picks was reachable through --dcp-comm-backend ag_rs before.

AI assistance

AI assistance (Claude) was used for the investigation, the patch and the tests.

cc @LucasWilkinson

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

…arse serves DCP

GlmMoeDsaForCausalLM defaults DCP to the a2a combine with query replication
(vllm-project#50382). FlashMLA sparse rejects anything but ag_rs under DCP (vllm-project#46514), and
on SM90 it is the only sparse MLA backend that supports DCP at all, so
GLM-5.x with DCP on Hopper defaulted into a configuration its only backend
refuses:

  NotImplementedError: DCP for FlashMLA sparse is only validated with the
  default 'ag_rs' DCP comm backend; got 'a2a'

Leave the stock defaults in place on SM90 and when FlashMLA sparse is
selected explicitly. An explicit --dcp-comm-backend still wins, and the
Blackwell default is unchanged.

Part of vllm-project#59306.

Co-Authored-By: Claude <noreply@anthropic.com>

Signed-off-by: Mikhail Kostryukov <mike@triptrack.net>
@drakosha
drakosha force-pushed the fix-glm-dcp-default-hopper branch from d7f5e30 to f6eb3d2 Compare October 10, 2026 16:43

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working glm

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant