Repository navigation
Conversation
…arse serves DCP GlmMoeDsaForCausalLM defaults DCP to the a2a combine with query replication (vllm-project#50382). FlashMLA sparse rejects anything but ag_rs under DCP (vllm-project#46514), and on SM90 it is the only sparse MLA backend that supports DCP at all, so GLM-5.x with DCP on Hopper defaulted into a configuration its only backend refuses: NotImplementedError: DCP for FlashMLA sparse is only validated with the default 'ag_rs' DCP comm backend; got 'a2a' Leave the stock defaults in place on SM90 and when FlashMLA sparse is selected explicitly. An explicit --dcp-comm-backend still wins, and the Blackwell default is unchanged. Part of vllm-project#59306. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: Mikhail Kostryukov <mike@triptrack.net>
drakosha
force-pushed
the
fix-glm-dcp-default-hopper
branch
from
October 10, 2026 16:43
d7f5e30 to
f6eb3d2
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Fixes the second failure in #59306, the DCP
a2adefault on Hopper. The MTP head failure from the same issue is fixed by #58209.GlmMoeDsaForCausalLM.verify_and_update_configdefaults DCP to thea2acombine with query replication (#50382). FlashMLA sparse rejects anything butag_rsunder DCP; that guard is ours, from #46514, merged two days before #50382. On SM90 FlashMLA sparse is also the only sparse MLA backend that supports DCP (FLASH_ATTN_MLA_SPARSEandFLASHINFER_MLA_SPARSE_SM90do not setsupports_dcp, so the selector drops them), which is why the startup log lists a single candidate:GLM-5.x with DCP on Hopper therefore defaulted into a configuration its only backend refuses and needed
--dcp-comm-backend ag_rs --no-dcp-q-replicateby hand.The hook now leaves the stock defaults in place where FlashMLA sparse serves DCP: on SM90 (CUDA), and when the backend is forced to
FLASHMLA_SPARSEon any platform. Unchanged:a2a+ qrep.--dcp-comm-backendstill wins.set_dcp_defaultsonly fills what the user left unset, so--dcp-comm-backend a2aon Hopper still reaches the guard and is refused as before.is_cuda().The other way to close this would be lifting the guard so FlashMLA sparse accepts the
a2acombine. The DSA path already merges throughMLADCPManager.combine, which implementsa2a, so it may work as is, but it has not been validated on Hopper and this PR does not touch the guard.The
q_replicate=Truehalf of the default is inert on the CUDA DSA pathSince #52861
GlmMoeDsaForCausalLMon CUDA isvllm/models/deepseek_v32/nvidia/model.py. ItsDeepseekV32Attentionbuildsq_b_projas a plainColumnParallelLinear, never passesdcp_q_replicatetoMLAAttention, and always all-gathers the query throughMLADCPManager.query_gather. Only the genericdeepseek_v2.pypath (anddots3_note,kimi_k3/amd) readsparallel_config.dcp_q_replicate. So on main the flag neither replicates the projection (no +4.57 GiB/rank) nor skips the all-gather, on Hopper and Blackwell alike; the qrep measurement in #50382 predates #52861 by two days. This PR leaves that default alone.dcp_q_replicate=Truestill shows up in the config log. @LucasWilkinson, should qrep be wired intoDeepseekV32Attention, or dropped from the GLM default?Not a duplicate
Searched open PRs for
dcp_comm_backend,a2a DCP,set_dcp_defaultsand59306 in:body. #56906 adds ana2adefault for Kimi-K3 on ROCm through the same hook, gated on the platform the same way, and does not touch GLM. #54472 is about direct A2A layouts. Nothing addresses the GLM default on Hopper.Test Plan
CPU-only unit tests next to the existing
set_dcp_defaultstests, withcurrent_platformpatched so both branches run on any runner:The parametrized test pins three cases: SM100 keeps
a2a+ qrep, SM90 gets the stockag_rswithout qrep,FLASHMLA_SPARSEforced on SM100 gets the stock defaults. A second test pins that an explicit--dcp-comm-backend a2ais not overridden.Both runs used the unpatched
vllm/vllm-openai:nightlyimage (af7f948) in a CPU container: first with only the test file swapped in, then with the patchedconfig.pymounted over the installed one.Test Result
Before, unpatched hook with the new tests:
After,
TestDCPCommBackendConfigplus the neighbouring CPU classTestLSEWeightedCombine:ruff check,ruff format --checkandtyposare clean on both files.Hardware: 4x H200 NVL,
Inferact/GLM-5.3-NVFP4, TP4 + DCP4 + EP, MTP 3,--kv-cache-dtype fp8_ds_mla, nightly af7f948 with this change and #58209 applied, and no DCP flags on the command line. The engine config logsdcp_comm_backend=ag_rs, the model starts withFLASHMLA_SPARSEand serves normally. Without this change the same command stops at theNotImplementedErrorabove. The change does not affect model output: it picks between configurations that already exist, and the one it picks was reachable through--dcp-comm-backend ag_rsbefore.AI assistance
AI assistance (Claude) was used for the investigation, the patch and the tests.
cc @LucasWilkinson