Repository navigation
[Model][GLM-5.3-Flash] Support DCP for the kpool sparse indexer - #59211
Conversation
|
This pull request has merge conflicts that must be resolved before it can be |
|
Thanks for adding the support, the sharding logic makes sense.
Could you run a eval with longer context? say gsm8k with |
LucasWilkinson
left a comment
There was a problem hiding this comment.
LGTM to me! thanks! left a couple (AI assisted) comments
| cp_interleave = vllm_config.parallel_config.cp_kv_cache_interleave_size | ||
| kpool = getattr(hf_text_config, "index_kpool", None) or 1 | ||
| if kpool > 1: | ||
| topk //= kpool |
There was a problem hiding this comment.
As far as I can tell, _PACK_DCP_TOPK_CANDIDATES_KERNEL / _STABLE_TOPK_FROM_GATHERED_CANDIDATES_KERNEL only get register_warmup() from SparseAttnIndexer.__init__. Does SparseAttnIndexerKpool need the same, so these keys are actually warmed for GLM-5.x?
|
@GirasoleY great point on the eval length, sorry I overlooked it! I increased the shots here on gsm8k and updated PR description |
292bfd4 to
33cc073
Compare
Shard the kpool indexer by pool: with --cp-kv-cache-interleave-size a multiple of index_kpool, each DCP rank owns whole pools, writes only its compressed states, scores them locally and merges the per-rank top-k with the existing DCP global top-k merge (in pool units). The raw tail ring stays replicated, like Mamba state, and no longer forces a block-outer KV cache layout. DeepSeek-V4 compression remains unsupported under DCP. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
33cc073 to
1f520bc
Compare
|
/ci run |
|
❌ This PR is 2 commits behind upstream |
…out DCP Only convert the interleave to compressed-state/pool units when DCP is enabled, so the non-DCP path keeps passing the token interleave (1) to BuildPrefillChunkMetadataKernel instead of 1 // compress_ratio = 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
NIXL P/D can adjust cp_kv_cache_interleave_size after the model is built, and the indexer metadata builder is created after that. Read it lazily, as SparseAttnIndexer does, so the layer and the builder agree on the local -> global pool mapping. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
3a5b08d to
c5931dd
Compare
|
/ci run --allow-stale |
|
✅ Triggered Buildkite CI #92865 for commit
|
CI selector (shadow): 152 test steps (215 jobs) instead of 76 (100 jobs)Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works. Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.
Selector would run (152)
Would skip (today's rules run them) (6)
Would add (today's rules do not run them) (82)
AMD mirrors: would skip (5)
AMD mirrors: would add (67)
9 changed files · base |
Conflicts with vllm-project#59211 (DCP for the GLM-5.3 kpool sparse indexer): - glm5next/common/attention.py: keep the CircularBufferSpec tail (no sliding_window) and take the new comment on dcp_sharded=False. - test_indexer_deepseek_v4_slot_mapping.py: take the new DCP builder fields; drop builder.kernel_block_size, which builders no longer have. - test_kv_cache_utils.py: build the replicated tail in the new DCP layout test as a CircularBufferSpec, which replaces KpoolTailSpec here. Signed-off-by: Dakai An <dakaian108@gmail.com>
Purpose
GLM-5.3-Flash cannot run with decode context parallelism. Its kpool sparse indexer compresses the indexer KV (
index_kpool=4), and the indexer metadata builder rejects DCP whenevercompress_ratio > 1. This check does not depend on the platform or the attention backend:This PR shards the kpool indexer across DCP ranks by pool:
--cp-kv-cache-interleave-sizemust be a multiple ofindex_kpool(e.g. 4), so each compressed pool lives entirely on one rank. The builder now raises a clear error when this does not hold.CompressedSlotMappingKernelis DCP-aware. Each rank writes only the compressed states it owns, at rank-local positions. Prefill chunk metadata uses the interleave measured in compressed units._merge_dcp_topk_global), using the interleave in pool units. The result is global token ids, which the sparse MLA backend already filters and localizes under DCP.KpoolTailSpec) stays replicated on every DCP rank, like Mamba state. Its pages are self-addressed (uses_slot_mapping=False), soresolve_kv_cache_layoutno longer treats it as a replicated draft group that requires a block-outer layout.Repro
On
main, this fails at engine startup with the error above. On Blackwell,FLASHINFER_MLA_SPARSEsupports DCP; the indexer builder check applies on any platform.With this PR, set the interleave to a multiple of
index_kpool:vllm serve zai-org/GLM-5.3-Flash -tp 8 --decode-context-parallel-size 8 \ --cp-kv-cache-interleave-size 4Test Plan
pytest tests/v1/attention/test_indexer_deepseek_v4_slot_mapping.py \ tests/v1/attention/test_indexer_dcp_localize.py \ tests/models/glm5next/test_sparse_indexer_topk_dispatch.py \ tests/v1/attention/test_kpool_tail_slot_mapping.py \ tests/v1/core/test_kv_cache_utils.pyNew tests:
test_compressed_slot_mapping_dcp_uses_rank_local_state_positions: each DCP rank writes only its own compressed states, at rank-local slots.test_dcp_replicated_kpool_tail_keeps_block_interior_layout: a replicated kpool tail does not force a block-outer layout.E2E: GSM8K (1319 questions) at 5, 25, 50 and 100 shots with
vllm serveat TP8 + DCP8 +--cp-kv-cache-interleave-size 4, compared with TP8 alone (bf16 KV, no MTP).Test Result
test_get_kv_cache_config_mamba_hybrid_sharing_pp_*, which need more than 1 GPU for their PP config and are unrelated to this change.5-shot GSM8K prompts are shorter than
index_topk(2048 tokens), so that eval only lightly exercises the cross-rank top-k selection. To cover longer contexts, I raised the shot count. Both configs ran on the same build:Accuracy matches TP8 at every length, including 100 shots (about 8×
index_topk). Prefix caching was on, so the shared few-shot prefix is mostly a cache hit, while decode and the question chunks run the top-k selection over the full context.Duplicate check
I searched open PRs and issues for GLM-5.3-Flash / kpool / sparse-indexer DCP and found none that enables DCP for compressed indexer KV. #57161 reworks kpool compression and scheduling but does not touch DCP.
AI assistance
This PR was written with AI assistance (Claude). I reviewed the changes and ran the tests above.