Skip to content

[Model][GLM-5.3-Flash] Support DCP for the kpool sparse indexer - #59211

Merged
NickLucche merged 3 commits into
vllm-project:mainfrom
NickLucche:glm-flash-kpool-indexer-dcp
Oct 5, 2026
Merged

NickLucche merged 3 commits into
vllm-project:mainfrom
NickLucche:glm-flash-kpool-indexer-dcp

Conversation

@NickLucche

@NickLucche NickLucche commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Purpose

GLM-5.3-Flash cannot run with decode context parallelism. Its kpool sparse indexer compresses the indexer KV (index_kpool=4), and the indexer metadata builder rejects DCP whenever compress_ratio > 1. This check does not depend on the platform or the attention backend:

NotImplementedError: DCP is not supported with sparse indexer KV compression (compress_ratio=4)

This PR shards the kpool indexer across DCP ranks by pool:

  • --cp-kv-cache-interleave-size must be a multiple of index_kpool (e.g. 4), so each compressed pool lives entirely on one rank. The builder now raises a clear error when this does not hold.
  • CompressedSlotMappingKernel is DCP-aware. Each rank writes only the compressed states it owns, at rank-local positions. Prefill chunk metadata uses the interleave measured in compressed units.
  • Each rank scores its local pools, takes a local top-k, and merges it across ranks with the existing DSA DCP global top-k merge (_merge_dcp_topk_global), using the interleave in pool units. The result is global token ids, which the sparse MLA backend already filters and localizes under DCP.
  • The CuteDSL DCP merge warmup keys use pool-level top-k and interleave for kpool models.
  • The raw kpool tail ring (KpoolTailSpec) stays replicated on every DCP rank, like Mamba state. Its pages are self-addressed (uses_slot_mapping=False), so resolve_kv_cache_layout no longer treats it as a replicated draft group that requires a block-outer layout.
  • Unchanged: DeepSeek-V4 compressed KV stays unsupported under DCP, and ROCm kpool raises explicitly.

Repro

On main, this fails at engine startup with the error above. On Blackwell, FLASHINFER_MLA_SPARSE supports DCP; the indexer builder check applies on any platform.

vllm serve zai-org/GLM-5.3-Flash -tp 8 --decode-context-parallel-size 8

With this PR, set the interleave to a multiple of index_kpool:

vllm serve zai-org/GLM-5.3-Flash -tp 8 --decode-context-parallel-size 8 \
    --cp-kv-cache-interleave-size 4

Test Plan

pytest tests/v1/attention/test_indexer_deepseek_v4_slot_mapping.py \
    tests/v1/attention/test_indexer_dcp_localize.py \
    tests/models/glm5next/test_sparse_indexer_topk_dispatch.py \
    tests/v1/attention/test_kpool_tail_slot_mapping.py \
    tests/v1/core/test_kv_cache_utils.py

New tests:

  • test_compressed_slot_mapping_dcp_uses_rank_local_state_positions: each DCP rank writes only its own compressed states, at rank-local slots.
  • test_dcp_replicated_kpool_tail_keeps_block_interior_layout: a replicated kpool tail does not force a block-outer layout.

E2E: GSM8K (1319 questions) at 5, 25, 50 and 100 shots with vllm serve at TP8 + DCP8 + --cp-kv-cache-interleave-size 4, compared with TP8 alone (bf16 KV, no MTP).

Test Result

  • Unit tests: 309 passed on 1×H100. The 2 failures are test_get_kv_cache_config_mamba_hybrid_sharing_pp_*, which need more than 1 GPU for their PP config and are unrelated to this change.
  • E2E was run on 8×H100 together with a separate local SM90 attention-backend change that Hopper needs for DCP at head_dim 512 (out of scope here). I did not have Blackwell hardware to run this PR on its own.
Config GSM8K accuracy
TP8 0.911 – 0.922
TP8 + DCP8 (interleave 4) 0.916

5-shot GSM8K prompts are shorter than index_topk (2048 tokens), so that eval only lightly exercises the cross-rank top-k selection. To cover longer contexts, I raised the shot count. Both configs ran on the same build:

Shots Prompt tokens TP8 TP8 + DCP8 (interleave 4)
5 ~650 0.916 0.913
25 ~4.3k 0.909 0.917
50 ~8.5k 0.920 0.920
100 ~16.9k 0.926 0.922

Accuracy matches TP8 at every length, including 100 shots (about 8× index_topk). Prefix caching was on, so the shared few-shot prefix is mostly a cache hit, while decode and the question chunks run the top-k selection over the full context.

Duplicate check

I searched open PRs and issues for GLM-5.3-Flash / kpool / sparse-indexer DCP and found none that enables DCP for compressed indexer KV. #57161 reworks kpool compression and scheduling but does not touch DCP.

AI assistance

This PR was written with AI assistance (Claude). I reviewed the changes and ran the tests above.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added deepseek Related to DeepSeek models glm kv-cache-manager labels Sep 29, 2026
@mergify

mergify Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @NickLucche.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@GirasoleY

Copy link
Copy Markdown
Contributor

Thanks for adding the support, the sharding logic makes sense.

GSM8K prompts are shorter than index_topk (2048 tokens)

Could you run a eval with longer context? say gsm8k with --num_fewshot or longbench for accuracy?

@LucasWilkinson LucasWilkinson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM to me! thanks! left a couple (AI assisted) comments

Comment thread vllm/v1/attention/backends/mla/indexer.py
Comment thread vllm/models/glm5next/nvidia/sparse_indexer.py Outdated
cp_interleave = vllm_config.parallel_config.cp_kv_cache_interleave_size
kpool = getattr(hf_text_config, "index_kpool", None) or 1
if kpool > 1:
topk //= kpool

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As far as I can tell, _PACK_DCP_TOPK_CANDIDATES_KERNEL / _STABLE_TOPK_FROM_GATHERED_CANDIDATES_KERNEL only get register_warmup() from SparseAttnIndexer.__init__. Does SparseAttnIndexerKpool need the same, so these keys are actually warmed for GLM-5.x?

@NickLucche

Copy link
Copy Markdown
Member Author

@GirasoleY great point on the eval length, sorry I overlooked it!

I increased the shots here on gsm8k and updated PR description

┌───────┬───────────────┬───────┬───────────────────────────┐
│ Shots │ Prompt tokens │  TP8  │ TP8 + DCP8 (interleave 4) │
├───────┼───────────────┼───────┼───────────────────────────┤
│ 5     │ ~650          │ 0.916 │ 0.913                     │
├───────┼───────────────┼───────┼───────────────────────────┤
│ 25    │ ~4.3k         │ 0.909 │ 0.917                     │
├───────┼───────────────┼───────┼───────────────────────────┤
│ 50    │ ~8.5k         │ 0.920 │ 0.920                     │
├───────┼───────────────┼───────┼───────────────────────────┤
│ 100   │ ~16.9k        │ 0.926 │ 0.922                     │
└───────┴───────────────┴───────┴───────────────────────────┘

@NickLucche
NickLucche force-pushed the glm-flash-kpool-indexer-dcp branch from 292bfd4 to 33cc073 Compare October 4, 2026 10:53
Shard the kpool indexer by pool: with --cp-kv-cache-interleave-size a
multiple of index_kpool, each DCP rank owns whole pools, writes only its
compressed states, scores them locally and merges the per-rank top-k with
the existing DCP global top-k merge (in pool units). The raw tail ring
stays replicated, like Mamba state, and no longer forces a block-outer KV
cache layout. DeepSeek-V4 compression remains unsupported under DCP.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
@NickLucche
NickLucche force-pushed the glm-flash-kpool-indexer-dcp branch from 33cc073 to 1f520bc Compare October 4, 2026 11:03
@mergify mergify Bot removed the needs-rebase label Oct 4, 2026
@NickLucche

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

❌ This PR is 2 commits behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

NickLucche and others added 2 commits October 5, 2026 08:49
…out DCP

Only convert the interleave to compressed-state/pool units when DCP is
enabled, so the non-DCP path keeps passing the token interleave (1) to
BuildPrefillChunkMetadataKernel instead of 1 // compress_ratio = 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
NIXL P/D can adjust cp_kv_cache_interleave_size after the model is built,
and the indexer metadata builder is created after that. Read it lazily, as
SparseAttnIndexer does, so the layer and the builder agree on the
local -> global pool mapping.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
@NickLucche

Copy link
Copy Markdown
Member Author

/ci run --allow-stale

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92865 for commit c5931ddcd5e3.

⚠️ This PR is 23 commits behind upstream main. Running CI at your own risk because --allow-stale was requested; outdated CI configuration may cause failures. Before merging, merge or rebase onto the latest main, then rerun /ci run on the latest PR commit.

@vllm-agent

Copy link
Copy Markdown
Contributor

CI selector (shadow): 152 test steps (215 jobs) instead of 76 (100 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 76 (100) 152 (215) 6 (6) 82 (121)
AMD mirrors 74 (101) 136 (191) 5 (8) 67 (98)
Selector would run (152)
  • amd-kernels-mi355
  • amd-lm-eval-small-models-harness
  • amd-qwen3-next-mtp-async-eplb-accuracy
  • arm-cpu-test ×3
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-test-other-cpu
  • basic-models-tests-extra-initialization ×14
  • basic-models-tests-initialization
  • basic-models-tests-other
  • batch-invariance-b200
  • batch-invariance-h100
  • benchmarks-cli-test
  • cpu-kernel-tests ×2
  • cpu-language-generation-and-pooling-model-tests ×3
  • cpu-multi-modal-model-tests-n ×4
  • cpu-multimodal-config
  • cpu-params-env-tokenizers-parser
  • cpu-spec-decode-tests
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • cudagraph
  • deepseek-v4-kernel-test-b200
  • distributed-compile-comm-4-gpus
  • distributed-compile-rpc-tests-2-gpus
  • distributed-dp-tests-2-gpus
  • distributed-dp-tests-4-gpus
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus
  • distributed-model-tests-2-gpus ×3
  • distributed-mooncakeconnector-pd-accuracy-4-gpus
  • distributed-nixlconnector-pd-accuracy-4-gpus
  • distributed-tests-8xh100
  • distributed-torchrun-examples-4-gpus
  • distributed-torchrun-shutdown-tests-2-gpus
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • e2e-core-1-gpu
  • e2e-core-large-memory
  • e2e-scheduling-1-gpu
  • e2e-scheduling-accuracy-1-gpu
  • elastic-ep-scaling-test
  • engine
  • engine-1-gpu
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • entrypoints-unit-tests
  • examples
  • extract-hidden-states-integration
  • extract-hidden-states-integration-2-gpus
  • fault-tolerance-e2e-2xh100
  • fusion-and-compile-unit-tests-2xb200
  • fusion-e2e-config-sweep-h100
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-ar-rms-config-sweep-h100
  • fusion-e2e-tp2-ar-rms-dynamo-partition-amd
  • fusion-e2e-tp2-ar-rms-inductor-partition-amd
  • fusion-e2e-tp2-asynctp-config-sweep-h100
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • glm5next-unit-tests
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus
  • kernels-attention-test ×7
  • kernels-b200 ×3
  • kernels-core-operation-test ×3
  • kernels-deepgemm-test-h100
  • kernels-flashmla-test-h100
  • kernels-root-misc-test-b200
  • kimi-k3-prefix-cache-4xb200 ×2
  • kimi-k3-unit-tests-b200
  • kv-offload-large
  • kv-offload-medium
  • kv-offload-small
  • language-models-tests-extra-standard ×2
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • lm-eval-dspark-watermark-2xh100
  • lm-eval-small-models
  • lm-eval-turboquant-k3v4nc
  • lm-eval-turboquant-k8v4
  • lm-eval-turboquant-t3nc
  • lm-eval-turboquant-t4nc
  • lm-eval-watermarking
  • lora ×4
  • lora-tp-distributed ×4
  • metrics-tracing-2-gpus
  • model-executor
  • mooncake-ec-tcp-e2e-2-gpus
  • mrcr-eval-small-models
  • multi-modal-accuracy-eval-small-models
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus
  • nixlconnector-pd-edge-cases-2-gpus
  • openai-api-correctness
  • pipeline-context-parallelism-4-gpus
  • plugin-tests-2-gpus
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus
  • pytorch-compilation-dynamic-shapes
  • pytorch-compilation-passes-unit-tests
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • pytorch-fullgraph-test
  • quantization ×4
  • quantized-models-test
  • quantized-moe-test-b200
  • rayexecutorv2-4-gpus
  • regression
  • replayssm-e2e
  • rust-frontend-core-correctness
  • rust-frontend-distributed
  • rust-frontend-openai-coverage
  • rust-frontend-serve-admin-coverage
  • rust-frontend-tool-use
  • samplers-multimodal-beam-search
  • samplers-test
  • scale-out-ec-e2e-2-gpus
  • sharded-rdt-weight-transfer
  • spec-decode-draft-model ×4
  • spec-decode-eagle-1-deepseek-qwen
  • spec-decode-eagle-2-llama3-qwen-vl-other
  • spec-decode-mtp-deepseek-mimo
  • spec-decode-mtp-gemma4
  • spec-decode-mtp-qwen3-5
  • spec-decode-ngram-suffix
  • spec-decode-speculators
  • v1-attention-b200 ×2
  • v1-attention-h100-mi300 ×2
  • v1-core
  • v1-executor-worker
  • v1-kv-connectors ×4
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-others-cpu
  • v1-sample
  • v1-spec-decode
Would skip (today's rules run them) (6)
  • ascend-npu-test
  • cpu-reasoning-renderers
  • cpu-tool-parsers
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • v1-kv-offload
Would add (today's rules do not run them) (82)
  • amd-kernels-mi355 (code map)
  • arm-cpu-test ×3 (code map)
  • basic-models-tests-extra-initialization ×14 (code map)
  • cpu-kernel-tests ×2 (code map)
  • cpu-multi-modal-model-tests-n ×4 (code map)
  • cpu-spec-decode-tests (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • cudagraph (Python record)
  • deepseek-v4-kernel-test-b200 (code map)
  • distributed-compile-comm-4-gpus (Python record)
  • distributed-compile-rpc-tests-2-gpus (Python record)
  • distributed-dp-tests-2-gpus (Python record)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-model-tests-2-gpus ×3 (code map)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (Python record)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-tests-8xh100 (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • distributed-torchrun-shutdown-tests-2-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • elastic-ep-scaling-test (Python record)
  • engine (code map)
  • engine-1-gpu (Python record)
  • entrypoints-unit-tests (Python record)
  • examples (Python record)
  • extract-hidden-states-integration (Python record)
  • extract-hidden-states-integration-2-gpus (Python record)
  • fault-tolerance-e2e-2xh100 (Python record)
  • fusion-and-compile-unit-tests-2xb200 (code map)
  • fusion-e2e-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-ar-rms-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (Python record)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • kernels-b200 ×3 (code map)
  • kernels-core-operation-test ×3 (code map)
  • kernels-deepgemm-test-h100 (Python record)
  • kernels-flashmla-test-h100 (code map)
  • kimi-k3-prefix-cache-4xb200 ×2 (code map)
  • kimi-k3-unit-tests-b200 (code map)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • kv-offload-small (Python record)
  • language-models-tests-extra-standard ×2 (code map)
  • lm-eval-dspark-watermark-2xh100 (Python record)
  • lm-eval-small-models (Python record)
  • lm-eval-turboquant-k3v4nc (Python record)
  • lm-eval-turboquant-k8v4 (Python record)
  • lm-eval-turboquant-t3nc (Python record)
  • lm-eval-turboquant-t4nc (Python record)
  • lm-eval-watermarking (Python record)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (Python record)
  • mooncake-ec-tcp-e2e-2-gpus (Python record)
  • mrcr-eval-small-models (Python record)
  • multi-modal-accuracy-eval-small-models (Python record)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (Python record)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-edge-cases-2-gpus (Python record)
  • openai-api-correctness (Python record)
  • pipeline-context-parallelism-4-gpus (code map)
  • plugin-tests-2-gpus (Python record)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • quantization ×4 (code map)
  • quantized-models-test (Python record)
  • quantized-moe-test-b200 (Python record)
  • rayexecutorv2-4-gpus (code map)
  • replayssm-e2e (Python record)
  • rust-frontend-core-correctness (Python record)
  • rust-frontend-distributed (Python record)
  • rust-frontend-openai-coverage (Python record)
  • rust-frontend-serve-admin-coverage (Python record)
  • rust-frontend-tool-use (Python record)
  • samplers-multimodal-beam-search (Python record)
  • samplers-test (Python record)
  • scale-out-ec-e2e-2-gpus (Python record)
  • sharded-rdt-weight-transfer (Python record)
  • spec-decode-draft-model ×4 (Python record)
  • spec-decode-eagle-1-deepseek-qwen (Python record)
  • spec-decode-eagle-2-llama3-qwen-vl-other (Python record)
  • spec-decode-ngram-suffix (Python record)
AMD mirrors: would skip (5)
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • multi-modal-processor ×4
  • platform-tests
  • v1-kv-offload
AMD mirrors: would add (67)
  • basic-models-tests-extra-initialization ×14 (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • cudagraph (code map)
  • distributed-compile-rpc-tests-2-gpus (code map)
  • distributed-dp-tests-2-gpus (code map)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (code map)
  • distributed-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-tests-8xh100 (code map)
  • distributed-torchrun-examples-4-gpus (code map)
  • distributed-torchrun-shutdown-tests-2-gpus (code map)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • engine (code map)
  • engine-1-gpu (code map)
  • entrypoints-unit-tests (code map)
  • examples (code map)
  • extract-hidden-states-integration (code map)
  • fusion-e2e-config-sweep-h100 (code map)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (code map)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (code map)
  • jit-monitor-no-runtime-jit (code map)
  • kernels-core-operation-test ×3 (code map)
  • kernels-flashmla-test-h100 (code map)
  • kernels-moe-test ×5 (code map)
  • kimi-k3-unit-tests-b200 (code map)
  • kv-offload-large (code map)
  • kv-offload-medium (code map)
  • kv-offload-small (code map)
  • lm-eval-dspark-watermark-2xh100 (code map)
  • lm-eval-turboquant-k3v4nc (code map)
  • lm-eval-turboquant-k8v4 (code map)
  • lm-eval-turboquant-t3nc (code map)
  • lm-eval-turboquant-t4nc (code map)
  • lm-eval-watermarking (code map)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (code map)
  • mooncake-ec-tcp-e2e-2-gpus (code map)
  • multi-modal-accuracy-eval-small-models (code map)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (code map)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-spec-decode-acceptance-2-gpus (code map)
  • plugin-tests-2-gpus (code map)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (code map)
  • pytorch-nightly-dependency-override-check (code map)
  • quantization ×4 (code map)
  • quantized-models-test (code map)
  • ray-dependency-compatibility-check (code map)
  • rayexecutorv2-4-gpus (code map)
  • rust-frontend-cargo-style-clippy (code map)
  • rust-frontend-cargo-tests (code map)
  • rust-frontend-core-correctness (code map)
  • rust-frontend-distributed (code map)
  • rust-frontend-openai-coverage (code map)
  • rust-frontend-serve-admin-coverage (code map)
  • rust-frontend-tool-use (code map)
  • samplers-multimodal-beam-search (code map)
  • samplers-test (code map)
  • scale-out-ec-e2e-2-gpus (code map)
  • sharded-rdt-weight-transfer (code map)
  • spec-decode-draft-model ×4 (code map)
  • spec-decode-eagle-1-deepseek-qwen (code map)
  • spec-decode-eagle-2-llama3-qwen-vl-other (code map)
  • spec-decode-ngram-suffix (code map)
  • torch-stable-abi-audit (code map)

9 changed files · base 155d23cb00 · head c5931ddcd5 · Python record: build 92817 at 7867d6c52d · kernel record: table 7867d6c (build 92817), map 7867d6c · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 126 optional steps the selector would also run

@NickLucche
NickLucche merged commit 30d4032 into vllm-project:main Oct 5, 2026
200 checks passed
andakai added a commit to andakai/vllm that referenced this pull request Oct 5, 2026
Conflicts with vllm-project#59211 (DCP for the GLM-5.3 kpool sparse indexer):
- glm5next/common/attention.py: keep the CircularBufferSpec tail (no
  sliding_window) and take the new comment on dcp_sharded=False.
- test_indexer_deepseek_v4_slot_mapping.py: take the new DCP builder
  fields; drop builder.kernel_block_size, which builders no longer have.
- test_kv_cache_utils.py: build the replicated tail in the new DCP layout
  test as a CircularBufferSpec, which replaces KpoolTailSpec here.

Signed-off-by: Dakai An <dakaian108@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek Related to DeepSeek models glm kv-cache-manager

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants