Skip to content

[Feature][PCP] Add DeepSeek-V4.1-Flash support - #59857

Open
wangyicong52 wants to merge 4 commits into
vllm-project:mainfrom
wangyicong52:feat/dsv4.1-flash-pcp
Open

wangyicong52 wants to merge 4 commits into
vllm-project:mainfrom
wangyicong52:feat/dsv4.1-flash-pcp

Conversation

@wangyicong52

@wangyicong52 wangyicong52 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Add text-only DeepSeek-V4.1-Flash Prefill Context Parallelism (PCP) support on NVIDIA CUDA to Model Runner V2 with FlashMLA sparse attention.

This PR intentionally targets a colocated single-model instance only. P/D disaggregation and KV transfer are not supported in this initial scope. The implementation keeps compressed KV, indexer, SWA, Engram lookback, MoE routing inputs, and padding consistent across PCP ranks while closing ratio-2 pairs split across PCP chunks with gathered predecessor rows.

Related to #25749 and #49109. This PR targets the DeepSeek-V4.1 ratio-1/2 cache path and does not add the V4 C4/C128 or MegaMoE work from #43809.

MoE Review Follow-up

The generic PCP MoE padding alignment fix has been extracted into draft #60946 and removed from this PR. PCP validation of this PR requires applying or merging that fix first.

Token-ID gathering and per-forward caching now live in the DeepSeek-V4.1 MoE model, covering both normal and deferred-finalization entries. Local input IDs remain available to embedding and Engram; this PR no longer changes the shared MoERunner.

Validation: TODO. The previously recorded validation results below belong to earlier revisions and do not establish validation of this reviewer-requested split.

Supported Scope

  • NVIDIA CUDA only: text-only DeepSeek-V4.1-Flash on Model Runner V2 with FLASHMLA_SPARSE_DSV41.
  • One colocated model instance with PP=DP=DCP=1.
  • Prefill PCP with arbitrary PCP chunk boundaries, including ranks with no local decode requests.
  • Tensor parallelism and PCP can be composed; the validation below covers TP8/PCP1, TP4/PCP2, and TP2/PCP4 on an 8-GPU node.
  • Expert parallel execution with the allgather_reducescatter all-to-all backend.
  • Eager execution with synchronous scheduling.

Not Supported Yet

  • P/D disaggregation or any KV connector.
  • ROCm.
  • Pipeline parallelism, data parallelism, decode context parallelism, or microbatching with PCP.
  • Prefix caching or SWA bounded replay with PCP.
  • Speculative decoding, CUDA graphs, or MegaMoE with PCP.
  • Multimodal inputs, encoder-decoder models, or LoRA with PCP.
  • Attention backends other than the DeepSeek-V4.1 FlashMLA sparse path.

These constraints are deliberate so the first upstream step establishes standalone PCP correctness without mixing in KV handoff or replay semantics. Follow-up work will expand one boundary at a time: first standalone cache/replay behavior, then additional single-instance parallel topologies and runtime features, and finally P/D disaggregation with a minimal direct connector before layering prefix caching, bounded replay, and speculative Decode support.

Implementation

  • Compress ratio-2 groups on each rank's local token rows. Gather only the last FP32 row of each local chunk to close cross-rank pairs and keep the open ring rows consistent.
  • Keep PCP linkage in model-specific compressor metadata; InputBatch and PCPManager remain unchanged relative to main.
  • Gather BF16 compressor latents in one cache-insert path shared with ratio 1, then use the aligned gathered positions and slots for main KV and Indexer K writes.
  • Keep Engram lookback, MoE routing token IDs, padding, compressed KV, indexer keys, and SWA metadata aligned across PCP ranks.
  • Keep PCP collectives ordered consistently across ranks, including empty local shards and requests with no local decode rows.

Engram embeddings remain TP-sharded, so increasing PCP can increase host memory requirements.

Rank-local Compressor Revision

Following LucasWilkinson's review, this revision adopts wangyicong52/vllm#1, authored by Lucas Wilkinson, at b06bea7abc4eb5cf6965fe704d16257ac7c70697. It removes the InputBatch.pcp_metadata / PCPBatchMetadata plumbing and the global FP32 kv_score gather.

Local validation of this revision: 101 passed, 6 CUDA-only skipped; applicable pre-commit hooks, mypy 3.10/3.12, syntax and diff checks passed. Additional CPU regressions cover PCP2/4 predecessor linkage, replicated short prefills, empty ranks, request isolation and newest open ring rows.

The GPU/model measurements below belong to the preceding global-order compressor implementation. They have not been rerun for this revision. The rank-local kernel, GSM8K and Prefill benchmark require renewed validation on the target H20 configuration before these historical results can support claims about this version.

Validation Environment

  • Hardware: one node with 8x NVIDIA H20 (96G, SM90) GPUs.
  • SM100 was not validated in this PR.
  • Image: vllm/vllm-openai:nightly-92044241a02f05de51420d654daf669de20b4691.

vLLM Launch Arguments

Only the TP/PCP pair changed across the accuracy arms: TP_SIZE=8, PCP_SIZE=1 for TP8/PCP1, TP_SIZE=4, PCP_SIZE=2 for TP4/PCP2, and TP_SIZE=2, PCP_SIZE=4 for TP2/PCP4. PP, DP, and DCP remained 1.

vllm serve /models/DeepSeek-V4.1-Flash \
  --host 0.0.0.0 \
  --port 18001 \
  --served-model-name deepseek-v4.1-flash \
  --language-model-only \
  --tokenizer-mode deepseek_v41 \
  --attention-backend FLASHMLA_SPARSE_DSV41 \
  --kv-cache-dtype fp8_ds_mla \
  --block-size 64 \
  --tensor-parallel-size "${TP_SIZE}" \
  --pipeline-parallel-size 1 \
  --decode-context-parallel-size 1 \
  --prefill-context-parallel-size "${PCP_SIZE}" \
  --enable-expert-parallel \
  --all2all-backend allgather_reducescatter \
  --moe-backend marlin \
  --engram-config '{"cpu_offload":true}' \
  --max-model-len auto \
  --max-num-batched-tokens 4096 \
  --max-num-seqs 64 \
  --gpu-memory-utilization 0.9 \
  --no-enable-prefix-caching \
  --no-swa-bounded-replay \
  --enforce-eager \
  --no-async-scheduling

Validation

evalscope eval \
  --eval-type openai_api \
  --api-url http://<HOST>:18001/v1 \
  --api-key EMPTY \
  --model deepseek-v4.1-flash \
  --model-id deepseek-v4.1-flash \
  --datasets gsm8k \
  --dataset-hub modelscope \
  --dataset-args '{"gsm8k":{"few_shot_num":4}}' \
  --seed 42 \
  --eval-batch-size 64 \
  --generation-config '{"temperature":0.0,"top_p":1.0,"max_tokens":65536,"reasoning_effort":"high","seed":42,"timeout":10000}' \
  --timeout 10000

EvalScope's GSM8K report exposes extracted-answer accuracy and per-request avg_output_tps; it does not report separate exact-match or flexible-match metrics.

Topology Correct Accuracy Avg output tok/s
TP8/PCP1 baseline 1281/1319 97.1190% 7.36
TP4/PCP2 1282/1319 97.1948% 7.04
TP2/PCP4 1284/1319 97.3465% 7.10

The Prefill benchmark used vllm bench serve with exact-length random prompts, one output token, seed 42, prefix caching disabled, one warmup wave, and four measured waves. Results below include input tok/s and mean TTFT in milliseconds; values in parentheses are changes relative to the TP8/PCP1 baseline.

Input tokens Max concurrency TP8/PCP1 input tok/s TP4/PCP2 input tok/s TP2/PCP4 input tok/s TP8/PCP1 mean TTFT (ms) TP4/PCP2 mean TTFT (ms) TP2/PCP4 mean TTFT (ms)
8192 1 6763.52 7914.17 (+17.01%) 8892.92 (+31.48%) 1210.98 1034.90 (-14.54%) 920.97 (-23.95%)
8192 4 6963.35 8220.16 (+18.05%) 9207.19 (+32.22%) 4267.48 3614.73 (-15.30%) 3227.61 (-24.37%)
8192 8 6967.68 8242.67 (+18.30%) 9215.61 (+32.26%) 8377.18 7082.04 (-15.46%) 6334.22 (-24.39%)
16384 1 6705.38 7929.32 (+18.25%) 8880.13 (+32.43%) 2443.16 2065.99 (-15.44%) 1844.79 (-24.49%)
16384 4 6897.41 8184.03 (+18.65%) 9184.73 (+33.16%) 8616.79 7264.02 (-15.70%) 6472.20 (-24.89%)
16384 8 6905.88 8204.16 (+18.80%) 9205.42 (+33.30%) 16906.55 14230.82 (-15.83%) 12682.85 (-24.98%)
32768 1 6629.52 7896.71 (+19.11%) 8818.36 (+33.02%) 4942.36 4149.20 (-16.05%) 3715.34 (-24.83%)
32768 4 6797.97 8121.66 (+19.47%) 9136.55 (+34.40%) 17484.04 14635.41 (-16.29%) 13010.39 (-25.59%)
32768 8 6816.65 8135.74 (+19.35%) 9163.40 (+34.43%) 34265.77 28708.44 (-16.22%) 25490.96 (-25.61%)

Humming Validation

Upstream PR #56997 ([Quantization] Prefer Humming before Marlin backends on SM90) changed main so SM90 MXFP4 automatic selection prefers Humming over Marlin for both linear and MoE kernels. Under PCP, TP4/PCP2 initially exposed a Humming permute-scratch capacity mismatch: PCP gathered 8,192 token rows while the old formula allocated 4,096 rows.

We revalidated this with upstream PR #60447, which replaces the configuration-derived capacity formula with a general input-sized scratch that grows during warmup and is fixed after the workspace manager locks. The validation used #60447 head 44ea227814 plus this PR's two commits, with no #60423 overlay. Both TP4/PCP2 and TP2/PCP4 selected HUMMING MXFP4 MoE and HummingMxfp8LinearKernel, matched all 206 overlaid runtime files, and completed startup with zero Pod restarts and no scratch or workspace-lock assertion.

The Humming run used vllm/vllm-openai:nightly-21d93d0d8c0e9627900020382bfce4730e61cab7 on one 8x H20 (SM90) node. The exact random Prefill matrix used one output token, seed 42, prefix caching disabled, one warmup wave, and four measured waves. EvalScope GSM8K used all 1,319 records, four-shot prompts, batch size 64, seed 42, temperature 0, top-p 1, and max_tokens=65536.

Humming GSM8K

Topology Correct Accuracy Avg output tok/s API errors Length finishes
TP4/PCP2 1282/1319 97.1948% 6.31 0 1
TP2/PCP4 1286/1319 97.4981% 6.19 0 0

Humming Prefill Performance

Input tokens Max concurrency TP4/PCP2 input tok/s TP4/PCP2 mean TTFT (ms) TP2/PCP4 input tok/s TP2/PCP4 mean TTFT (ms)
8192 1 5467.28 1498.15 5814.07 1408.78
8192 4 5596.76 5307.06 5967.52 4977.04
8192 8 5607.23 10412.47 5984.13 9757.40
16384 1 5478.95 2990.12 5833.22 2808.50
16384 4 5601.20 10607.57 5971.07 9949.80
16384 8 5613.44 20795.95 5990.78 19487.53
32768 1 5464.95 5995.72 5842.98 5607.76
32768 4 5579.34 21293.21 5962.29 19924.71
32768 8 5583.31 41820.42 5964.95 39143.57
Topology Nine-point mean input tok/s Previous Humming run Delta
TP4/PCP2 5554.72 5548.94 +0.10%
TP2/PCP4 5925.67 5917.19 +0.14%

The #60447 run has no measurable Prefill regression relative to the preceding Humming validation. Humming remains materially below the corresponding Marlin measurements, so the documented launch command remains pinned to Marlin.

One TP4/PCP2 GSM8K sample is a quality and latency caveat: sample 119 reached max_tokens=65536 after 9,701 seconds, while the matching TP2/PCP4 sample stopped naturally at 10,096 tokens. The TP4 result should therefore not be read as token-identical or as a clean quality-equivalence claim. This did not reproduce as a runtime failure: both topologies completed with zero API errors, zero Pod restarts, and zero critical server errors.

AI assistance was used to analyze, implement, and validate this draft. The human submitter retains responsibility for line-by-line review and final validation.


Pull Request Checklist
  • I used vLLM's /pr-checklist skill. (Mandatory for agents, optional for humans).

  • AI assistance was used during the creation of this PR.

  • Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.

  • Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.

  • Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.

  • Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.

@mergify

mergify Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--59857.org.readthedocs.build/en/59857/

@mergify mergify Bot added documentation Improvements or additions to documentation deepseek Related to DeepSeek models DSv4.1 Related to DeepSeek-V4.1 models mrv2 Model Runner V2 specific labels Oct 3, 2026
@wangyicong52
wangyicong52 force-pushed the feat/dsv4.1-flash-pcp branch 2 times, most recently from d9ef97a to ecd1cf1 Compare October 4, 2026 16:23
@mergify mergify Bot added kv-connector mooncake Mooncake KV-transfer / EC-transfer labels Oct 4, 2026
@wangyicong52 wangyicong52 changed the title [Feature][PCP] Add DeepSeek-V4.1-Flash support [Feature][PCP] Add DeepSeek-V4.1-Flash and Mooncake P/D support Oct 4, 2026
@wangyicong52
wangyicong52 force-pushed the feat/dsv4.1-flash-pcp branch from ded124f to aba224a Compare October 5, 2026 23:22
@wangyicong52 wangyicong52 changed the title [Feature][PCP] Add DeepSeek-V4.1-Flash and Mooncake P/D support [Feature][PCP] Add DeepSeek-V4.1-Flash support Oct 5, 2026
@wangyicong52
wangyicong52 force-pushed the feat/dsv4.1-flash-pcp branch 2 times, most recently from 85ee8dd to a7e9008 Compare October 6, 2026 03:57
@mergify

mergify Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @wangyicong52.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

✅ @wangyicong52, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • Your branch must contain every commit currently on its upstream target branch. Merge or rebase onto the latest target branch, then rerun the command. Append --allow-stale to a run command to test an outdated branch at your own risk.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@wangyicong52
wangyicong52 force-pushed the feat/dsv4.1-flash-pcp branch from 2203915 to 866ff6f Compare October 7, 2026 05:56
@wangyicong52

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

❌ This PR is 2 commits behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@wangyicong52
wangyicong52 force-pushed the feat/dsv4.1-flash-pcp branch from 866ff6f to ea9ce46 Compare October 7, 2026 07:42
@wangyicong52

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93253 for commit ea9ce46b067e.

@vllm-agent

vllm-agent commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

CI selector (shadow): 93 test steps (154 jobs) instead of 97 (128 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 97 (128) 93 (154) 38 (39) 34 (65)
AMD mirrors 96 (141) 142 (198) 5 (5) 51 (62)
Selector would run (93)
  • amd-kernels-mi355
  • amd-lm-eval-small-models-harness
  • amd-qwen3-next-mtp-async-eplb-accuracy
  • arm-cpu-test ×3
  • basic-models-tests-extra-initialization ×14
  • basic-models-tests-initialization
  • basic-models-tests-other
  • batch-invariance-b200
  • batch-invariance-h100
  • cpu-kernel-tests ×2
  • cpu-language-generation-and-pooling-model-tests ×3
  • cpu-multi-modal-model-tests-n ×4
  • cpu-multimodal-config
  • cpu-params-env-tokenizers-parser
  • cpu-reasoning-renderers
  • cpu-spec-decode-tests
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • deepseek-v4-kernel-test-b200
  • deepseek-v4-kernel-test-h100
  • distributed-dp-tests-2-gpus
  • distributed-dp-tests-4-gpus
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus
  • distributed-model-tests-2-gpus ×3
  • distributed-nixlconnector-pd-accuracy-4-gpus
  • distributed-tests-8xh100
  • distributed-torchrun-examples-4-gpus
  • distributed-torchrun-shutdown-tests-2-gpus
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • e2e-core-1-gpu
  • e2e-core-large-memory
  • elastic-ep-scaling-test
  • engine
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • fault-tolerance-e2e-2xh100
  • fusion-and-compile-unit-tests-2xb200
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-ar-rms-dynamo-partition-amd
  • fusion-e2e-tp2-ar-rms-inductor-partition-amd
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus
  • kernels-attention-test ×7
  • kernels-b200 ×3
  • kernels-flashmla-test-h100
  • kernels-fusedmoe-layer-test-2-b200s
  • kernels-fusedmoe-layer-test-2-h100s
  • kernels-mhc-test-b200
  • kernels-moe-test ×5
  • kernels-root-misc-test-b200
  • kimi-k3-prefix-cache-4xb200 ×2
  • kimi-k3-unit-tests-b200
  • kv-offload-large
  • kv-offload-medium
  • language-models-tests-extra-standard ×2
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • lm-eval-small-models
  • lora-tp-distributed ×4
  • model-executor
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • pipeline-context-parallelism-4-gpus
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus
  • pytorch-compilation-passes-unit-tests
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-test
  • quantization ×4
  • quantized-models-test
  • quantized-moe-test-b200
  • rayexecutorv2-4-gpus
  • replayssm-e2e
  • rust-frontend-distributed
  • sharded-rdt-weight-transfer
  • spec-decode-draft-model ×4
  • spec-decode-eagle-1-deepseek-qwen
  • spec-decode-mtp-deepseek-mimo
  • v1-attention-b200 ×2
  • v1-attention-h100-mi300 ×2
  • v1-executor-worker
  • v1-kv-connectors ×4
  • v1-others-cpu
  • v1-sample
  • v1-spec-decode
Would skip (today's rules run them) (38)
  • amd-fp8-moe-kernels-mi355
  • ascend-npu-test
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-test-other-cpu
  • benchmarks-cli-test
  • cpu-tool-parsers
  • distributed-compile-rpc-tests-2-gpus
  • distributed-compile-unit-tests-2xh100
  • e2e-scheduling-1-gpu
  • e2e-scheduling-accuracy-1-gpu
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • kernels-deepgemm-test-h100
  • kernels-fla-ops-test-b200
  • lm-eval-dspark-watermark-2xh100
  • lm-eval-watermarking
  • metrics-tracing-2-gpus
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • pytorch-compilation-dynamic-shapes
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • regression
  • rl-entrypoints-tests
  • samplers-multimodal-beam-search
  • samplers-test
  • spec-decode-mtp-gemma4
  • spec-decode-mtp-qwen3-5
  • spec-decode-speculators
  • v1-core
  • v1-kv-offload
  • v1-logits-oracle
  • v1-metrics-lmeval
Would add (today's rules do not run them) (34)
  • amd-kernels-mi355 (code map)
  • arm-cpu-test ×3 (code map)
  • basic-models-tests-extra-initialization ×14 (code map)
  • cpu-kernel-tests ×2 (code map)
  • cpu-multi-modal-model-tests-n ×4 (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • deepseek-v4-kernel-test-b200 (code map)
  • deepseek-v4-kernel-test-h100 (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-model-tests-2-gpus ×3 (code map)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-tests-8xh100 (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • elastic-ep-scaling-test (Python record)
  • engine (code map)
  • fusion-and-compile-unit-tests-2xb200 (code map)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • kernels-b200 ×3 (code map)
  • kernels-flashmla-test-h100 (code map)
  • kimi-k3-prefix-cache-4xb200 ×2 (Python record)
  • kimi-k3-unit-tests-b200 (code map)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • language-models-tests-extra-standard ×2 (code map)
  • lm-eval-small-models (Python record)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • quantization ×4 (code map)
  • quantized-models-test (Python record)
  • rayexecutorv2-4-gpus (code map)
  • sharded-rdt-weight-transfer (Python record)
  • spec-decode-draft-model ×4 (Python record)
  • spec-decode-eagle-1-deepseek-qwen (Python record)
AMD mirrors: would skip (5)
  • kernels-fla-ops-test-b200
  • kernels-fp4-moe-test-b200
  • kernels-fusedmoe-layer-test-2-b200s
  • platform-tests
  • v1-kv-offload
AMD mirrors: would add (51)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • cudagraph (code map)
  • deepseek-v4-kernel-test-h100 (code map)
  • distributed-comm-ops (code map)
  • distributed-compile-comm-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (code map)
  • distributed-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-tests-8xh100 (code map)
  • distributed-torchrun-examples-4-gpus (code map)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • engine (code map)
  • engine-1-gpu (code map)
  • entrypoints-unit-tests (code map)
  • examples (code map)
  • extract-hidden-states-integration (code map)
  • fusion-e2e-config-sweep-h100 (code map)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (code map)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (code map)
  • kernels-flashmla-test-h100 (code map)
  • kernels-quantization-test ×6 (code map)
  • kimi-k3-unit-tests-b200 (code map)
  • kv-offload-large (code map)
  • kv-offload-medium (code map)
  • kv-offload-small (code map)
  • lm-eval-turboquant-k3v4nc (code map)
  • lm-eval-turboquant-k8v4 (code map)
  • lm-eval-turboquant-t3nc (code map)
  • lm-eval-turboquant-t4nc (code map)
  • lora ×4 (code map)
  • mooncake-ec-tcp-e2e-2-gpus (code map)
  • multi-modal-accuracy-eval-small-models (code map)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (code map)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-spec-decode-acceptance-2-gpus (code map)
  • plugin-tests-2-gpus (code map)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (code map)
  • pytorch-nightly-dependency-override-check (code map)
  • quantization ×4 (code map)
  • quantized-models-test (code map)
  • ray-dependency-compatibility-check (code map)
  • rayexecutorv2-4-gpus (code map)
  • rust-frontend-core-correctness (code map)
  • rust-frontend-openai-coverage (code map)
  • rust-frontend-serve-admin-coverage (code map)
  • rust-frontend-tool-use (code map)
  • scale-out-ec-e2e-2-gpus (code map)
  • sharded-rdt-weight-transfer (code map)
  • torch-stable-abi-audit (code map)

16 changed files · base fd0216d3b5 · head a8d037f0a7 · Python record: build 93039 at 1e5d0ea888 · kernel record: table a982c81 (build 93415), map a982c81 · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 111 optional steps the selector would also run

@wangyicong52

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93542 for commit a8d037f0a7ce.

@LucasWilkinson LucasWilkinson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Id like to review this before it lands, we've been working hard to avoid pcp bleeding into InputBatch

@LucasWilkinson

Copy link
Copy Markdown
Contributor

I think ideally we would remove the input batch changes, maybe something like: wangyicong52#1

@wangyicong52

Copy link
Copy Markdown
Contributor Author

I think ideally we would remove the input batch changes, maybe something like: wangyicong52#1

@LucasWilkinson, Thanks for your feedback. I will try to follow your suggestion and retest it.

@wangyicong52
wangyicong52 force-pushed the feat/dsv4.1-flash-pcp branch from 6450d08 to 781c90d Compare October 9, 2026 02:48
wangyicong52 and others added 3 commits October 9, 2026 11:13
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: wangyicong <wangyicong@bytedance.com>
Signed-off-by: wangyicong <wangyicong@bytedance.com>
Adopt Lucas Wilkinson's proposal from #1. Link cross-chunk predecessor rows in compressor metadata and gather chunk boundary rows instead of the full FP32 token batch.

Remove PCPBatchMetadata and InputBatch plumbing, share the latent cache insert path across compression ratios, and drop unused ROCm Indexer slot-index arguments. Cover PCP2/4 boundary linkage, replicated short prefills, empty ranks and open ring writes with CPU regressions.

Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: wangyicong <wangyicong@bytedance.com>
@wangyicong52

wangyicong52 commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi @LucasWilkinson, I've retested this PR after incorporating your suggested changes, and the results look good.

Test environment:

GSM8K results:

Topology Correct Accuracy Avg output tok/s API errors Length finishes
TP8/PCP1 baseline 1283/1319 97.2707% 6.53 0 3 (119, 749, 982)
TP4/PCP2 1282/1319 97.1948% 6.10 0 2 (951, 1035)
TP2/PCP4 1284/1319 97.3465% 6.07 0 0

Prefill Performance result:

The Prefill benchmark used vllm bench serve with exact-length random prompts, one output token, seed 42, prefix caching disabled, one warmup wave, and four measured waves. Values in parentheses are changes relative to the TP8/PCP1 baseline from the same run.

Input tokens Max concurrency TP8/PCP1 input tok/s TP4/PCP2 input tok/s TP2/PCP4 input tok/s TP8/PCP1 mean TTFT (ms) TP4/PCP2 mean TTFT (ms) TP2/PCP4 mean TTFT (ms)
8192 1 4937.72 5444.05 (+10.25%) 5798.59 (+17.43%) 1658.83 1504.52 (-9.30%) 1412.55 (-14.85%)
8192 4 5040.89 5580.31 (+10.70%) 5946.76 (+17.97%) 5891.46 5322.78 (-9.65%) 4993.70 (-15.24%)
8192 8 5054.20 5606.39 (+10.93%) 5973.30 (+18.18%) 11553.44 10415.94 (-9.85%) 9775.99 (-15.38%)
16384 1 4911.40 5472.84 (+11.43%) 5824.22 (+18.59%) 3335.68 2993.47 (-10.26%) 2812.85 (-15.67%)
16384 4 5029.91 5597.32 (+11.28%) 5972.18 (+18.73%) 11811.53 10614.15 (-10.14%) 9947.94 (-15.78%)
16384 8 5042.89 5612.04 (+11.29%) 5982.51 (+18.63%) 23149.56 20803.19 (-10.14%) 19513.36 (-15.71%)
32768 1 4901.01 5469.72 (+11.60%) 5835.01 (+19.06%) 6685.64 5990.53 (-10.40%) 5615.42 (-16.01%)
32768 4 4988.05 5575.62 (+11.78%) 5956.64 (+19.42%) 23815.06 21307.28 (-10.53%) 19944.55 (-16.25%)
32768 8 4989.90 5579.51 (+11.82%) 5959.63 (+19.43%) 46791.28 41849.40 (-10.56%) 39178.76 (-16.27%)

Could you give another look when you have time? Thanks a lot!

Comment thread vllm/model_executor/layers/fused_moe/runner/moe_runner.py Outdated
Comment thread vllm/model_executor/layers/fused_moe/runner/moe_runner.py Outdated
Resolve LucasWilkinson's review comments by extracting the generic PCP MoE padding-mask fix into draft vllm-project#60946 and reverting the corresponding shared MoERunner changes from this PR.

Move PCP token-ID gathering and per-forward caching into the DeepSeek-V4.1 MoE model, covering both the normal and deferred-finalization entry points while retaining local IDs for embedding and Engram.

Review discussions: vllm-project#59857 (comment) and vllm-project#59857 (comment).

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: wangyicong <wangyicong@bytedance.com>
@wangyicong52

wangyicong52 commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi @LucasWilkinson, I've addressed both review requests in commit 7908ee8.

  • The generic PCP MoE padding alignment fix is now in [Bugfix][PCP] Align MoE routing padding with gathered tokens #60946. Its changes have been removed from this PR. By the way, after further testing, I found that this PR does not actually depend on this fix, although the fix still addresses a reproducible PCP bug in other models.
  • PCP token-ID gathering and caching now live in the DeepSeek-V4.1 MoE model.

Now vllm/model_executor/layers/fused_moe/runner/moe_runner.py remains unchanged.

Test environment:

TP8/PCP1 was not rerun; the matching previous baseline is retained below.

GSM8K results:

Topology Correct Accuracy Avg output tok/s API errors Length finishes
TP8/PCP1 retained baseline 1283/1319 97.2707% 6.53 0 3
TP4/PCP2 1281/1319 97.1190% 6.10 0 0
TP2/PCP4 1285/1319 97.4223% 6.09 0 4

Prefill performance:

The benchmark used exact-length random prompts, one output token, seed 42, prefix caching disabled, one warmup wave, and four measured waves. Values in parentheses are changes relative to the retained TP8/PCP1 baseline.

Input tokens Max concurrency TP8/PCP1 input tok/s TP4/PCP2 input tok/s TP2/PCP4 input tok/s TP8/PCP1 mean TTFT (ms) TP4/PCP2 mean TTFT (ms) TP2/PCP4 mean TTFT (ms)
8192 1 4937.72 5449.89 (+10.37%) 5795.10 (+17.36%) 1658.83 1502.93 (-9.40%) 1413.38 (-14.80%)
8192 4 5040.89 5596.34 (+11.02%) 5949.85 (+18.03%) 5891.46 5307.48 (-9.91%) 4991.95 (-15.27%)
8192 8 5054.20 5597.65 (+10.75%) 5964.54 (+18.01%) 11553.44 10431.43 (-9.71%) 9788.74 (-15.27%)
16384 1 4911.40 5443.87 (+10.84%) 5799.08 (+18.07%) 3335.68 3009.43 (-9.78%) 2825.05 (-15.31%)
16384 4 5029.91 5588.62 (+11.11%) 5958.02 (+18.45%) 11811.53 10632.74 (-9.98%) 9972.24 (-15.57%)
16384 8 5042.89 5610.05 (+11.25%) 5982.85 (+18.64%) 23149.56 20812.06 (-10.10%) 19513.88 (-15.71%)
32768 1 4901.01 5456.16 (+11.33%) 5829.14 (+18.94%) 6685.64 6005.35 (-10.18%) 5621.03 (-15.92%)
32768 4 4988.05 5576.66 (+11.80%) 5960.29 (+19.49%) 23815.06 21303.42 (-10.55%) 19931.88 (-16.31%)
32768 8 4989.90 5576.77 (+11.76%) 5960.96 (+19.46%) 46791.28 41868.81 (-10.52%) 39173.39 (-16.28%)

Both PCP arms completed smoke, the full Prefill benchmark, and full GSM8K with zero API errors and zero Pod restarts. Prefill shows no measurable regression relative to the previous matching runs.

Could you give it another look when you have time? Thanks a lot!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek Related to DeepSeek models documentation Improvements or additions to documentation DSv4.1 Related to DeepSeek-V4.1 models kv-connector mooncake Mooncake KV-transfer / EC-transfer mrv2 Model Runner V2 specific ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants