Repository navigation
[Feature][PCP] Add DeepSeek-V4.1-Flash support - #59857
wangyicong52 wants to merge 4 commits into
Conversation
|
Documentation preview: https://vllm--59857.org.readthedocs.build/en/59857/ |
d9ef97a to
ecd1cf1
Compare
ded124f to
aba224a
Compare
85ee8dd to
a7e9008
Compare
|
This pull request has merge conflicts that must be resolved before it can be |
a7e9008 to
11a1d1e
Compare
|
✅ @wangyicong52, CI is now available for this PR.
|
2203915 to
866ff6f
Compare
|
/ci run |
|
❌ This PR is 2 commits behind upstream |
866ff6f to
ea9ce46
Compare
|
/ci run |
|
✅ Triggered Buildkite CI #93253 for commit |
CI selector (shadow): 93 test steps (154 jobs) instead of 97 (128 jobs)Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works. Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.
Selector would run (93)
Would skip (today's rules run them) (38)
Would add (today's rules do not run them) (34)
AMD mirrors: would skip (5)
AMD mirrors: would add (51)
16 changed files · base |
7f8c659 to
ea9ce46
Compare
ea9ce46 to
a8d037f
Compare
|
/ci run |
|
✅ Triggered Buildkite CI #93542 for commit |
LucasWilkinson
left a comment
There was a problem hiding this comment.
Id like to review this before it lands, we've been working hard to avoid pcp bleeding into InputBatch
|
I think ideally we would remove the input batch changes, maybe something like: wangyicong52#1 |
@LucasWilkinson, Thanks for your feedback. I will try to follow your suggestion and retest it. |
6450d08 to
781c90d
Compare
Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: wangyicong <wangyicong@bytedance.com>
Signed-off-by: wangyicong <wangyicong@bytedance.com>
Adopt Lucas Wilkinson's proposal from #1. Link cross-chunk predecessor rows in compressor metadata and gather chunk boundary rows instead of the full FP32 token batch. Remove PCPBatchMetadata and InputBatch plumbing, share the latent cache insert path across compression ratios, and drop unused ROCm Indexer slot-index arguments. Cover PCP2/4 boundary linkage, replicated short prefills, empty ranks and open ring writes with CPU regressions. Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: wangyicong <wangyicong@bytedance.com>
781c90d to
68600e9
Compare
|
Hi @LucasWilkinson, I've retested this PR after incorporating your suggested changes, and the results look good. Test environment:
GSM8K results:
Prefill Performance result: The Prefill benchmark used
Could you give another look when you have time? Thanks a lot! |
Resolve LucasWilkinson's review comments by extracting the generic PCP MoE padding-mask fix into draft vllm-project#60946 and reverting the corresponding shared MoERunner changes from this PR. Move PCP token-ID gathering and per-forward caching into the DeepSeek-V4.1 MoE model, covering both the normal and deferred-finalization entry points while retaining local IDs for embedding and Engram. Review discussions: vllm-project#59857 (comment) and vllm-project#59857 (comment). Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: wangyicong <wangyicong@bytedance.com>
|
Hi @LucasWilkinson, I've addressed both review requests in commit 7908ee8.
Now Test environment:
TP8/PCP1 was not rerun; the matching previous baseline is retained below. GSM8K results:
Prefill performance: The benchmark used exact-length random prompts, one output token, seed 42, prefix caching disabled, one warmup wave, and four measured waves. Values in parentheses are changes relative to the retained TP8/PCP1 baseline.
Both PCP arms completed smoke, the full Prefill benchmark, and full GSM8K with zero API errors and zero Pod restarts. Prefill shows no measurable regression relative to the previous matching runs. Could you give it another look when you have time? Thanks a lot! |
Overview
Add text-only DeepSeek-V4.1-Flash Prefill Context Parallelism (PCP) support on NVIDIA CUDA to Model Runner V2 with FlashMLA sparse attention.
This PR intentionally targets a colocated single-model instance only. P/D disaggregation and KV transfer are not supported in this initial scope. The implementation keeps compressed KV, indexer, SWA, Engram lookback, MoE routing inputs, and padding consistent across PCP ranks while closing ratio-2 pairs split across PCP chunks with gathered predecessor rows.
Related to #25749 and #49109. This PR targets the DeepSeek-V4.1 ratio-1/2 cache path and does not add the V4 C4/C128 or MegaMoE work from #43809.
MoE Review Follow-up
The generic PCP MoE padding alignment fix has been extracted into draft #60946 and removed from this PR. PCP validation of this PR requires applying or merging that fix first.
Token-ID gathering and per-forward caching now live in the DeepSeek-V4.1 MoE model, covering both normal and deferred-finalization entries. Local input IDs remain available to embedding and Engram; this PR no longer changes the shared MoERunner.
Validation: TODO. The previously recorded validation results below belong to earlier revisions and do not establish validation of this reviewer-requested split.
Supported Scope
FLASHMLA_SPARSE_DSV41.allgather_reducescatterall-to-all backend.Not Supported Yet
These constraints are deliberate so the first upstream step establishes standalone PCP correctness without mixing in KV handoff or replay semantics. Follow-up work will expand one boundary at a time: first standalone cache/replay behavior, then additional single-instance parallel topologies and runtime features, and finally P/D disaggregation with a minimal direct connector before layering prefix caching, bounded replay, and speculative Decode support.
Implementation
InputBatchandPCPManagerremain unchanged relative tomain.Engram embeddings remain TP-sharded, so increasing PCP can increase host memory requirements.
Rank-local Compressor Revision
Following LucasWilkinson's review, this revision adopts wangyicong52/vllm#1, authored by Lucas Wilkinson, at
b06bea7abc4eb5cf6965fe704d16257ac7c70697. It removes theInputBatch.pcp_metadata/PCPBatchMetadataplumbing and the global FP32kv_scoregather.Local validation of this revision: 101 passed, 6 CUDA-only skipped; applicable pre-commit hooks, mypy 3.10/3.12, syntax and diff checks passed. Additional CPU regressions cover PCP2/4 predecessor linkage, replicated short prefills, empty ranks, request isolation and newest open ring rows.
The GPU/model measurements below belong to the preceding global-order compressor implementation. They have not been rerun for this revision. The rank-local kernel, GSM8K and Prefill benchmark require renewed validation on the target H20 configuration before these historical results can support claims about this version.
Validation Environment
vllm/vllm-openai:nightly-92044241a02f05de51420d654daf669de20b4691.vLLM Launch Arguments
Only the TP/PCP pair changed across the accuracy arms:
TP_SIZE=8, PCP_SIZE=1for TP8/PCP1,TP_SIZE=4, PCP_SIZE=2for TP4/PCP2, andTP_SIZE=2, PCP_SIZE=4for TP2/PCP4. PP, DP, and DCP remained 1.Validation
EvalScope's GSM8K report exposes extracted-answer
accuracyand per-requestavg_output_tps; it does not report separate exact-match or flexible-match metrics.The Prefill benchmark used
vllm bench servewith exact-length random prompts, one output token, seed 42, prefix caching disabled, one warmup wave, and four measured waves. Results below include input tok/s and mean TTFT in milliseconds; values in parentheses are changes relative to the TP8/PCP1 baseline.Humming Validation
Upstream PR #56997 (
[Quantization] Prefer Humming before Marlin backends on SM90) changedmainso SM90 MXFP4 automatic selection prefers Humming over Marlin for both linear and MoE kernels. Under PCP, TP4/PCP2 initially exposed a Humming permute-scratch capacity mismatch: PCP gathered 8,192 token rows while the old formula allocated 4,096 rows.We revalidated this with upstream PR #60447, which replaces the configuration-derived capacity formula with a general input-sized scratch that grows during warmup and is fixed after the workspace manager locks. The validation used #60447 head
44ea227814plus this PR's two commits, with no #60423 overlay. Both TP4/PCP2 and TP2/PCP4 selectedHUMMINGMXFP4 MoE andHummingMxfp8LinearKernel, matched all 206 overlaid runtime files, and completed startup with zero Pod restarts and no scratch or workspace-lock assertion.The Humming run used
vllm/vllm-openai:nightly-21d93d0d8c0e9627900020382bfce4730e61cab7on one 8x H20 (SM90) node. The exact random Prefill matrix used one output token, seed 42, prefix caching disabled, one warmup wave, and four measured waves. EvalScope GSM8K used all 1,319 records, four-shot prompts, batch size 64, seed 42, temperature 0, top-p 1, andmax_tokens=65536.Humming GSM8K
Humming Prefill Performance
The #60447 run has no measurable Prefill regression relative to the preceding Humming validation. Humming remains materially below the corresponding Marlin measurements, so the documented launch command remains pinned to Marlin.
One TP4/PCP2 GSM8K sample is a quality and latency caveat: sample 119 reached
max_tokens=65536after 9,701 seconds, while the matching TP2/PCP4 sample stopped naturally at 10,096 tokens. The TP4 result should therefore not be read as token-identical or as a clean quality-equivalence claim. This did not reproduce as a runtime failure: both topologies completed with zero API errors, zero Pod restarts, and zero critical server errors.AI assistance was used to analyze, implement, and validate this draft. The human submitter retains responsibility for line-by-line review and final validation.
Pull Request Checklist
I used vLLM's
/pr-checklistskill. (Mandatory for agents, optional for humans).AI assistance was used during the creation of this PR.
Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.
Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.
Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.
Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.