Skip to content

[DeepSeekV4][PCP] Enable context-parallel prefill for hybrid KV and MegaMoE - #43809

Open
baonudesifeizhai wants to merge 6 commits into
vllm-project:mainfrom
baonudesifeizhai:depskpcpsupportv3
Open

baonudesifeizhai wants to merge 6 commits into
vllm-project:mainfrom
baonudesifeizhai:depskpcpsupportv3

Conversation

@baonudesifeizhai

@baonudesifeizhai baonudesifeizhai commented May 27, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

#40902
Implementation summary:

Added PCP-aware hybrid KV handling for DeepSeekV4, covering MLA / SWA / C4 / C128 block and slot mapping.
Added PCP merge for sparse prefill attention: each PCP rank computes local partial attention, then merges LSE/output globally and applies attn_sink only once.
Added a MegaMoE PCP token-shard path: after attention, each PCP rank runs router / MegaMoE / shared expert only for its local token shard, then all-gathers full hidden states for the next layer.
Main challenges:

DeepSeekV4 uses multiple KV layouts with different block sizes and compression ratios.
C4/C128 compressor states need global PCP merge, not local-only updates.
attn_sink must not be applied independently on each PCP rank.
Naive PCP duplicated MoE work across ranks, causing large overhead.
The MoE token-shard path had to be wrapped as a custom op to work cleanly with CUDA graph / torch compile.

Aligned common-length results:

**TTFT**

| Length | TP8+EP+MegaMoE TTFT | PCP2 TTFT | Improvement |
|---|---:|---:|---:|
| 128K | 11.17s | 8.73s | ~21.8% faster |
| 256K | 21.20s | 18.38s | ~13.3% faster |
| 512K | 49.19s | 41.01s | ~16.6% faster |
| 1M | 126.68s | 95.59s | ~24.5% faster |

**Throughput**

| Length | TP8+EP+MegaMoE tok/s | PCP2 tok/s | Improvement |
|---|---:|---:|---:|
| 128K | 11,739 | 15,005 | +27.8% |
| 256K | 12,365 | 14,264 | +15.4% |
| 512K | 10,657 | 12,784 | +20.0% |
| 1M | 8,277 | 10,970 | +32.5% | 

tp =8 +maga moe:
start:

MODEL=/root/model/DeepSeek-V4-Pro

CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 \
.venv/bin/python -m vllm.entrypoints.cli.main serve "$MODEL" \
  --host 0.0.0.0 \
  --port 8000 \
  --served-model-name dsv4pro \
  --tensor-parallel-size 8 \
  --decode-context-parallel-size 1 \
  --prefill-context-parallel-size 1 \
  --enable-expert-parallel \
  --tokenizer-mode deepseek_v4 \
  --trust-remote-code \
  --dtype bfloat16 \
  --kv-cache-dtype fp8 \
  --max-model-len 1048576 \
  --max-num-batched-tokens 32768 \
  --enable-chunked-prefill \
  --gpu-memory-utilization 0.90 \
  --moe-backend deep_gemm_mega_moe \
  --no-enable-prefix-caching

res : https://paste.ubuntu.com/p/mcQrbwqVN3/
tp =4 pcp =2 maga moe:
MODEL=/root/model/DeepSeek-V4-Pro

CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 \
.venv/bin/python -m vllm.entrypoints.cli.main serve "$MODEL" \
  --host 0.0.0.0 \
  --port 8000 \
  --served-model-name dsv4pro \
  --tensor-parallel-size 4 \
  --decode-context-parallel-size 1 \
  --prefill-context-parallel-size 2 \
  --enable-expert-parallel \
  --tokenizer-mode deepseek_v4 \
  --trust-remote-code \
  --dtype bfloat16 \
  --kv-cache-dtype fp8 \
  --max-model-len 1048576 \
  --max-num-batched-tokens 32768 \
  --enable-chunked-prefill \
  --gpu-memory-utilization 0.90 \
  --moe-backend deep_gemm_mega_moe \
  --no-enable-prefix-caching
 

res: https://docs.google.com/document/d/1upWKu1zhqfJpRG9n-fdQA8k3RhNeIxIuh9FJNysN_gQ/edit?tab=t.0

tp=4 dp =2 +maga moe : https://paste.ubuntu.com/p/mG2dDQrVKp/
lm eval:

m8k_250/dsv4pro/*.jsonl
local-completions ({'model': 'dsv4pro', 'base_url': 'http://127.0.0.1:8000/v1/completions', 'tokenizer_backend': 'none', 'tokenized_requests': False, 'max_length': 1048576, 'max_gen_toks': 512, 'num_concurrent': 4, 'timeout': 1200}), gen_kwargs: ({}), limit: 250.0, num_fewshot: None, batch_size: 4
|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value|   |Stderr|
|-----|------:|----------------|-----:|-----------|---|----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  | 0.94|±  |0.0151|
|     |       |strict-match    |     5|exact_match|↑  | 0.94|±  |0.0151

Test Result


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

@baonudesifeizhai

Copy link
Copy Markdown
Contributor Author

maybe break into several prs?

@mergify

mergify Bot commented May 27, 2026

Copy link
Copy Markdown
Contributor

Hi @baonudesifeizhai, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

Tip

Is mypy failing?
mypy is run differently in CI. If the failure is related to this check, please use the following command to run it locally:
# For mypy (substitute "3.10" with the failing version if needed)
pre-commit run --hook-stage manual mypy-3.10

@baonudesifeizhai

baonudesifeizhai commented May 28, 2026 •

Copy link
Copy Markdown
Contributor Author

https://paste.ubuntu.com/p/6hzxMGRfMH/ this is for turning on enable cache + pcp2 +maga moe (enable prefix cache support on continues pr did not submit yet)

https://paste.ubuntu.com/p/swyG56pXbQ/ baseline

image

@mergify

mergify Bot commented May 29, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @baonudesifeizhai.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify

mergify Bot commented Jun 4, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @baonudesifeizhai.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Jun 4, 2026
@baonudesifeizhai
baonudesifeizhai marked this pull request as draft June 4, 2026 18:49
@baonudesifeizhai
baonudesifeizhai marked this pull request as ready for review June 4, 2026 19:05

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot removed the needs-rebase label Jun 4, 2026
@mergify

mergify Bot commented Jun 5, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @baonudesifeizhai.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Jun 5, 2026
Signed-off-by: roG0d <baonudesifeizhai@gmail.com>
Signed-off-by: roG0d <baonudesifeizhai@gmail.com>
Signed-off-by: roG0d <baonudesifeizhai@gmail.com>
Signed-off-by: roG0d <baonudesifeizhai@gmail.com>
Signed-off-by: roG0d <baonudesifeizhai@gmail.com>
Signed-off-by: roG0d <baonudesifeizhai@gmail.com>
@mergify

mergify Bot commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @baonudesifeizhai.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant