Skip to content

[Core][PCP] Support data parallelism with MRV2 PCP - #49246

Draft
LucasWilkinson wants to merge 2 commits into
vllm-project:mainfrom
LucasWilkinson:codex/dp-pcp-support-pr
Draft

LucasWilkinson wants to merge 2 commits into
vllm-project:mainfrom
LucasWilkinson:codex/dp-pcp-support-pr

Conversation

@LucasWilkinson

Copy link
Copy Markdown
Contributor

Summary

  • allow MRV2 prefill context parallelism (PCP) with data parallelism (DP)
  • allocate each local DP engine's TP/PP/PCP device range correctly
  • rebuild DP token metadata after PCP partitions the scheduled batch
  • compose AgRs MoE dispatch across DP then PCP, reversing the order for combine
  • replicate sequence-parallel token sizes across PCP ranks in flattened EP order
  • mask empty MLA context chunks with -inf LSE before merging

Related to #46358.

Why this is not a duplicate

I checked issue #46358 and searched open PRs by issue number and DP/PCP/MLA
keywords. No open PR enables generic MRV2 DP+PCP composition. #49109 is scoped to
Mooncake KV-connector PCP replica selection, while #43809 is model-specific
DeepSeek-V4 hybrid-KV/MegaMoE PCP work. Neither changes the generic DP supervisor,
post-PCP DP metadata, or hierarchical AgRs path implemented here.

Validation

.venv/bin/python -m pytest \
  tests/test_config.py \
  tests/entrypoints/openai/test_dp_supervisor.py -q

Result: 178 passed.

Commit-time pre-commit hooks passed, including Ruff, formatting, mypy, SPDX,
configuration validation, and forbidden-import checks.

Runtime configuration (4 GPUs):

CUDA_VISIBLE_DEVICES=0,1,2,3 \
VLLM_USE_V2_MODEL_RUNNER=1 \
.venv/bin/vllm serve nvidia/GLM-5.2-NVFP4 \
  --enforce-eager \
  --max-model-len 4096 \
  --max-num-batched-tokens 32768 \
  --safetensors-load-strategy prefetch \
  --moe-backend flashinfer_cutlass \
  --data-parallel-size 2 \
  --prefill-context-parallel-size 2 \
  --enable-expert-parallel \
  --kv-cache-dtype fp8 \
  --trust-remote-code

Both DP ranks returned the expected answer (18) when addressed with
X-data-parallel-rank.

100-sample GSM8K smoke evaluation:

.venv/bin/python tests/evals/gsm8k/gsm8k_eval.py \
  --num-questions 100 \
  --num-shots 5 \
  --max-tokens 256 \
  --max-concurrency 100 \
  --port 8000
Deployment Accuracy Invalid responses
DP2+PCP2+EP4 94/100 0/100

The same rank-pinned runtime probe passed with VLLM_GPU_SYNC_CHECK=error, with
no synchronization detected during model execution.

AI assistance

AI assistance was used to diagnose and implement this draft. This PR should
remain a draft until the human submitter reviews every changed line, confirms
the test/evaluation results, and can explain and defend the change end-to-end.

LucasWilkinson and others added 2 commits July 20, 2026 21:53
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
@yiminghub2024

Copy link
Copy Markdown

this is very userful pr for prefill , but i want to know , if we can pcp not include in worldsize ? fg :
CUDA_VISIBLE_DEVICES=0,1,2,3
VLLM_USE_V2_MODEL_RUNNER=1
.venv/bin/vllm serve nvidia/GLM-5.2-NVFP4
--enforce-eager
--max-model-len 4096
--max-num-batched-tokens 32768
--safetensors-load-strategy prefetch
--moe-backend flashinfer_cutlass
--data-parallel-size 2
--prefill-context-parallel-size 2
--enable-expert-parallel
--kv-cache-dtype fp8
--trust-remote-code

world size not 4 is 2 , only compute tp/pp/dp or only use pcp if setting pcp ,setting tp/pp/dp to 1 ,so worldsize always 8 with pcp 8 @LucasWilkinson

refer:
#47846

@mergify

mergify Bot commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @LucasWilkinson.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants