Skip to content

[Feat] Support PP with PCP in GPU Model Runner V2 - #59139

Merged
LucasWilkinson merged 3 commits into
vllm-project:mainfrom
pisceskkk:codex/pcp-pp-main
Sep 30, 2026
Merged

LucasWilkinson merged 3 commits into
vllm-project:mainfrom
pisceskkk:codex/pcp-pp-main

Conversation

@pisceskkk

@pisceskkk pisceskkk commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Enable pipeline parallelism (PP) together with Prefill Context Parallelism (PCP) in GPU Model Runner V2 for MLA models.

The last PP stage restores PCP-local hidden states to global request order before sampling. Earlier PP stages use the global request batch when receiving sampled tokens and updating request state, while model forward continues on the PCP-local batch.

Related to #25749 and #50853. Open PRs #33403, #49109, #49246, and #53948 address other PCP or PP paths; none covers this Model Runner V2 request mapping for PP feedback.

AI assistance was used to prepare the implementation, tests, and validation.

Acceptance criteria

  • PP2+PCP2 initializes and completes generation for an MLA model in Model Runner V2.
  • Sampled-token feedback on earlier PP stages updates the correct global requests. Hidden-state restoration occurs on the last PP stage before sampling.
  • Mixed-length requests, including chunked prefill, preserve generated tokens and logprob counts relative to PCP2 in the functional topology check.
  • A full real-model dataset evaluation in FULL_AND_PIECEWISE CUDA graph mode completes every request and reports accuracy alongside a PCP2 control.

Test Plan

.venv/bin/python -m pytest -q tests/v1/worker/test_gpu_model_runner_v2.py::test_non_last_pp_rank_uses_global_batch_for_sample_feedback

Run Ruff check and format on the three changed files. Compare PCP2 and PP2+PCP2 using a two-layer, randomly initialized DeepSeek-V3.2 FP8 model with one 607-token and one 7-token prompt, eight output tokens each, and prompt/output logprobs.

Evaluate DeepSeek-V2-Lite-Chat on all 1,319 GSM8K test questions with the same five-shot chat prompts, greedy decoding, and 256-token output limit in PCP2 and PP2+PCP2. Both runs use BF16, expert parallelism, FULL_AND_PIECEWISE CUDA graphs, and VLLM_MOE_SKIP_PADDING=0 on NVIDIA RTX 5090 GPUs.

Test Result

  • The focused test, Ruff check and format, and git diff --check passed on base aedaba8.

  • The DeepSeek-V3.2 functional check produced identical token IDs, text, and logprob counts in PCP2 and PP2+PCP2 for both requests.

  • DeepSeek-V2-Lite-Chat completed all 1,319 GSM8K questions in both topologies:

    Topology Correct Accuracy
    PCP2 873/1,319 66.19%
    PP2+PCP2 886/1,319 67.17%

    PP2+PCP2 scored 13 questions (0.99 percentage points) higher. Extracted answers matched on 1,070/1,319 questions; the cause of the cross-topology differences has not been isolated.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added the mrv2 Model Runner V2 specific label Sep 29, 2026
@pisceskkk pisceskkk changed the title Support PP with PCP in GPU Model Runner V2 [Feat] Support PP with PCP in GPU Model Runner V2 Sep 29, 2026
Assisted-by: OpenAI Codex

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>

@LucasWilkinson LucasWilkinson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM thanks!

@LucasWilkinson

Copy link
Copy Markdown
Contributor

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91872 for commit fa3f88cbc556.

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
@pisceskkk

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91903 for commit 893bb90a4f81, running 3 failed step(s) from Buildkite CI #91872.

@LucasWilkinson
LucasWilkinson merged commit 98b45c1 into vllm-project:main Sep 30, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

mrv2 Model Runner V2 specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants