Skip to content

[Misc][MRV2] Revert obsolete vLLM 0.28 PCP+DP adapters - #17174

Merged
realliujiaxu merged 1 commit into
vllm-project:mainfrom
wzx0726:codex/revert-16853-028-compat
Sep 23, 2026
Merged

realliujiaxu merged 1 commit into
vllm-project:mainfrom
wzx0726:codex/revert-16853-028-compat

Conversation

@wzx0726

@wzx0726 wzx0726 commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Remove the obsolete vLLM 0.28.0 PCP+DP compatibility introduced by #16853 (8f2e3fed73328148ea603ccfe5764542fa7cbbd2) after #17004 upgraded the supported release to v0.29.0. Both v0.29.0 and the paired vLLM main provide native PCP-local token counting before dispatch.

This is a conflict-resolved revert with targeted cleanup:

  • Remove the 0.28-only ParallelConfig validator shim, dispatch ContextVar/wrapper, the legacy dispatch block inside gather_batch_req_state, and PCP token-count override.

  • Preserve the independent v0.29.0 patch_parallel_config.py: the release still rejects PCP+DP without that patch.

  • Preserve [Bugfix][MRV2][P/D] Preserve decode graph with recompute scheduler #17128's PD decode recompute logic in gather_batch_req_state and its regression tests.

  • Preserve subsequent changes, including nullcontext and the v0.29.0 argument gates, generic PCP count/feature tests, and slot-buffer dtype coverage.

  • Remove obsolete adapter-specific tests and patch documentation. Remove imports made unused by the cleanup, including vllm_version_is in pcp_manager.py; that import predates [BugFix][MRV2] Adapt PCP+DP dispatch for vLLM 0.28.0 #16853 but has no remaining references after this revert and would trigger Ruff F401.

No new tests or unrelated formatting changes. The execute_model body is unindented only to remove its obsolete context wrapper.

Does this PR introduce any user-facing change?

No intended behavior change on the supported v0.29.0 and fixed-main lanes. The retired v0.28.0 compatibility is removed; v0.29.0 PCP+DP configuration support remains provided by the independent validator patch.

How was this patch tested?

On rebased commit d6b9f29e01ee2daa2d594b8c5c828e94dac0b128:

  • Ruff lint/format, AST syntax, conflict-marker and git diff --check checks passed for the seven changed Python files.
  • Preserve the new upstream step_eplb_after import, PD decode recompute logic, and o_proj TP graph checks. Range-diff against 5034869 shows only upstream context changes; the cleanup scope is unchanged.
  • DCO Signed-off-by is retained. Fresh CI is pending; no local Ascend runtime tests were run on this revision.
  • Previous revision 5034869 passed E2E run 35741354764, including pre-commit, CPU UT, selected NPU jobs, upstream tests and ci-gate. Skipped jobs are not passes; these results do not establish validation of this rebased revision.

Base: vllm-ascend main 972fcd5c974d55a7970ebe630329daa4b4049c49.
Release: vLLM v0.29.0 (98dff2a81d747d1dba01a47f939f48c3526d4206).

vLLM main: vllm-project/vllm@84030bb

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request performs a cleanup by removing compatibility code and workarounds specifically introduced for vLLM 0.28.0. With the upgrade to vLLM 0.29.0, these adapters are no longer necessary as the newer version provides native support for PCP+DP configurations. The changes focus on removing shims, context wrappers, and associated tests, while preserving necessary v0.29.0-specific patches.

Highlights

  • Removal of vLLM 0.28.0 compatibility: Removed the obsolete ParallelConfig validator shim, dispatch ContextVar/wrapper, and PCP token-count overrides that were required for vLLM 0.28.0.
  • Cleanup of unused code and documentation: Cleaned up unused imports, such as vllm_version_is, and removed documentation related to the retired 0.28.0 compatibility workarounds.
  • Test suite refinement: Removed obsolete adapter-specific tests and updated existing test files to reflect the removal of 0.28.0 support.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request removes obsolete vLLM 0.28.0 PCP+DP compatibility adapters and workarounds across several files, including model runner, PCP manager, and parallel configuration patches, as well as their associated unit tests, following the upgrade to v0.29.0. The review feedback correctly identifies that the PR title and summary should be updated to comply with the Repository Style Guide, providing a well-structured markdown template for the title and description.

Comment thread vllm_ascend/patch/__init__.py
@wzx0726
wzx0726 marked this pull request as ready for review September 22, 2026 06:40
realliujiaxu
realliujiaxu previously approved these changes Sep 22, 2026
@wzx0726 wzx0726 changed the title [Cleanup][MRV2] Revert obsolete vLLM 0.28 PCP+DP adapters [Misc][MRV2] Revert obsolete vLLM 0.28 PCP+DP adapters Sep 22, 2026
@wzx0726
wzx0726 force-pushed the codex/revert-16853-028-compat branch from 3227f2d to 990fdd5 Compare September 22, 2026 11:31
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Revert the obsolete compatibility portions of
8f2e3fe (vllm-project#16853).

Both supported vLLM pins provide PCP-local dispatch counting. Remove
the 0.28 validator shim, dispatch context and count override, and their
release-only tests and documentation. Keep the separate 0.29 validator,
generic PCP count/feature coverage, and later slot-buffer dtype tests.
Resolve conflicts preserving 0.29 argument gating and nullcontext.

Validation: targeted Ruff lint/format, AST syntax, unchanged-function
comparison and git diff --check passed. Ascend runtime UT and smoke
validation were not run.

Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
@wzx0726
wzx0726 force-pushed the codex/revert-16853-028-compat branch from 5034869 to d6b9f29 Compare September 23, 2026 06:30
@realliujiaxu
realliujiaxu enabled auto-merge (squash) September 23, 2026 07:43
@realliujiaxu
realliujiaxu merged commit dcdf99f into vllm-project:main Sep 23, 2026
31 checks passed
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…17174)

### What this PR does / why we need it?

Remove the obsolete vLLM 0.28.0 PCP+DP compatibility introduced by
vllm-project#16853 (`8f2e3fed73328148ea603ccfe5764542fa7cbbd2`) after vllm-project#17004
upgraded the supported release to v0.29.0. Both v0.29.0 and the paired
vLLM main provide native PCP-local token counting before dispatch.

This is a conflict-resolved revert with targeted cleanup:

- Remove the 0.28-only ParallelConfig validator shim, dispatch
ContextVar/wrapper, the legacy dispatch block inside
gather_batch_req_state, and PCP token-count override.
- Preserve the independent v0.29.0 patch_parallel_config.py: the release
still rejects PCP+DP without that patch.
- Preserve vllm-project#17128's PD decode recompute logic in gather_batch_req_state
and its regression tests.
- Preserve subsequent changes, including nullcontext and the v0.29.0
argument gates, generic PCP count/feature tests, and slot-buffer dtype
coverage.
- Remove obsolete adapter-specific tests and patch documentation. Remove
imports made unused by the cleanup, including vllm_version_is in
pcp_manager.py; that import predates vllm-project#16853 but has no remaining
references after this revert and would trigger Ruff F401.

No new tests or unrelated formatting changes. The execute_model body is
unindented only to remove its obsolete context wrapper.

### Does this PR introduce _any_ user-facing change?

No intended behavior change on the supported v0.29.0 and fixed-main
lanes. The retired v0.28.0 compatibility is removed; v0.29.0 PCP+DP
configuration support remains provided by the independent validator
patch.

### How was this patch tested?

On rebased commit `d6b9f29e01ee2daa2d594b8c5c828e94dac0b128`:

- Ruff lint/format, AST syntax, conflict-marker and git diff --check
checks passed for the seven changed Python files.
- Preserve the new upstream step_eplb_after import, PD decode recompute
logic, and o_proj TP graph checks. Range-diff against 5034869 shows
only upstream context changes; the cleanup scope is unchanged.
- DCO Signed-off-by is retained. Fresh CI is pending; no local Ascend
runtime tests were run on this revision.
- Previous revision 5034869 passed [E2E run
35741354764](https://github.com/vllm-project/vllm-ascend/actions/runs/35741354764),
including pre-commit, CPU UT, selected NPU jobs, upstream tests and
ci-gate. Skipped jobs are not passes; these results do not establish
validation of this rebased revision.

Base: vllm-ascend main `972fcd5c974d55a7970ebe630329daa4b4049c49`.
Release: vLLM v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`).
vLLM main:
vllm-project/vllm@84030bb



- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants