Skip to content

[BugFix][MoE] Align hash routing input IDs for PCP - #16834

Merged
yiz-liu merged 1 commit into
vllm-project:mainfrom
li1how:bugfix/pcp-hash-router-input-ids
Sep 20, 2026
Merged

yiz-liu merged 1 commit into
vllm-project:mainfrom
li1how:bugfix/pcp-hash-router-input-ids

Conversation

@li1how

@li1how li1how commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

The AllGather MoE path gathers hidden states and router logits in DP-then-PCP order, while hash routing previously gathered input_ids only across DP. With PCP enabled, this leaves fewer token IDs than router-logit rows and causes hash routing to fail.

This PR:

  • gathers hash-routing input IDs with the same DP-then-PCP padding and communication layout used by the AllGather prepare path;
  • renames the helper to reflect that it covers every required parallel dimension;
  • extends the existing prepare/finalize and hash-router regression coverage for no parallelism, DP, PCP, and combined DP+PCP layouts.

Does this PR introduce any user-facing change?

No API or configuration change. Hash-based MoE routing now works with the existing PCP configuration.

How was this patch tested?

  • A5 PCP4 (DP1 × TP1 × PCP4): service started and ran normally. Full GPQA Diamond (enable_thinking=false) achieved 76.26% accuracy (151/198), with 198/198 requests successful and a 56.68% DSpark acceptance rate.

  • vLLM main: vllm-project/vllm@84030bb

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@li1how
li1how marked this pull request as ready for review September 18, 2026 02:45
Copilot AI lite review requested due to automatic review settings September 18, 2026 02:45

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The PCP gather path can fail for SP and duplicate collectives for SP+DP/PCP, preventing reliable ID alignment.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

This PR aligns hash-routing input_ids with AllGather MoE layouts across DP and PCP parallelism.

Changes:

  • Renames and extends the input-ID gathering helper.
  • Updates hash routing to use DP-then-PCP gathering.
  • Expands regression coverage across parallel layouts.
File summaries
File Summary
vllm_ascend/ops/fused_moe/router/fused_topk_router.py Uses the renamed input-ID gather helper.
vllm_ascend/ops/fused_moe/prepare_finalize.py Adds DP/PCP-aligned input-ID gathering; the SP+PCP path has a critical unresolved issue.
tests/ut/ops/test_prepare_finalize.py Tests input-ID alignment across parallel layouts.
tests/ut/ops/test_fused_moe.py Updates hash-router regression coverage.
Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread vllm_ascend/ops/fused_moe/prepare_finalize.py
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a bug in the MoE hash routing mechanism where input IDs were not correctly aligned with gathered router logits when using PCP (Parallel Column Partitioning). By synchronizing the communication layout for input IDs with the existing AllGather prepare path, the fix ensures that hash-based routing functions correctly across various parallel configurations.

Highlights

  • Hash Routing Alignment: Updated the hash routing input ID gathering process to use the same DP-then-PCP communication layout as the AllGather prepare path, ensuring consistent token ID counts.
  • Helper Method Refactoring: Renamed 'all_gather_input_id_with_dp_group' to 'all_gather_input_ids' to accurately reflect its role in handling both DP and PCP parallel dimensions.
  • Regression Coverage: Extended unit tests to include comprehensive coverage for no parallelism, DP, PCP, and combined DP+PCP configurations.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Support PCP gathering for input IDs in fused MoE

Suggested PR Summary:

### What this PR does / why we need it?
This pull request updates the fused MoE routing logic to support PCP (Pipeline Cell Parallelism) gathering for input IDs. Specifically, it renames `all_gather_input_id_with_dp_group` to `all_gather_input_ids` in `PrepareAndFinalizeWithAllGather` and adds support for gathering input IDs across the PCP group when `pcp_size > 1`.

However, there is a critical issue when sequence parallelism is enabled (`self._use_ep_sequence_parallel()` is `True`). In this case, `self.num_tokens_pcp` is undefined, and `self.num_tokens` is set to the post-gather token count, which will cause incorrect padding or an `AttributeError`. We should handle sequence parallelism by using `torch.ops.vllm.maybe_all_gather_and_maybe_unpad` to gather `input_ids` across the EP group.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
The changes are tested by updating existing unit tests in `tests/ut/ops/test_fused_moe.py` and expanding `test_allgather_prepare_finalize` in `tests/ut/ops/test_prepare_finalize.py` to cover various combinations of DP and PCP sizes.

Comment thread vllm_ascend/ops/fused_moe/prepare_finalize.py
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@li1how
li1how force-pushed the bugfix/pcp-hash-router-input-ids branch from 8add7b2 to 528573c Compare September 18, 2026 07:09
@nwpu-zxr nwpu-zxr added the ready-precise run selected e2e test for pr label Sep 18, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: leolee <yihao.li@huawei.com>
@li1how
li1how force-pushed the bugfix/pcp-hash-router-input-ids branch from 528573c to 4c7eca1 Compare September 19, 2026 06:06
@li1how

li1how commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@yiz-liu
yiz-liu merged commit 8038f64 into vllm-project:main Sep 20, 2026
53 of 55 checks passed
@li1how
li1how deleted the bugfix/pcp-hash-router-input-ids branch September 20, 2026 02:00
ZT-AIA pushed a commit that referenced this pull request Sep 20, 2026
)

### What this PR does / why we need it?

Revert #16393 at the author's request. This reverses squash commit
`af277b2681443089e4d29684b1101707e2eb6426` on upstream main
`8038f64129a1f8751ac6684f47c8d058208e4ff3`, restoring the previous
fixed-main + v0.28.0 support contract.

The revert applies without conflicts. Subsequent upstream patches are
retained, including #16775 (PCP/DCP slots), #16923 (A3 SFA prefill
all-to-all), and #16834 (PCP hash-routing input IDs). No unrelated fixes
are included.

| Reverted change | Reason and source | Lane |
| --- | --- | --- |
| Release marker | Reverse
[#16393](https://github.com/vllm-project/vllm-ascend/pull/16393/files):
v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`) to v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) | Release |
| Version branches and interface adapters across scheduler, KV cache,
attention, model runners and speculative decoding | Restore the
pre-#16393 v0.28.0 contracts by reversing the exact merged diff; keep
main implementations behind their original version selection | Main +
release |
| v0.29.0-only parallel-config patch, registration and documentation |
Remove the patch introduced by #16393; retain the independent upstream
v0.28.0 PCP+DP workaround | Release |
| Existing unit/e2e tests and release-only skips | Restore the tests and
conditions changed by #16393, without introducing additional skip scope
beyond the reverted baseline | Main + release |

The vLLM main pin remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. No
main interface scan is needed for this release-only revert. The only
changes beyond the inverse patch are formatting in
`tests/ut/test_utils.py` and line-ending normalization in
`categorical_sample.py` needed for formatting/whitespace checks.
Workflows are unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. The supported release returns to vLLM v0.28.0; v0.29.0 support and
its specific compatibility changes from #16393 are reverted. The
vllm-ascend package version and fixed vLLM main pin are unchanged.

### How was this patch tested?

- `git revert --no-commit af277b2`
applied cleanly on the latest fetched main.
- Verified that the complete post-#16393 upstream patch can still
reverse-apply to the staged result (`git apply --reverse --check
--cached`), confirming subsequent changes are preserved.
- `git diff --cached --check`, Python AST parsing, Ruff lint and
formatting checks passed for all 94 remaining changed Python files.
- Ran the equivalent of `format.sh ci`: `pre-commit run --all-files
--hook-stage manual`. Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-based hooks could not execute on
this Windows host; two Python launcher hooks exited 9009. Full
pre-commit validation remains for CI.
- Actual CPU/NPU tests have not been run locally because the required
vLLM/NPU runtime is unavailable. New PR CI is pending; the original
upgrade PR's green run is not validation of this revert.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
ningjingbengxiaohai pushed a commit that referenced this pull request Sep 20, 2026
…#17004)

### What this PR does / why we need it?



Reintroduce the vLLM v0.29.0 release upgrade from #16393, reverted by
#16949, with container defaults aligned to the supported release. Based
on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves
later merged changes.


The image/source mismatch is confirmed in [nightly job
106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020):
the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded
Ascend failed importing `_get_packed_kv_cache_groups`. All eight root
Dockerfiles still defaulted to v0.28.0. The reusable image workflow
passes no `VLLM_TAG` override, so those defaults govern release builds.
Changing only the release marker does not update those images.


- Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

- Release changes from v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0
(`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the
remote tag).
- Restore #16393's source compatibility and existing UT changes by
reversing #16949, then reconcile current upstream changes. No main
interface scan is rerun for this release-only upgrade.
- Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P,
Ubuntu/openEuler. Preserve existing exact-commit build overrides. No
workflow changes.


| Current-base adaptation | Exact cause and evidence | Lane | Validation
|
| --- | --- | --- | --- |

| Eight Dockerfile `VLLM_TAG` defaults and release marker | #16393
changed the release contract but omitted image defaults; the linked
nightly log proves installation of 0.28.0 and missing
`_get_packed_kv_cache_groups`. | Release images | All eight defaults
match marker; image-build CI requested, pending |
| `worker/v2/attn_utils.py::get_kv_cache_spec` and existing
`_make_mla_layer` UT fixture | Ascend
[#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files),
`df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field
selector. vLLM
[#51718](https://github.com/vllm-project/vllm/pull/51718/files),
`8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio`
with `tokens_per_state`; both exact supported pins use the latter. Use
the common field, retaining metadata and cache-view assertions. | Both |
Source inspection and static checks passed; actual CI pending |
| `attention/attention_v1.py` import conflict | Preserve
`attention_transfer_window` from Ascend
[#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files),
`5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from
#16798; keep the common relocated PCP import from #16393. Do not restore
the old compute-start import or unused weak-reference import. | Both |
Conflict resolved; static checks passed |
| `worker/v2/aclgraph_utils.py` import conflict | Preserve
`ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend
[#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files),
`34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete
0.28 version selector restored by the revert. | Both | Conflict
resolved; static checks passed |


| Existing 310P and Mamba model-runner UT imports | Preserve
hardware-profile imports and mocks from Ascend
[#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files),
`d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused
old-release selector import. Hardware capability routing remains
unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format
passed; fresh CI pending |

Newly merged changes were reviewed for version-contract impact:
#16775/#16834 PCP metadata and routing, #16923 A3 SFA,
#16426/#16924/#16955 custom ops, #16081 MTP/SP, #16669 xlite,
#15636/#16747 transfer, #16913 A5 pages, #16798 graph updates, #16755
GLM PD, #16952 C8 config, #16673 operator removal and #16320 Kimi-K3 KV
pool. Preserve these changes; no additional source-proven version branch
was identified beyond the entries above. In particular, #16747's
`UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size`
exist in both exact pins; #16320 adds Ascend connector hooks. CI remains
necessary to validate runtime interactions. CI/documentation-only PRs
are retained unchanged.


<details>

<summary>Inherited per-file compatibility evidence from #16393</summary>


The following source-contract ledger is inherited from #16393. Any
historical verification wording refers only to that earlier PR; it does
not certify this new head. New-head validation is listed below.


| Reference | Commit |





|---|---|





| vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` |
| PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` |





| Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` |





| Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a`
|
| Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` |






- Use common implementations where v0.29.0 and the fixed main share
KV-cache layouts, Mamba copy/group APIs, PCP handling and
speculative-decoding contracts.
- Retain explicit `vllm_version_is("0.29.0")` branches for contracts
that still differ, including RoPE, scheduler block snapshots,
InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.
- Remove obsolete v0.28.0 compatibility and adapt existing test
fixtures. Version detection uses package versions and the explicit
`VLLM_VERSION` override, with local-version suffix handling and UT
environment isolation retained; no hard-coded release-SHA inference
remains.
- Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving
other validations and EPLB platform binding. Rebuild dependent Pydantic
schemas so nested configuration validation uses the patched validator.
The global patch documentation records its rationale and removal
criteria.






- Preserve #15747's Spec+PP protocol/partition handling after rebase;
use the common exact-release selector and the real function-local DSpark
sharing import.






This release-only upgrade does not require a new main old-to-new
interface scan. The latest rebase incorporates
[#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files),
which reverted #16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The
now-unnecessary V4.1 drafter import gate and tuple annotation have been
removed. Other release adaptations, including #15747 Spec+PP handling,
remain. Detailed contract evidence is retained below for review.






<details>





<summary>Per-file adaptations and exact upstream evidence</summary>






#### Per-file adaptation ledger











Evidence IDs refer to the exact source contract and upstream diff table
below. Every row is syntax checked; branch-normalized AST comparison
confirms unchanged main function bodies except the version identity
helper and the explicitly retired propose argument/type annotations.
Both supported versions completed CPU and NPU execution as recorded
below.






| Ascend file / symbols | Disposition and upstream evidence |
Verification |
|---|---|---|





| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` /
`_load_dspark_model_with_target_quant`;
`tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased
[Ascend
#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a)
(`82df9d871`), including manual PP partition masking. vLLM [#52809
diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share`
inside the loader on both supported pins, so retain the earlier release
fix: patch `eagle_utils`, never read/patch a nonexistent
`dspark_utils._should_share`. v0.29 keeps its global PP guard; main has
native PP via
[#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks
the real function-local import, absent module alias, and restoration on
success/failure for both lanes; partition and PP assertions retained.
Ruff/syntax pass; actual CPU/NPU pending. |
| `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`;
`tests/ut/worker/v2/test_pp_utils.py` | #15747 added broad 0.28/0.29
routing and an obsolete 0.28 dev-build recognition path. For the two
supported pins, #50514 exists only on fixed main. Route through
`vllm_version_is("0.29.0")`; retain package/local-suffix and explicit
environment-override semantics. No release-SHA inference or third
release lane. | Existing UTs exercise the real uncached version helper
with monkeypatch isolation, local suffix, fixed-main dev string,
explicit override, and non-target versions. Isolated routing checks and
Ruff/syntax pass; full CPU UT pending. |
| `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`,
`_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass`
| #56078; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`,
`_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436;
use shared standardized layouts and retain the v0.29.0 InputBatch gate.
The rebase preserves vllm-ascend #16043's MTP copy tracking while
removing only legacy v0.28.0 allocation paths. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` |
#56078; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` |
#52839; both lanes use the common PCP import. Rebase keeps current
upstream graph code and removes only the v0.28.0 import branch. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module
imports/dispatch` | #52839; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch`
| #51358, #54853; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/core/recompute_scheduler.py`<br>`module
imports/dispatch`, `schedule` | #51358, #54853; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`,
`AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker`
| #52615; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__`
| #53614; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module
imports/dispatch`, `_get_max_layers_per_page_size`,
`_ascend_max_memory_usage_bytes_from_groups`,
`_ascend_get_kv_cache_config_from_groups` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module
imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` |
v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it
after KV binding. Evidence: [#52506,
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils
diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c).
Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. |
Failure reproduced on v0.29.0; Historical release/main NPU validation
passed; current-head CI pending |
| `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module
imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module
imports/dispatch` | #51718; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant`
| #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports
`_should_share` locally from Eagle utilities; fixed main removes the PP
guard. Keep the release `get_pp_group` patch, share through the common
Eagle utility, and delete the obsolete v0.28.0
`dspark_utils._should_share` patch. | Exact release failure reproduced;
Historical release/main CPU/NPU validation passed; current-head CI
pending |
|
`vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`,
`__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`,
`register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`,
`_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`,
`_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers |
#51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and
import the NaN helpers directly because v0.29.0 and fixed main expose
the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static
checked; historical dual-version CPU/NPU validation passed; current-head
CI pending |
| `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes
`pcp_manager` common to both lanes. The rebase preserves vllm-ascend
#16409's host-parameter-update revert and removes only the obsolete
v0.28.0 capture branch. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`,
`_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/worker/v2/block_table.py`<br>`__init__`,
`init_block_table_layout_tensors`, `compute_slot_mappings` | #51718,
#51031; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`,
`prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`,
`execute_model` | #50514, #54436, #52506, #55212, #53515; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch`
| #54282; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` |
#49811; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
|
`vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module
imports/dispatch`, `propose` | #53694, #52188; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`,
`wake_up` | #51718, #53508; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |






#### Exact upstream evidence











| Upstream change | Full commit SHA / direct diff | Actual supported
contracts and branch decision |
|---|---|---|





| #51718 |
`8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Both use layers/layer_stride/block_stride/offset, tokens_per_state,
CircularBufferSpec and standardized backing; retire
shared_by/compress_ratio allocation branches. |
| #52839 |
`58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8)
| Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
`e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec)
| Both use Mamba copy-function dictionaries and unwrap
UniformTypeKVCacheSpecs. |
| #53106 |
`1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21)
| Both use WeightsMapper instead of AutoWeightsLoader
skip_prefixes/skip_substrs. |
| #53906 |
`98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Only pinned main has the optional MLA storage_block_size dataclass
field; release keeps the Ascend derived property, using
tokens_per_state. |
| #56078 |
`719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417)
| Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned
main uses mrope_num_dims and unified RoPE. |
| #52615 |
`138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045)
| Release uses num_blocks/kv_bytes_per_block; main uses
num_chunks/kv_bytes_per_chunk. |
| #51358 |
`6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release now has boundary_state_offloads and KVConnectorBlockState;
remove partial_tail_offloads plumbing. |
| #54853 |
`0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release constructor takes block_ids snapshots; main takes req_ids and
resolve_block_ids. Keep exact release snapshot membership and main
lazy-resolution membership. |
| #53614 |
`144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725)
| Only main configures drop_eagle_checkpoint_block for replay-aligned
Mamba checkpoints. |
| #50514 |
`d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Release retains module-level PP/share symbols and Ascend PP
workaround; main has the subsequent PP integration. |
| #52809 |
`91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark
module binding to a function-local import from Eagle utilities. The
shared Eagle patch remains effective; the old DSpark-module read/write
must be removed. |
| #54436 |
`6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902)
| Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Only main accepts valid_dummy_state_slots/valid_state_slots capture
arguments. |
| #55212 |
`83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Release prepares DCP local sequence lengths before partitioning; main
initializes DCP metadata afterwards. |
| #53515 |
`b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127)
| Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
`b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212)
| Both accept pcp_manager during graph capture. |
| #51031 |
`0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1)
| Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
`fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8)
| Both gumbel sampling APIs include is_drafting. |
| #52188 |
`d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0)
| Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
`5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6)
| Both propose APIs take DPSyncState; remove the obsolete token-count
argument and retain replicated-PCP synchronization. |
| #49811 |
`01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e)
| Both support extract_hidden_states on MRV2; remove old unsupported
dispatch/skip. |
| #53508 |
`479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29)
| Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
`3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1)
| Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain
release exclusion. |
| #52861 |
`b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28)
| Both include DeepseekV32MTPModel in the two-hidden-state architecture
set. |
| #54713 |
`b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3)
| Only main takes replay_boundaries in compressed-prefix hit lookup;
preserve release calls without that keyword. |
| #42785 |
`442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18)
| Only main capture callers pass axis_keys; preserve the existing Ascend
rejection of nonempty axes. |
| #52358 |
`8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
`9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4)
| Release already has mamba_has_prefill_checkpoint_blocks; later main
also has fine-grained prefix-cache state. |
| #51251 |
`7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| v0.29.0 and main expose ec_manager_config; retire the old release-only
ScoreEncoder configuration skip. |
| #53240 |
`b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c)
| v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache
groups; retire the old release-only replay skip. |
| #53853 |
`e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
`4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both expose the V1 unsupported-feature helper used by the existing
Ascend MRV1 feature filter. |






#### Additional inherited contracts











| vllm-ascend change | Why it is required | Upstream cause and direct
link | Lane | Verification |
|---|---|---|---|---|





| `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM
v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec,
list[int]]` from `get_mamba_groups` and both initialize `recoverssm`;
keeping the old fallback would preserve an unsupported third contract |
[v0.29.0 mamba
groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704),
[fixed-main mamba
groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704),
[v0.29.0
RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103),
[fixed-main
RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103)
| both, common implementation | Existing constructor UT now asserts the
parent-created RecoverSSM value is retained; Ruff and compileall pass |
| `worker/v2/model_runner.py`: always forward
`kv_cache_allocation_context` | v0.29.0 and fixed main both accept this
keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0
signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538),
[fixed-main
signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566),
[vllm-ascend
#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) |
both, common implementation | Existing UT continues to assert the exact
context object reaches the parent; Ruff and compileall pass |
| `_310p/worker/v2/model_runner.py`: select the release KV-zeroing
contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes
`KVCacheGroupSpec.is_eagle_group` but lacks
`SpeculativeConfig.use_eagle_block_drop`; fixed main added the method |
vLLM [#53388
diff](https://github.com/vllm-project/vllm/pull/53388/files), commit
[`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a);
[v0.29.0 group
field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200),
[fixed-main
helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876)
| release differs from main | Existing two-path KV-zeroing UT retained
and renamed for v0.29.0; Ruff and compileall pass |
| existing DFlash kernel UT: remove v0.28-only kwarg omission | the
current Ascend kernel accepts the CP arguments and the only excluded
lane was v0.28.0, which this PR replaces | [vllm-ascend
#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) |
both, common invocation | Existing NPU test remains enabled with all
assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |





|---|---|---|---|





| `patch/platform/patch_parallel_config.py`, registration, and global
patch documentation | Allow Ascend PCP+DP by removing the generic GPU
restriction, following vLLM
[#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit
`7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic
`parallel_config.current_platform` lookup and rebuild ParallelConfig →
SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing
configuration/EPLB UTs and historical PCP+DP NPU execution passed;
current-head CI pending. |
| `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT |
Both supported contracts require `is_drafting`, from
[#54282](https://github.com/vllm-project/vllm/pull/54282/files),
`fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited
optimized kernel and use a common wrapper. | Both | Existing positive
drafting assertion retained; current-head CI pending. |
| `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache
UT | Both support `cache_hit_alignment_tokens`, introduced by
[#53598](https://github.com/vllm-project/vllm/pull/53598/files),
`2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only
write-mask branch. | Both | Existing assertions retained; current-head
CI pending. |






#### Rebase and retired-fallback evidence











| Change | Exact source evidence | Decision |





|---|---|---|





| `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM
[#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95),
`d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL
before v0.29.0; both exact supported sources lack it. The old
conditional came from vllm-ascend
[#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0
selector and use the existing exclusion for both supported lanes. This
does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy
skip. |
| `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and
`tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM
[#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29),
`12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export
`nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The
v0.28.0 fallback originated in vllm-ascend
[#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete
import-failure/`None` fallbacks and the existing UT's obsolete
availability skip; assertions remain unchanged and execute on both
lanes. |
| `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend
[#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3),
`799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream
revert while resolving the real rebase conflict; do not reintroduce the
reverted graph-update behavior. |
| `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend
[#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2),
`d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged
310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation
branch. |






</details>













</details>



### Does this PR introduce _any_ user-facing change?



Yes. The supported vLLM release and default container builds move to
v0.29.0, while the fixed main remains supported. The vllm-ascend package
version does not change. Existing 0.29-only PCP+DP compatibility is
restored.


### How was this patch tested?

- Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`:
[run
35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654)
succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected
NPU jobs confirm the requested Ascend head, integration base
`2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM
pins/installations: fixed main
`84030bbe3d74d99bad477a3d2e37a973ccd8865c` /
`0.1.dev1+g84030bbe3.empty`, release
`98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane
totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest
invocations. Skips/xfails are not passes. The resulting rebased
integration commit is not printed and is not inferred. Actual release
CPU and failed/cancelled image variants remain gaps.


- Current-head CI update (2026-09-20): [main CPU
UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955)
passed **5157 tests**, with **67 skipped**; actual installed vLLM was
`0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head;
the log does not print the full resulting integration head/base.
Pre-commit and mypy passed. NPU/E2E results are recorded above; actual
release CPU remains unverified.
- Image build is partially blocked by infrastructure: A5 amd64
[Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005)
and
[openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068)
failed before reading the Dockerfile because BuildKit could not create a
snapshot temporary directory (`no space left on device`). No release
compatibility code change is justified by this failure; cancelled
variants remain unverified. The successful [310P openEuler arm64
build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987)
explicitly checked out release
`98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`;
this is build evidence, not runtime UT coverage.


- Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto
the baseline above. Previous-head CPU [job
106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458)
stopped during Ascend integration rebase after #16803 merged, before any
UT ran. Its vLLM checkout was the fixed main and installation reported
`0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved;
fresh-head CI was retriggered by the push.


- Remote official v0.29.0 tag resolved to the exact SHA above; both
source pins inspected.
- All 95 changed Python files pass Ruff lint, Ruff format and AST
parsing; `git diff --check` passes. Dockerfile tag/marker consistency
checked across all eight variants.
- Full pre-commit invocation: Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-dependent hooks cannot run on this
Windows host; Python launcher hooks exit 9009. Full CI lint remains
pending.
- Actual main/release CPU/NPU and image builds are not claimed locally:
the required Linux/NPU/container environment is unavailable. New PR E2E
and image-build CI are requested. Existing CPU workflow runs fixed main
only, so actual release CPU remains a validation gap.
- #16393's historical green CI is not a substitute for this new head. No
new test functions, workflow changes, golden/threshold changes or
additional skips are introduced beyond restoring #16393. Inherited
0.28-only tests/skips are not counted as passes.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
### What this PR does / why we need it?

The AllGather MoE path gathers hidden states and router logits in
DP-then-PCP order, while hash routing previously gathered `input_ids`
only across DP. With PCP enabled, this leaves fewer token IDs than
router-logit rows and causes hash routing to fail.

This PR:

- gathers hash-routing input IDs with the same DP-then-PCP padding and
communication layout used by the AllGather prepare path;
- renames the helper to reflect that it covers every required parallel
dimension;
- extends the existing prepare/finalize and hash-router regression
coverage for no parallelism, DP, PCP, and combined DP+PCP layouts.

### Does this PR introduce _any_ user-facing change?

No API or configuration change. Hash-based MoE routing now works with
the existing PCP configuration.

### How was this patch tested?

- A5 PCP4 (`DP1 × TP1 × PCP4`): service started and ran normally. Full
GPQA Diamond (`enable_thinking=false`) achieved 76.26% accuracy
(151/198), with 198/198 requests successful and a 56.68% DSpark
acceptance rate.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: leolee <yihao.li@huawei.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
…#16393) (vllm-project#16949)

### What this PR does / why we need it?

Revert vllm-project#16393 at the author's request. This reverses squash commit
`af277b2681443089e4d29684b1101707e2eb6426` on upstream main
`8038f64129a1f8751ac6684f47c8d058208e4ff3`, restoring the previous
fixed-main + v0.28.0 support contract.

The revert applies without conflicts. Subsequent upstream patches are
retained, including vllm-project#16775 (PCP/DCP slots), vllm-project#16923 (A3 SFA prefill
all-to-all), and vllm-project#16834 (PCP hash-routing input IDs). No unrelated fixes
are included.

| Reverted change | Reason and source | Lane |
| --- | --- | --- |
| Release marker | Reverse
[vllm-project#16393](https://github.com/vllm-project/vllm-ascend/pull/16393/files):
v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`) to v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) | Release |
| Version branches and interface adapters across scheduler, KV cache,
attention, model runners and speculative decoding | Restore the
pre-vllm-project#16393 v0.28.0 contracts by reversing the exact merged diff; keep
main implementations behind their original version selection | Main +
release |
| v0.29.0-only parallel-config patch, registration and documentation |
Remove the patch introduced by vllm-project#16393; retain the independent upstream
v0.28.0 PCP+DP workaround | Release |
| Existing unit/e2e tests and release-only skips | Restore the tests and
conditions changed by vllm-project#16393, without introducing additional skip scope
beyond the reverted baseline | Main + release |

The vLLM main pin remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. No
main interface scan is needed for this release-only revert. The only
changes beyond the inverse patch are formatting in
`tests/ut/test_utils.py` and line-ending normalization in
`categorical_sample.py` needed for formatting/whitespace checks.
Workflows are unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. The supported release returns to vLLM v0.28.0; v0.29.0 support and
its specific compatibility changes from vllm-project#16393 are reverted. The
vllm-ascend package version and fixed vLLM main pin are unchanged.

### How was this patch tested?

- `git revert --no-commit af277b2`
applied cleanly on the latest fetched main.
- Verified that the complete post-vllm-project#16393 upstream patch can still
reverse-apply to the staged result (`git apply --reverse --check
--cached`), confirming subsequent changes are preserved.
- `git diff --cached --check`, Python AST parsing, Ruff lint and
formatting checks passed for all 94 remaining changed Python files.
- Ran the equivalent of `format.sh ci`: `pre-commit run --all-files
--hook-stage manual`. Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-based hooks could not execute on
this Windows host; two Python launcher hooks exited 9009. Full
pre-commit validation remains for CI.
- Actual CPU/NPU tests have not been run locally because the required
vLLM/NPU runtime is unavailable. New PR CI is pending; the original
upgrade PR's green run is not validation of this revert.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
…vllm-project#17004)

### What this PR does / why we need it?



Reintroduce the vLLM v0.29.0 release upgrade from vllm-project#16393, reverted by
vllm-project#16949, with container defaults aligned to the supported release. Based
on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves
later merged changes.


The image/source mismatch is confirmed in [nightly job
106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020):
the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded
Ascend failed importing `_get_packed_kv_cache_groups`. All eight root
Dockerfiles still defaulted to v0.28.0. The reusable image workflow
passes no `VLLM_TAG` override, so those defaults govern release builds.
Changing only the release marker does not update those images.


- Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

- Release changes from v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0
(`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the
remote tag).
- Restore vllm-project#16393's source compatibility and existing UT changes by
reversing vllm-project#16949, then reconcile current upstream changes. No main
interface scan is rerun for this release-only upgrade.
- Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P,
Ubuntu/openEuler. Preserve existing exact-commit build overrides. No
workflow changes.


| Current-base adaptation | Exact cause and evidence | Lane | Validation
|
| --- | --- | --- | --- |

| Eight Dockerfile `VLLM_TAG` defaults and release marker | vllm-project#16393
changed the release contract but omitted image defaults; the linked
nightly log proves installation of 0.28.0 and missing
`_get_packed_kv_cache_groups`. | Release images | All eight defaults
match marker; image-build CI requested, pending |
| `worker/v2/attn_utils.py::get_kv_cache_spec` and existing
`_make_mla_layer` UT fixture | Ascend
[vllm-project#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files),
`df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field
selector. vLLM
[#51718](https://github.com/vllm-project/vllm/pull/51718/files),
`8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio`
with `tokens_per_state`; both exact supported pins use the latter. Use
the common field, retaining metadata and cache-view assertions. | Both |
Source inspection and static checks passed; actual CI pending |
| `attention/attention_v1.py` import conflict | Preserve
`attention_transfer_window` from Ascend
[vllm-project#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files),
`5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from
vllm-project#16798; keep the common relocated PCP import from vllm-project#16393. Do not restore
the old compute-start import or unused weak-reference import. | Both |
Conflict resolved; static checks passed |
| `worker/v2/aclgraph_utils.py` import conflict | Preserve
`ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend
[vllm-project#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files),
`34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete
0.28 version selector restored by the revert. | Both | Conflict
resolved; static checks passed |


| Existing 310P and Mamba model-runner UT imports | Preserve
hardware-profile imports and mocks from Ascend
[vllm-project#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files),
`d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused
old-release selector import. Hardware capability routing remains
unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format
passed; fresh CI pending |

Newly merged changes were reviewed for version-contract impact:
vllm-project#16775/vllm-project#16834 PCP metadata and routing, vllm-project#16923 A3 SFA,
vllm-project#16426/vllm-project#16924/vllm-project#16955 custom ops, vllm-project#16081 MTP/SP, vllm-project#16669 xlite,
vllm-project#15636/vllm-project#16747 transfer, vllm-project#16913 A5 pages, vllm-project#16798 graph updates, vllm-project#16755
GLM PD, vllm-project#16952 C8 config, vllm-project#16673 operator removal and vllm-project#16320 Kimi-K3 KV
pool. Preserve these changes; no additional source-proven version branch
was identified beyond the entries above. In particular, vllm-project#16747's
`UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size`
exist in both exact pins; vllm-project#16320 adds Ascend connector hooks. CI remains
necessary to validate runtime interactions. CI/documentation-only PRs
are retained unchanged.


<details>

<summary>Inherited per-file compatibility evidence from vllm-project#16393</summary>


The following source-contract ledger is inherited from vllm-project#16393. Any
historical verification wording refers only to that earlier PR; it does
not certify this new head. New-head validation is listed below.


| Reference | Commit |





|---|---|





| vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` |
| PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` |





| Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` |





| Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a`
|
| Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` |






- Use common implementations where v0.29.0 and the fixed main share
KV-cache layouts, Mamba copy/group APIs, PCP handling and
speculative-decoding contracts.
- Retain explicit `vllm_version_is("0.29.0")` branches for contracts
that still differ, including RoPE, scheduler block snapshots,
InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.
- Remove obsolete v0.28.0 compatibility and adapt existing test
fixtures. Version detection uses package versions and the explicit
`VLLM_VERSION` override, with local-version suffix handling and UT
environment isolation retained; no hard-coded release-SHA inference
remains.
- Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving
other validations and EPLB platform binding. Rebuild dependent Pydantic
schemas so nested configuration validation uses the patched validator.
The global patch documentation records its rationale and removal
criteria.






- Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase;
use the common exact-release selector and the real function-local DSpark
sharing import.






This release-only upgrade does not require a new main old-to-new
interface scan. The latest rebase incorporates
[vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files),
which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The
now-unnecessary V4.1 drafter import gate and tuple annotation have been
removed. Other release adaptations, including vllm-project#15747 Spec+PP handling,
remain. Detailed contract evidence is retained below for review.






<details>





<summary>Per-file adaptations and exact upstream evidence</summary>






#### Per-file adaptation ledger











Evidence IDs refer to the exact source contract and upstream diff table
below. Every row is syntax checked; branch-normalized AST comparison
confirms unchanged main function bodies except the version identity
helper and the explicitly retired propose argument/type annotations.
Both supported versions completed CPU and NPU execution as recorded
below.






| Ascend file / symbols | Disposition and upstream evidence |
Verification |
|---|---|---|





| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` /
`_load_dspark_model_with_target_quant`;
`tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased
[Ascend
vllm-project#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a)
(`82df9d871`), including manual PP partition masking. vLLM [#52809
diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share`
inside the loader on both supported pins, so retain the earlier release
fix: patch `eagle_utils`, never read/patch a nonexistent
`dspark_utils._should_share`. v0.29 keeps its global PP guard; main has
native PP via
[#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks
the real function-local import, absent module alias, and restoration on
success/failure for both lanes; partition and PP assertions retained.
Ruff/syntax pass; actual CPU/NPU pending. |
| `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`;
`tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29
routing and an obsolete 0.28 dev-build recognition path. For the two
supported pins, #50514 exists only on fixed main. Route through
`vllm_version_is("0.29.0")`; retain package/local-suffix and explicit
environment-override semantics. No release-SHA inference or third
release lane. | Existing UTs exercise the real uncached version helper
with monkeypatch isolation, local suffix, fixed-main dev string,
explicit override, and non-target versions. Isolated routing checks and
Ruff/syntax pass; full CPU UT pending. |
| `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`,
`_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass`
| #56078; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`,
`_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436;
use shared standardized layouts and retain the v0.29.0 InputBatch gate.
The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while
removing only legacy v0.28.0 allocation paths. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` |
#56078; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` |
#52839; both lanes use the common PCP import. Rebase keeps current
upstream graph code and removes only the v0.28.0 import branch. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module
imports/dispatch` | #52839; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch`
| #51358, #54853; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/core/recompute_scheduler.py`<br>`module
imports/dispatch`, `schedule` | #51358, #54853; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`,
`AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker`
| #52615; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__`
| #53614; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module
imports/dispatch`, `_get_max_layers_per_page_size`,
`_ascend_max_memory_usage_bytes_from_groups`,
`_ascend_get_kv_cache_config_from_groups` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module
imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` |
v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it
after KV binding. Evidence: [#52506,
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils
diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c).
Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. |
Failure reproduced on v0.29.0; Historical release/main NPU validation
passed; current-head CI pending |
| `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module
imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module
imports/dispatch` | #51718; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant`
| #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports
`_should_share` locally from Eagle utilities; fixed main removes the PP
guard. Keep the release `get_pp_group` patch, share through the common
Eagle utility, and delete the obsolete v0.28.0
`dspark_utils._should_share` patch. | Exact release failure reproduced;
Historical release/main CPU/NPU validation passed; current-head CI
pending |
|
`vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`,
`__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`,
`register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`,
`_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`,
`_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers |
#51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and
import the NaN helpers directly because v0.29.0 and fixed main expose
the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static
checked; historical dual-version CPU/NPU validation passed; current-head
CI pending |
| `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes
`pcp_manager` common to both lanes. The rebase preserves vllm-ascend
vllm-project#16409's host-parameter-update revert and removes only the obsolete
v0.28.0 capture branch. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`,
`_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/worker/v2/block_table.py`<br>`__init__`,
`init_block_table_layout_tensors`, `compute_slot_mappings` | #51718,
#51031; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`,
`prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`,
`execute_model` | #50514, #54436, #52506, #55212, #53515; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch`
| #54282; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` |
#49811; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
|
`vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module
imports/dispatch`, `propose` | #53694, #52188; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`,
`wake_up` | #51718, #53508; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |






#### Exact upstream evidence











| Upstream change | Full commit SHA / direct diff | Actual supported
contracts and branch decision |
|---|---|---|





| #51718 |
`8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Both use layers/layer_stride/block_stride/offset, tokens_per_state,
CircularBufferSpec and standardized backing; retire
shared_by/compress_ratio allocation branches. |
| #52839 |
`58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8)
| Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
`e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec)
| Both use Mamba copy-function dictionaries and unwrap
UniformTypeKVCacheSpecs. |
| #53106 |
`1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21)
| Both use WeightsMapper instead of AutoWeightsLoader
skip_prefixes/skip_substrs. |
| #53906 |
`98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Only pinned main has the optional MLA storage_block_size dataclass
field; release keeps the Ascend derived property, using
tokens_per_state. |
| #56078 |
`719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417)
| Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned
main uses mrope_num_dims and unified RoPE. |
| #52615 |
`138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045)
| Release uses num_blocks/kv_bytes_per_block; main uses
num_chunks/kv_bytes_per_chunk. |
| #51358 |
`6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release now has boundary_state_offloads and KVConnectorBlockState;
remove partial_tail_offloads plumbing. |
| #54853 |
`0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release constructor takes block_ids snapshots; main takes req_ids and
resolve_block_ids. Keep exact release snapshot membership and main
lazy-resolution membership. |
| #53614 |
`144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725)
| Only main configures drop_eagle_checkpoint_block for replay-aligned
Mamba checkpoints. |
| #50514 |
`d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Release retains module-level PP/share symbols and Ascend PP
workaround; main has the subsequent PP integration. |
| #52809 |
`91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark
module binding to a function-local import from Eagle utilities. The
shared Eagle patch remains effective; the old DSpark-module read/write
must be removed. |
| #54436 |
`6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902)
| Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Only main accepts valid_dummy_state_slots/valid_state_slots capture
arguments. |
| #55212 |
`83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Release prepares DCP local sequence lengths before partitioning; main
initializes DCP metadata afterwards. |
| #53515 |
`b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127)
| Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
`b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212)
| Both accept pcp_manager during graph capture. |
| #51031 |
`0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1)
| Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
`fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8)
| Both gumbel sampling APIs include is_drafting. |
| #52188 |
`d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0)
| Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
`5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6)
| Both propose APIs take DPSyncState; remove the obsolete token-count
argument and retain replicated-PCP synchronization. |
| #49811 |
`01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e)
| Both support extract_hidden_states on MRV2; remove old unsupported
dispatch/skip. |
| #53508 |
`479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29)
| Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
`3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1)
| Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain
release exclusion. |
| #52861 |
`b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28)
| Both include DeepseekV32MTPModel in the two-hidden-state architecture
set. |
| #54713 |
`b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3)
| Only main takes replay_boundaries in compressed-prefix hit lookup;
preserve release calls without that keyword. |
| #42785 |
`442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18)
| Only main capture callers pass axis_keys; preserve the existing Ascend
rejection of nonempty axes. |
| #52358 |
`8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
`9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4)
| Release already has mamba_has_prefill_checkpoint_blocks; later main
also has fine-grained prefix-cache state. |
| #51251 |
`7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| v0.29.0 and main expose ec_manager_config; retire the old release-only
ScoreEncoder configuration skip. |
| #53240 |
`b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c)
| v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache
groups; retire the old release-only replay skip. |
| #53853 |
`e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
`4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both expose the V1 unsupported-feature helper used by the existing
Ascend MRV1 feature filter. |






#### Additional inherited contracts











| vllm-ascend change | Why it is required | Upstream cause and direct
link | Lane | Verification |
|---|---|---|---|---|





| `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM
v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec,
list[int]]` from `get_mamba_groups` and both initialize `recoverssm`;
keeping the old fallback would preserve an unsupported third contract |
[v0.29.0 mamba
groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704),
[fixed-main mamba
groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704),
[v0.29.0
RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103),
[fixed-main
RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103)
| both, common implementation | Existing constructor UT now asserts the
parent-created RecoverSSM value is retained; Ruff and compileall pass |
| `worker/v2/model_runner.py`: always forward
`kv_cache_allocation_context` | v0.29.0 and fixed main both accept this
keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0
signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538),
[fixed-main
signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566),
[vllm-ascend
vllm-project#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) |
both, common implementation | Existing UT continues to assert the exact
context object reaches the parent; Ruff and compileall pass |
| `_310p/worker/v2/model_runner.py`: select the release KV-zeroing
contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes
`KVCacheGroupSpec.is_eagle_group` but lacks
`SpeculativeConfig.use_eagle_block_drop`; fixed main added the method |
vLLM [#53388
diff](https://github.com/vllm-project/vllm/pull/53388/files), commit
[`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a);
[v0.29.0 group
field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200),
[fixed-main
helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876)
| release differs from main | Existing two-path KV-zeroing UT retained
and renamed for v0.29.0; Ruff and compileall pass |
| existing DFlash kernel UT: remove v0.28-only kwarg omission | the
current Ascend kernel accepts the CP arguments and the only excluded
lane was v0.28.0, which this PR replaces | [vllm-ascend
vllm-project#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) |
both, common invocation | Existing NPU test remains enabled with all
assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |





|---|---|---|---|





| `patch/platform/patch_parallel_config.py`, registration, and global
patch documentation | Allow Ascend PCP+DP by removing the generic GPU
restriction, following vLLM
[#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit
`7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic
`parallel_config.current_platform` lookup and rebuild ParallelConfig →
SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing
configuration/EPLB UTs and historical PCP+DP NPU execution passed;
current-head CI pending. |
| `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT |
Both supported contracts require `is_drafting`, from
[#54282](https://github.com/vllm-project/vllm/pull/54282/files),
`fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited
optimized kernel and use a common wrapper. | Both | Existing positive
drafting assertion retained; current-head CI pending. |
| `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache
UT | Both support `cache_hit_alignment_tokens`, introduced by
[#53598](https://github.com/vllm-project/vllm/pull/53598/files),
`2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only
write-mask branch. | Both | Existing assertions retained; current-head
CI pending. |






#### Rebase and retired-fallback evidence











| Change | Exact source evidence | Decision |





|---|---|---|





| `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM
[#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95),
`d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL
before v0.29.0; both exact supported sources lack it. The old
conditional came from vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0
selector and use the existing exclusion for both supported lanes. This
does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy
skip. |
| `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and
`tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM
[#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29),
`12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export
`nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The
v0.28.0 fallback originated in vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete
import-failure/`None` fallbacks and the existing UT's obsolete
availability skip; assertions remain unchanged and execute on both
lanes. |
| `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend
[vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3),
`799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream
revert while resolving the real rebase conflict; do not reintroduce the
reverted graph-update behavior. |
| `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend
[vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2),
`d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged
310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation
branch. |






</details>













</details>



### Does this PR introduce _any_ user-facing change?



Yes. The supported vLLM release and default container builds move to
v0.29.0, while the fixed main remains supported. The vllm-ascend package
version does not change. Existing 0.29-only PCP+DP compatibility is
restored.


### How was this patch tested?

- Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`:
[run
35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654)
succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected
NPU jobs confirm the requested Ascend head, integration base
`2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM
pins/installations: fixed main
`84030bbe3d74d99bad477a3d2e37a973ccd8865c` /
`0.1.dev1+g84030bbe3.empty`, release
`98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane
totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest
invocations. Skips/xfails are not passes. The resulting rebased
integration commit is not printed and is not inferred. Actual release
CPU and failed/cancelled image variants remain gaps.


- Current-head CI update (2026-09-20): [main CPU
UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955)
passed **5157 tests**, with **67 skipped**; actual installed vLLM was
`0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head;
the log does not print the full resulting integration head/base.
Pre-commit and mypy passed. NPU/E2E results are recorded above; actual
release CPU remains unverified.
- Image build is partially blocked by infrastructure: A5 amd64
[Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005)
and
[openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068)
failed before reading the Dockerfile because BuildKit could not create a
snapshot temporary directory (`no space left on device`). No release
compatibility code change is justified by this failure; cancelled
variants remain unverified. The successful [310P openEuler arm64
build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987)
explicitly checked out release
`98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`;
this is build evidence, not runtime UT coverage.


- Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto
the baseline above. Previous-head CPU [job
106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458)
stopped during Ascend integration rebase after vllm-project#16803 merged, before any
UT ran. Its vLLM checkout was the fixed main and installation reported
`0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved;
fresh-head CI was retriggered by the push.


- Remote official v0.29.0 tag resolved to the exact SHA above; both
source pins inspected.
- All 95 changed Python files pass Ruff lint, Ruff format and AST
parsing; `git diff --check` passes. Dockerfile tag/marker consistency
checked across all eight variants.
- Full pre-commit invocation: Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-dependent hooks cannot run on this
Windows host; Python launcher hooks exit 9009. Full CI lint remains
pending.
- Actual main/release CPU/NPU and image builds are not claimed locally:
the required Linux/NPU/container environment is unavailable. New PR E2E
and image-build CI are requested. Existing CPU workflow runs fixed main
only, so actual release CPU remains a validation gap.
- vllm-project#16393's historical green CI is not a substitute for this new head. No
new test functions, workflow changes, golden/threshold changes or
additional skips are introduced beyond restoring vllm-project#16393. Inherited
0.28-only tests/skips are not counted as passes.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
### What this PR does / why we need it?

The AllGather MoE path gathers hidden states and router logits in
DP-then-PCP order, while hash routing previously gathered `input_ids`
only across DP. With PCP enabled, this leaves fewer token IDs than
router-logit rows and causes hash routing to fail.

This PR:

- gathers hash-routing input IDs with the same DP-then-PCP padding and
communication layout used by the AllGather prepare path;
- renames the helper to reflect that it covers every required parallel
dimension;
- extends the existing prepare/finalize and hash-router regression
coverage for no parallelism, DP, PCP, and combined DP+PCP layouts.

### Does this PR introduce _any_ user-facing change?

No API or configuration change. Hash-based MoE routing now works with
the existing PCP configuration.

### How was this patch tested?

- A5 PCP4 (`DP1 × TP1 × PCP4`): service started and ran normally. Full
GPQA Diamond (`enable_thinking=false`) achieved 76.26% accuracy
(151/198), with 198/198 requests successful and a 56.68% DSpark
acceptance rate.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: leolee <yihao.li@huawei.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
…#16393) (vllm-project#16949)

Revert vllm-project#16393 at the author's request. This reverses squash commit
`af277b2681443089e4d29684b1101707e2eb6426` on upstream main
`8038f64129a1f8751ac6684f47c8d058208e4ff3`, restoring the previous
fixed-main + v0.28.0 support contract.

The revert applies without conflicts. Subsequent upstream patches are
retained, including vllm-project#16775 (PCP/DCP slots), vllm-project#16923 (A3 SFA prefill
all-to-all), and vllm-project#16834 (PCP hash-routing input IDs). No unrelated fixes
are included.

| Reverted change | Reason and source | Lane |
| --- | --- | --- |
| Release marker | Reverse
[vllm-project#16393](https://github.com/vllm-project/vllm-ascend/pull/16393/files):
v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`) to v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) | Release |
| Version branches and interface adapters across scheduler, KV cache,
attention, model runners and speculative decoding | Restore the
pre-vllm-project#16393 v0.28.0 contracts by reversing the exact merged diff; keep
main implementations behind their original version selection | Main +
release |
| v0.29.0-only parallel-config patch, registration and documentation |
Remove the patch introduced by vllm-project#16393; retain the independent upstream
v0.28.0 PCP+DP workaround | Release |
| Existing unit/e2e tests and release-only skips | Restore the tests and
conditions changed by vllm-project#16393, without introducing additional skip scope
beyond the reverted baseline | Main + release |

The vLLM main pin remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. No
main interface scan is needed for this release-only revert. The only
changes beyond the inverse patch are formatting in
`tests/ut/test_utils.py` and line-ending normalization in
`categorical_sample.py` needed for formatting/whitespace checks.
Workflows are unchanged.

Yes. The supported release returns to vLLM v0.28.0; v0.29.0 support and
its specific compatibility changes from vllm-project#16393 are reverted. The
vllm-ascend package version and fixed vLLM main pin are unchanged.

- `git revert --no-commit af277b2`
applied cleanly on the latest fetched main.
- Verified that the complete post-vllm-project#16393 upstream patch can still
reverse-apply to the staged result (`git apply --reverse --check
--cached`), confirming subsequent changes are preserved.
- `git diff --cached --check`, Python AST parsing, Ruff lint and
formatting checks passed for all 94 remaining changed Python files.
- Ran the equivalent of `format.sh ci`: `pre-commit run --all-files
--hook-stage manual`. Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-based hooks could not execute on
this Windows host; two Python launcher hooks exited 9009. Full
pre-commit validation remains for CI.
- Actual CPU/NPU tests have not been run locally because the required
vLLM/NPU runtime is unavailable. New PR CI is pending; the original
upgrade PR's green run is not validation of this revert.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…vllm-project#17004)

### What this PR does / why we need it?



Reintroduce the vLLM v0.29.0 release upgrade from vllm-project#16393, reverted by
vllm-project#16949, with container defaults aligned to the supported release. Based
on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves
later merged changes.


The image/source mismatch is confirmed in [nightly job
106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020):
the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded
Ascend failed importing `_get_packed_kv_cache_groups`. All eight root
Dockerfiles still defaulted to v0.28.0. The reusable image workflow
passes no `VLLM_TAG` override, so those defaults govern release builds.
Changing only the release marker does not update those images.


- Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

- Release changes from v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0
(`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the
remote tag).
- Restore vllm-project#16393's source compatibility and existing UT changes by
reversing vllm-project#16949, then reconcile current upstream changes. No main
interface scan is rerun for this release-only upgrade.
- Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P,
Ubuntu/openEuler. Preserve existing exact-commit build overrides. No
workflow changes.


| Current-base adaptation | Exact cause and evidence | Lane | Validation
|
| --- | --- | --- | --- |

| Eight Dockerfile `VLLM_TAG` defaults and release marker | vllm-project#16393
changed the release contract but omitted image defaults; the linked
nightly log proves installation of 0.28.0 and missing
`_get_packed_kv_cache_groups`. | Release images | All eight defaults
match marker; image-build CI requested, pending |
| `worker/v2/attn_utils.py::get_kv_cache_spec` and existing
`_make_mla_layer` UT fixture | Ascend
[vllm-project#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files),
`df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field
selector. vLLM
[#51718](https://github.com/vllm-project/vllm/pull/51718/files),
`8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio`
with `tokens_per_state`; both exact supported pins use the latter. Use
the common field, retaining metadata and cache-view assertions. | Both |
Source inspection and static checks passed; actual CI pending |
| `attention/attention_v1.py` import conflict | Preserve
`attention_transfer_window` from Ascend
[vllm-project#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files),
`5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from
vllm-project#16798; keep the common relocated PCP import from vllm-project#16393. Do not restore
the old compute-start import or unused weak-reference import. | Both |
Conflict resolved; static checks passed |
| `worker/v2/aclgraph_utils.py` import conflict | Preserve
`ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend
[vllm-project#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files),
`34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete
0.28 version selector restored by the revert. | Both | Conflict
resolved; static checks passed |


| Existing 310P and Mamba model-runner UT imports | Preserve
hardware-profile imports and mocks from Ascend
[vllm-project#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files),
`d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused
old-release selector import. Hardware capability routing remains
unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format
passed; fresh CI pending |

Newly merged changes were reviewed for version-contract impact:
vllm-project#16775/vllm-project#16834 PCP metadata and routing, vllm-project#16923 A3 SFA,
vllm-project#16426/vllm-project#16924/vllm-project#16955 custom ops, vllm-project#16081 MTP/SP, vllm-project#16669 xlite,
vllm-project#15636/vllm-project#16747 transfer, vllm-project#16913 A5 pages, vllm-project#16798 graph updates, vllm-project#16755
GLM PD, vllm-project#16952 C8 config, vllm-project#16673 operator removal and vllm-project#16320 Kimi-K3 KV
pool. Preserve these changes; no additional source-proven version branch
was identified beyond the entries above. In particular, vllm-project#16747's
`UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size`
exist in both exact pins; vllm-project#16320 adds Ascend connector hooks. CI remains
necessary to validate runtime interactions. CI/documentation-only PRs
are retained unchanged.


<details>

<summary>Inherited per-file compatibility evidence from vllm-project#16393</summary>


The following source-contract ledger is inherited from vllm-project#16393. Any
historical verification wording refers only to that earlier PR; it does
not certify this new head. New-head validation is listed below.


| Reference | Commit |





|---|---|





| vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` |
| PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` |





| Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` |





| Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a`
|
| Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` |






- Use common implementations where v0.29.0 and the fixed main share
KV-cache layouts, Mamba copy/group APIs, PCP handling and
speculative-decoding contracts.
- Retain explicit `vllm_version_is("0.29.0")` branches for contracts
that still differ, including RoPE, scheduler block snapshots,
InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.
- Remove obsolete v0.28.0 compatibility and adapt existing test
fixtures. Version detection uses package versions and the explicit
`VLLM_VERSION` override, with local-version suffix handling and UT
environment isolation retained; no hard-coded release-SHA inference
remains.
- Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving
other validations and EPLB platform binding. Rebuild dependent Pydantic
schemas so nested configuration validation uses the patched validator.
The global patch documentation records its rationale and removal
criteria.






- Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase;
use the common exact-release selector and the real function-local DSpark
sharing import.






This release-only upgrade does not require a new main old-to-new
interface scan. The latest rebase incorporates
[vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files),
which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The
now-unnecessary V4.1 drafter import gate and tuple annotation have been
removed. Other release adaptations, including vllm-project#15747 Spec+PP handling,
remain. Detailed contract evidence is retained below for review.






<details>





<summary>Per-file adaptations and exact upstream evidence</summary>






#### Per-file adaptation ledger











Evidence IDs refer to the exact source contract and upstream diff table
below. Every row is syntax checked; branch-normalized AST comparison
confirms unchanged main function bodies except the version identity
helper and the explicitly retired propose argument/type annotations.
Both supported versions completed CPU and NPU execution as recorded
below.






| Ascend file / symbols | Disposition and upstream evidence |
Verification |
|---|---|---|





| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` /
`_load_dspark_model_with_target_quant`;
`tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased
[Ascend
vllm-project#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a)
(`82df9d871`), including manual PP partition masking. vLLM [#52809
diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share`
inside the loader on both supported pins, so retain the earlier release
fix: patch `eagle_utils`, never read/patch a nonexistent
`dspark_utils._should_share`. v0.29 keeps its global PP guard; main has
native PP via
[#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks
the real function-local import, absent module alias, and restoration on
success/failure for both lanes; partition and PP assertions retained.
Ruff/syntax pass; actual CPU/NPU pending. |
| `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`;
`tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29
routing and an obsolete 0.28 dev-build recognition path. For the two
supported pins, #50514 exists only on fixed main. Route through
`vllm_version_is("0.29.0")`; retain package/local-suffix and explicit
environment-override semantics. No release-SHA inference or third
release lane. | Existing UTs exercise the real uncached version helper
with monkeypatch isolation, local suffix, fixed-main dev string,
explicit override, and non-target versions. Isolated routing checks and
Ruff/syntax pass; full CPU UT pending. |
| `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`,
`_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass`
| #56078; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`,
`_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436;
use shared standardized layouts and retain the v0.29.0 InputBatch gate.
The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while
removing only legacy v0.28.0 allocation paths. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` |
#56078; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` |
#52839; both lanes use the common PCP import. Rebase keeps current
upstream graph code and removes only the v0.28.0 import branch. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module
imports/dispatch` | #52839; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch`
| #51358, #54853; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/core/recompute_scheduler.py`<br>`module
imports/dispatch`, `schedule` | #51358, #54853; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`,
`AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker`
| #52615; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__`
| #53614; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module
imports/dispatch`, `_get_max_layers_per_page_size`,
`_ascend_max_memory_usage_bytes_from_groups`,
`_ascend_get_kv_cache_config_from_groups` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module
imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` |
v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it
after KV binding. Evidence: [#52506,
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils
diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c).
Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. |
Failure reproduced on v0.29.0; Historical release/main NPU validation
passed; current-head CI pending |
| `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module
imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module
imports/dispatch` | #51718; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant`
| #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports
`_should_share` locally from Eagle utilities; fixed main removes the PP
guard. Keep the release `get_pp_group` patch, share through the common
Eagle utility, and delete the obsolete v0.28.0
`dspark_utils._should_share` patch. | Exact release failure reproduced;
Historical release/main CPU/NPU validation passed; current-head CI
pending |
|
`vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`,
`__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`,
`register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`,
`_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`,
`_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers |
#51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and
import the NaN helpers directly because v0.29.0 and fixed main expose
the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static
checked; historical dual-version CPU/NPU validation passed; current-head
CI pending |
| `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes
`pcp_manager` common to both lanes. The rebase preserves vllm-ascend
vllm-project#16409's host-parameter-update revert and removes only the obsolete
v0.28.0 capture branch. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`,
`_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/worker/v2/block_table.py`<br>`__init__`,
`init_block_table_layout_tensors`, `compute_slot_mappings` | #51718,
#51031; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`,
`prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`,
`execute_model` | #50514, #54436, #52506, #55212, #53515; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch`
| #54282; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` |
#49811; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
|
`vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module
imports/dispatch`, `propose` | #53694, #52188; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`,
`wake_up` | #51718, #53508; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |






#### Exact upstream evidence











| Upstream change | Full commit SHA / direct diff | Actual supported
contracts and branch decision |
|---|---|---|





| #51718 |
`8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Both use layers/layer_stride/block_stride/offset, tokens_per_state,
CircularBufferSpec and standardized backing; retire
shared_by/compress_ratio allocation branches. |
| #52839 |
`58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8)
| Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
`e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec)
| Both use Mamba copy-function dictionaries and unwrap
UniformTypeKVCacheSpecs. |
| #53106 |
`1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21)
| Both use WeightsMapper instead of AutoWeightsLoader
skip_prefixes/skip_substrs. |
| #53906 |
`98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Only pinned main has the optional MLA storage_block_size dataclass
field; release keeps the Ascend derived property, using
tokens_per_state. |
| #56078 |
`719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417)
| Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned
main uses mrope_num_dims and unified RoPE. |
| #52615 |
`138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045)
| Release uses num_blocks/kv_bytes_per_block; main uses
num_chunks/kv_bytes_per_chunk. |
| #51358 |
`6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release now has boundary_state_offloads and KVConnectorBlockState;
remove partial_tail_offloads plumbing. |
| #54853 |
`0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release constructor takes block_ids snapshots; main takes req_ids and
resolve_block_ids. Keep exact release snapshot membership and main
lazy-resolution membership. |
| #53614 |
`144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725)
| Only main configures drop_eagle_checkpoint_block for replay-aligned
Mamba checkpoints. |
| #50514 |
`d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Release retains module-level PP/share symbols and Ascend PP
workaround; main has the subsequent PP integration. |
| #52809 |
`91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark
module binding to a function-local import from Eagle utilities. The
shared Eagle patch remains effective; the old DSpark-module read/write
must be removed. |
| #54436 |
`6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902)
| Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Only main accepts valid_dummy_state_slots/valid_state_slots capture
arguments. |
| #55212 |
`83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Release prepares DCP local sequence lengths before partitioning; main
initializes DCP metadata afterwards. |
| #53515 |
`b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127)
| Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
`b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212)
| Both accept pcp_manager during graph capture. |
| #51031 |
`0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1)
| Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
`fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8)
| Both gumbel sampling APIs include is_drafting. |
| #52188 |
`d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0)
| Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
`5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6)
| Both propose APIs take DPSyncState; remove the obsolete token-count
argument and retain replicated-PCP synchronization. |
| #49811 |
`01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e)
| Both support extract_hidden_states on MRV2; remove old unsupported
dispatch/skip. |
| #53508 |
`479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29)
| Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
`3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1)
| Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain
release exclusion. |
| #52861 |
`b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28)
| Both include DeepseekV32MTPModel in the two-hidden-state architecture
set. |
| #54713 |
`b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3)
| Only main takes replay_boundaries in compressed-prefix hit lookup;
preserve release calls without that keyword. |
| #42785 |
`442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18)
| Only main capture callers pass axis_keys; preserve the existing Ascend
rejection of nonempty axes. |
| #52358 |
`8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
`9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4)
| Release already has mamba_has_prefill_checkpoint_blocks; later main
also has fine-grained prefix-cache state. |
| #51251 |
`7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| v0.29.0 and main expose ec_manager_config; retire the old release-only
ScoreEncoder configuration skip. |
| #53240 |
`b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c)
| v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache
groups; retire the old release-only replay skip. |
| #53853 |
`e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
`4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both expose the V1 unsupported-feature helper used by the existing
Ascend MRV1 feature filter. |






#### Additional inherited contracts











| vllm-ascend change | Why it is required | Upstream cause and direct
link | Lane | Verification |
|---|---|---|---|---|





| `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM
v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec,
list[int]]` from `get_mamba_groups` and both initialize `recoverssm`;
keeping the old fallback would preserve an unsupported third contract |
[v0.29.0 mamba
groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704),
[fixed-main mamba
groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704),
[v0.29.0
RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103),
[fixed-main
RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103)
| both, common implementation | Existing constructor UT now asserts the
parent-created RecoverSSM value is retained; Ruff and compileall pass |
| `worker/v2/model_runner.py`: always forward
`kv_cache_allocation_context` | v0.29.0 and fixed main both accept this
keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0
signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538),
[fixed-main
signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566),
[vllm-ascend
vllm-project#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) |
both, common implementation | Existing UT continues to assert the exact
context object reaches the parent; Ruff and compileall pass |
| `_310p/worker/v2/model_runner.py`: select the release KV-zeroing
contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes
`KVCacheGroupSpec.is_eagle_group` but lacks
`SpeculativeConfig.use_eagle_block_drop`; fixed main added the method |
vLLM [#53388
diff](https://github.com/vllm-project/vllm/pull/53388/files), commit
[`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a);
[v0.29.0 group
field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200),
[fixed-main
helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876)
| release differs from main | Existing two-path KV-zeroing UT retained
and renamed for v0.29.0; Ruff and compileall pass |
| existing DFlash kernel UT: remove v0.28-only kwarg omission | the
current Ascend kernel accepts the CP arguments and the only excluded
lane was v0.28.0, which this PR replaces | [vllm-ascend
vllm-project#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) |
both, common invocation | Existing NPU test remains enabled with all
assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |





|---|---|---|---|





| `patch/platform/patch_parallel_config.py`, registration, and global
patch documentation | Allow Ascend PCP+DP by removing the generic GPU
restriction, following vLLM
[#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit
`7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic
`parallel_config.current_platform` lookup and rebuild ParallelConfig →
SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing
configuration/EPLB UTs and historical PCP+DP NPU execution passed;
current-head CI pending. |
| `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT |
Both supported contracts require `is_drafting`, from
[#54282](https://github.com/vllm-project/vllm/pull/54282/files),
`fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited
optimized kernel and use a common wrapper. | Both | Existing positive
drafting assertion retained; current-head CI pending. |
| `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache
UT | Both support `cache_hit_alignment_tokens`, introduced by
[#53598](https://github.com/vllm-project/vllm/pull/53598/files),
`2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only
write-mask branch. | Both | Existing assertions retained; current-head
CI pending. |






#### Rebase and retired-fallback evidence











| Change | Exact source evidence | Decision |





|---|---|---|





| `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM
[#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95),
`d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL
before v0.29.0; both exact supported sources lack it. The old
conditional came from vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0
selector and use the existing exclusion for both supported lanes. This
does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy
skip. |
| `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and
`tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM
[#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29),
`12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export
`nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The
v0.28.0 fallback originated in vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete
import-failure/`None` fallbacks and the existing UT's obsolete
availability skip; assertions remain unchanged and execute on both
lanes. |
| `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend
[vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3),
`799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream
revert while resolving the real rebase conflict; do not reintroduce the
reverted graph-update behavior. |
| `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend
[vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2),
`d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged
310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation
branch. |






</details>













</details>



### Does this PR introduce _any_ user-facing change?



Yes. The supported vLLM release and default container builds move to
v0.29.0, while the fixed main remains supported. The vllm-ascend package
version does not change. Existing 0.29-only PCP+DP compatibility is
restored.


### How was this patch tested?

- Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`:
[run
35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654)
succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected
NPU jobs confirm the requested Ascend head, integration base
`2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM
pins/installations: fixed main
`84030bbe3d74d99bad477a3d2e37a973ccd8865c` /
`0.1.dev1+g84030bbe3.empty`, release
`98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane
totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest
invocations. Skips/xfails are not passes. The resulting rebased
integration commit is not printed and is not inferred. Actual release
CPU and failed/cancelled image variants remain gaps.


- Current-head CI update (2026-09-20): [main CPU
UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955)
passed **5157 tests**, with **67 skipped**; actual installed vLLM was
`0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head;
the log does not print the full resulting integration head/base.
Pre-commit and mypy passed. NPU/E2E results are recorded above; actual
release CPU remains unverified.
- Image build is partially blocked by infrastructure: A5 amd64
[Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005)
and
[openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068)
failed before reading the Dockerfile because BuildKit could not create a
snapshot temporary directory (`no space left on device`). No release
compatibility code change is justified by this failure; cancelled
variants remain unverified. The successful [310P openEuler arm64
build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987)
explicitly checked out release
`98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`;
this is build evidence, not runtime UT coverage.


- Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto
the baseline above. Previous-head CPU [job
106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458)
stopped during Ascend integration rebase after vllm-project#16803 merged, before any
UT ran. Its vLLM checkout was the fixed main and installation reported
`0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved;
fresh-head CI was retriggered by the push.


- Remote official v0.29.0 tag resolved to the exact SHA above; both
source pins inspected.
- All 95 changed Python files pass Ruff lint, Ruff format and AST
parsing; `git diff --check` passes. Dockerfile tag/marker consistency
checked across all eight variants.
- Full pre-commit invocation: Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-dependent hooks cannot run on this
Windows host; Python launcher hooks exit 9009. Full CI lint remains
pending.
- Actual main/release CPU/NPU and image builds are not claimed locally:
the required Linux/NPU/container environment is unavailable. New PR E2E
and image-build CI are requested. Existing CPU workflow runs fixed main
only, so actual release CPU remains a validation gap.
- vllm-project#16393's historical green CI is not a substitute for this new head. No
new test functions, workflow changes, golden/threshold changes or
additional skips are introduced beyond restoring vllm-project#16393. Inherited
0.28-only tests/skips are not counted as passes.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants