[Misc][Main2Main] Upgrade vLLM release dependency to v0.29.0 - #16393
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request updates the vLLM release dependency to v0.29.0 to align with upstream changes. It refactors internal Ascend-specific logic to support the new release contracts while preserving compatibility with the pinned main branch. The changes focus on removing legacy version-gated code paths and standardizing KV cache management and attention backend interfaces to match the v0.29.0 release. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[CI][Misc] Bump verified vLLM version to v0.29.0 and remove v0.28.0 compatibility codeSuggested PR Summary:
### What this PR does / why we need it?
This PR bumps the verified vLLM version from `v0.28.0` to `v0.29.0` and cleans up legacy compatibility code, workarounds, and test skips that were specific to `v0.28.0`. It simplifies the codebase by removing version-gated branches (e.g., `vllm_version_is("0.28.0")`) across various modules including attention, schedulers, speculators, and model runners.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested via existing unit tests and end-to-end tests updated to align with `v0.29.0` specifications.I have reviewed the changes and have no additional feedback to provide.
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
dd2de16 to
7ed85bc
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
46dfd05 to
0f1a0cc
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
dbe6f6c to
ccbf61e
Compare
Rollup of unrelated failures across the last four runs: - R1: pytest collection of test_qwen3_0.6B_basic.py (dot in filename, fixed here) - R2: runner git-checkout network timeout (a2-1 part 2/5) - R3/R5: test_glm5_2_dspark_eager[fixed] acceptance threshold - measured mean 3.07-3.16 vs floor 3.15 (3.5 * 0.9); passed on other PRs the same day (e.g. vllm-project#16393 a3-8), no code-path overlap with this PR's diff. Signed-off-by: AceCoder0 <hellozkr@163.com>
Rollup of unrelated failures across the last four runs: - R1: pytest collection of test_qwen3_0.6B_basic.py (dot in filename, fixed here) - R2: runner git-checkout network timeout (a2-1 part 2/5) - R3/R5: test_glm5_2_dspark_eager[fixed] acceptance threshold - measured mean 3.07-3.16 vs floor 3.15 (3.5 * 0.9); passed on other PRs the same day (e.g. vllm-project#16393 a3-8), no code-path overlap with this PR's diff. Signed-off-by: AceCoder0 <hellozkr@163.com>
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
This reverts commit 2560da9. Signed-off-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: shenzhao <shenzhao9@huawei.com>
8c5d785 to
fd876c5
Compare
Resolve conflicts from vllm-project#16393 (vLLM 0.28.0 -> 0.29.0). Keep this branch's default-V2 + blacklist selection, adopt 0.29 propose(dp_sync), and always bind the V1 unsupported-feature helper. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
…llm-project#16393)" This reverts commit af277b2. Signed-off-by: chenzeyu <2978509328@qq.com>
) ### What this PR does / why we need it? Revert #16393 at the author's request. This reverses squash commit `af277b2681443089e4d29684b1101707e2eb6426` on upstream main `8038f64129a1f8751ac6684f47c8d058208e4ff3`, restoring the previous fixed-main + v0.28.0 support contract. The revert applies without conflicts. Subsequent upstream patches are retained, including #16775 (PCP/DCP slots), #16923 (A3 SFA prefill all-to-all), and #16834 (PCP hash-routing input IDs). No unrelated fixes are included. | Reverted change | Reason and source | Lane | | --- | --- | --- | | Release marker | Reverse [#16393](https://github.com/vllm-project/vllm-ascend/pull/16393/files): v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`) to v0.28.0 (`2cf0a6915ce544dc493a0990f2ea38d81601128a`) | Release | | Version branches and interface adapters across scheduler, KV cache, attention, model runners and speculative decoding | Restore the pre-#16393 v0.28.0 contracts by reversing the exact merged diff; keep main implementations behind their original version selection | Main + release | | v0.29.0-only parallel-config patch, registration and documentation | Remove the patch introduced by #16393; retain the independent upstream v0.28.0 PCP+DP workaround | Release | | Existing unit/e2e tests and release-only skips | Restore the tests and conditions changed by #16393, without introducing additional skip scope beyond the reverted baseline | Main + release | The vLLM main pin remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. No main interface scan is needed for this release-only revert. The only changes beyond the inverse patch are formatting in `tests/ut/test_utils.py` and line-ending normalization in `categorical_sample.py` needed for formatting/whitespace checks. Workflows are unchanged. ### Does this PR introduce _any_ user-facing change? Yes. The supported release returns to vLLM v0.28.0; v0.29.0 support and its specific compatibility changes from #16393 are reverted. The vllm-ascend package version and fixed vLLM main pin are unchanged. ### How was this patch tested? - `git revert --no-commit af277b2` applied cleanly on the latest fetched main. - Verified that the complete post-#16393 upstream patch can still reverse-apply to the staged result (`git apply --reverse --check --cached`), confirming subsequent changes are preserved. - `git diff --cached --check`, Python AST parsing, Ruff lint and formatting checks passed for all 94 remaining changed Python files. - Ran the equivalent of `format.sh ci`: `pre-commit run --all-files --hook-stage manual`. Ruff, codespell, typos, clang-format, markdownlint, actionlint, package-init, forbidden-import and boolean-context checks passed. Bash-based hooks could not execute on this Windows host; two Python launcher hooks exited 9009. Full pre-commit validation remains for CI. - Actual CPU/NPU tests have not been run locally because the required vLLM/NPU runtime is unavailable. New PR CI is pending; the original upgrade PR's green run is not validation of this revert. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…llm-project#16393)" This reverts commit af277b2. Signed-off-by: chenzeyu <2978509328@qq.com>
Restore vllm-project#16393 on current main, align all eight Dockerfile release defaults, and reconcile GLM MRV2 cache metadata with the shared v0.29.0/main contract. Preserve subsequent upstream graph and transfer changes. Signed-off-by: shenzhao <shenzhao9@huawei.com>
…#17004) ### What this PR does / why we need it? Reintroduce the vLLM v0.29.0 release upgrade from #16393, reverted by #16949, with container defaults aligned to the supported release. Based on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves later merged changes. The image/source mismatch is confirmed in [nightly job 106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020): the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded Ascend failed importing `_get_packed_kv_cache_groups`. All eight root Dockerfiles still defaulted to v0.28.0. The reusable image workflow passes no `VLLM_TAG` override, so those defaults govern release builds. Changing only the release marker does not update those images. - Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. - Release changes from v0.28.0 (`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the remote tag). - Restore #16393's source compatibility and existing UT changes by reversing #16949, then reconcile current upstream changes. No main interface scan is rerun for this release-only upgrade. - Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P, Ubuntu/openEuler. Preserve existing exact-commit build overrides. No workflow changes. | Current-base adaptation | Exact cause and evidence | Lane | Validation | | --- | --- | --- | --- | | Eight Dockerfile `VLLM_TAG` defaults and release marker | #16393 changed the release contract but omitted image defaults; the linked nightly log proves installation of 0.28.0 and missing `_get_packed_kv_cache_groups`. | Release images | All eight defaults match marker; image-build CI requested, pending | | `worker/v2/attn_utils.py::get_kv_cache_spec` and existing `_make_mla_layer` UT fixture | Ascend [#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files), `df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field selector. vLLM [#51718](https://github.com/vllm-project/vllm/pull/51718/files), `8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio` with `tokens_per_state`; both exact supported pins use the latter. Use the common field, retaining metadata and cache-view assertions. | Both | Source inspection and static checks passed; actual CI pending | | `attention/attention_v1.py` import conflict | Preserve `attention_transfer_window` from Ascend [#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files), `5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from #16798; keep the common relocated PCP import from #16393. Do not restore the old compute-start import or unused weak-reference import. | Both | Conflict resolved; static checks passed | | `worker/v2/aclgraph_utils.py` import conflict | Preserve `ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend [#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files), `34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete 0.28 version selector restored by the revert. | Both | Conflict resolved; static checks passed | | Existing 310P and Mamba model-runner UT imports | Preserve hardware-profile imports and mocks from Ascend [#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files), `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused old-release selector import. Hardware capability routing remains unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format passed; fresh CI pending | Newly merged changes were reviewed for version-contract impact: #16775/#16834 PCP metadata and routing, #16923 A3 SFA, #16426/#16924/#16955 custom ops, #16081 MTP/SP, #16669 xlite, #15636/#16747 transfer, #16913 A5 pages, #16798 graph updates, #16755 GLM PD, #16952 C8 config, #16673 operator removal and #16320 Kimi-K3 KV pool. Preserve these changes; no additional source-proven version branch was identified beyond the entries above. In particular, #16747's `UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size` exist in both exact pins; #16320 adds Ascend connector hooks. CI remains necessary to validate runtime interactions. CI/documentation-only PRs are retained unchanged. <details> <summary>Inherited per-file compatibility evidence from #16393</summary> The following source-contract ledger is inherited from #16393. Any historical verification wording refers only to that earlier PR; it does not certify this new head. New-head validation is listed below. | Reference | Commit | |---|---| | vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` | | PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` | | Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` | | Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a` | | Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` | - Use common implementations where v0.29.0 and the fixed main share KV-cache layouts, Mamba copy/group APIs, PCP handling and speculative-decoding contracts. - Retain explicit `vllm_version_is("0.29.0")` branches for contracts that still differ, including RoPE, scheduler block snapshots, InputBatch, ReplaySSM, KV zeroing and DSpark PP handling. - Remove obsolete v0.28.0 compatibility and adapt existing test fixtures. Version detection uses package versions and the explicit `VLLM_VERSION` override, with local-version suffix handling and UT environment isolation retained; no hard-coded release-SHA inference remains. - Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving other validations and EPLB platform binding. Rebuild dependent Pydantic schemas so nested configuration validation uses the patched validator. The global patch documentation records its rationale and removal criteria. - Preserve #15747's Spec+PP protocol/partition handling after rebase; use the common exact-release selector and the real function-local DSpark sharing import. This release-only upgrade does not require a new main old-to-new interface scan. The latest rebase incorporates [#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files), which reverted #16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The now-unnecessary V4.1 drafter import gate and tuple annotation have been removed. Other release adaptations, including #15747 Spec+PP handling, remain. Detailed contract evidence is retained below for review. <details> <summary>Per-file adaptations and exact upstream evidence</summary> #### Per-file adaptation ledger Evidence IDs refer to the exact source contract and upstream diff table below. Every row is syntax checked; branch-normalized AST comparison confirms unchanged main function bodies except the version identity helper and the explicitly retired propose argument/type annotations. Both supported versions completed CPU and NPU execution as recorded below. | Ascend file / symbols | Disposition and upstream evidence | Verification | |---|---|---| | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` / `_load_dspark_model_with_target_quant`; `tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased [Ascend #15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a) (`82df9d871`), including manual PP partition masking. vLLM [#52809 diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share` inside the loader on both supported pins, so retain the earlier release fix: patch `eagle_utils`, never read/patch a nonexistent `dspark_utils._should_share`. v0.29 keeps its global PP guard; main has native PP via [#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks the real function-local import, absent module alias, and restoration on success/failure for both lanes; partition and PP assertions retained. Ruff/syntax pass; actual CPU/NPU pending. | | `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`; `tests/ut/worker/v2/test_pp_utils.py` | #15747 added broad 0.28/0.29 routing and an obsolete 0.28 dev-build recognition path. For the two supported pins, #50514 exists only on fixed main. Route through `vllm_version_is("0.29.0")`; retain package/local-suffix and explicit environment-override semantics. No release-SHA inference or third release lane. | Existing UTs exercise the real uncached version helper with monkeypatch isolation, local suffix, fixed-main dev string, explicit override, and non-target versions. Isolated routing checks and Ruff/syntax pass; full CPU UT pending. | | `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`, `_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`, `_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436; use shared standardized layouts and retain the v0.29.0 InputBatch gate. The rebase preserves vllm-ascend #16043's MTP copy tracking while removing only legacy v0.28.0 allocation paths. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` | #52839; both lanes use the common PCP import. Rebase keeps current upstream graph code and removes only the v0.28.0 import branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/recompute_scheduler.py`<br>`module imports/dispatch`, `schedule` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`, `AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker` | #52615; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__` | #53614; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module imports/dispatch`, `_get_max_layers_per_page_size`, `_ascend_max_memory_usage_bytes_from_groups`, `_ascend_get_kv_cache_config_from_groups` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` | v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it after KV binding. Evidence: [#52506, `adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c). Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. | Failure reproduced on v0.29.0; Historical release/main NPU validation passed; current-head CI pending | | `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module imports/dispatch` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant` | #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports `_should_share` locally from Eagle utilities; fixed main removes the PP guard. Keep the release `get_pp_group` patch, share through the common Eagle utility, and delete the obsolete v0.28.0 `dspark_utils._should_share` patch. | Exact release failure reproduced; Historical release/main CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`, `__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`, `register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`, `_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`, `_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers | #51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and import the NaN helpers directly because v0.29.0 and fixed main expose the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static checked; historical dual-version CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes `pcp_manager` common to both lanes. The rebase preserves vllm-ascend #16409's host-parameter-update revert and removes only the obsolete v0.28.0 capture branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`, `_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/block_table.py`<br>`__init__`, `init_block_table_layout_tensors`, `compute_slot_mappings` | #51718, #51031; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`, `prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`, `execute_model` | #50514, #54436, #52506, #55212, #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch` | #54282; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` | #49811; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module imports/dispatch`, `propose` | #53694, #52188; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`, `wake_up` | #51718, #53508; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | #### Exact upstream evidence | Upstream change | Full commit SHA / direct diff | Actual supported contracts and branch decision | |---|---|---| | #51718 | `8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Both use layers/layer_stride/block_stride/offset, tokens_per_state, CircularBufferSpec and standardized backing; retire shared_by/compress_ratio allocation branches. | | #52839 | `58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8) | Both import PCP operations from vllm.v1.attention.ops.pcp. | | #53896 | `e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec) | Both use Mamba copy-function dictionaries and unwrap UniformTypeKVCacheSpecs. | | #53106 | `1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21) | Both use WeightsMapper instead of AutoWeightsLoader skip_prefixes/skip_substrs. | | #53906 | `98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Only pinned main has the optional MLA storage_block_size dataclass field; release keeps the Ascend derived property, using tokens_per_state. | | #56078 | `719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417) | Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned main uses mrope_num_dims and unified RoPE. | | #52615 | `138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045) | Release uses num_blocks/kv_bytes_per_block; main uses num_chunks/kv_bytes_per_chunk. | | #51358 | `6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release now has boundary_state_offloads and KVConnectorBlockState; remove partial_tail_offloads plumbing. | | #54853 | `0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release constructor takes block_ids snapshots; main takes req_ids and resolve_block_ids. Keep exact release snapshot membership and main lazy-resolution membership. | | #53614 | `144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725) | Only main configures drop_eagle_checkpoint_block for replay-aligned Mamba checkpoints. | | #50514 | `d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Release retains module-level PP/share symbols and Ascend PP workaround; main has the subsequent PP integration. | | #52809 | `91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark module binding to a function-local import from Eagle utilities. The shared Eagle patch remains effective; the old DSpark-module read/write must be removed. | | #54436 | `6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902) | Release InputBatch requires max_seq_len_np; main removed it. | | #52506 | `adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Only main accepts valid_dummy_state_slots/valid_state_slots capture arguments. | | #55212 | `83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Release prepares DCP local sequence lengths before partitioning; main initializes DCP metadata afterwards. | | #53515 | `b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127) | Both accept padded_num_tokens for persistent PCP input buffers. | | #53869 | `b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212) | Both accept pcp_manager during graph capture. | | #51031 | `0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1) | Both distinguish KV and kernel block sizes during DCP slot mapping. | | #54282 | `fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8) | Both gumbel sampling APIs include is_drafting. | | #52188 | `d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0) | Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. | | #53694 | `5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6) | Both propose APIs take DPSyncState; remove the obsolete token-count argument and retain replicated-PCP synchronization. | | #49811 | `01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e) | Both support extract_hidden_states on MRV2; remove old unsupported dispatch/skip. | | #53508 | `479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29) | Both remove post_kv_cache_wake_up; retire release-only call. | | #52494 | `3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1) | Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain release exclusion. | | #52861 | `b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28) | Both include DeepseekV32MTPModel in the two-hidden-state architecture set. | | #54713 | `b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3) | Only main takes replay_boundaries in compressed-prefix hit lookup; preserve release calls without that keyword. | | #42785 | `442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18) | Only main capture callers pass axis_keys; preserve the existing Ascend rejection of nonempty axes. | | #52358 | `8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Both ExecuteModelState have dp_sync; only main has cudagraph_stats. | | #52789 | `9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4) | Release already has mamba_has_prefill_checkpoint_blocks; later main also has fine-grained prefix-cache state. | | #51251 | `7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | v0.29.0 and main expose ec_manager_config; retire the old release-only ScoreEncoder configuration skip. | | #53240 | `b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c) | v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache groups; retire the old release-only replay skip. | | #53853 | `e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both delegate PCP compatibility validation to the PCP manager. | | #53183 | `4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both expose the V1 unsupported-feature helper used by the existing Ascend MRV1 feature filter. | #### Additional inherited contracts | vllm-ascend change | Why it is required | Upstream cause and direct link | Lane | Verification | |---|---|---|---|---| | `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec, list[int]]` from `get_mamba_groups` and both initialize `recoverssm`; keeping the old fallback would preserve an unsupported third contract | [v0.29.0 mamba groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704), [fixed-main mamba groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704), [v0.29.0 RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103), [fixed-main RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103) | both, common implementation | Existing constructor UT now asserts the parent-created RecoverSSM value is retained; Ruff and compileall pass | | `worker/v2/model_runner.py`: always forward `kv_cache_allocation_context` | v0.29.0 and fixed main both accept this keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0 signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538), [fixed-main signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566), [vllm-ascend #16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) | both, common implementation | Existing UT continues to assert the exact context object reaches the parent; Ruff and compileall pass | | `_310p/worker/v2/model_runner.py`: select the release KV-zeroing contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes `KVCacheGroupSpec.is_eagle_group` but lacks `SpeculativeConfig.use_eagle_block_drop`; fixed main added the method | vLLM [#53388 diff](https://github.com/vllm-project/vllm/pull/53388/files), commit [`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a); [v0.29.0 group field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200), [fixed-main helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876) | release differs from main | Existing two-path KV-zeroing UT retained and renamed for v0.29.0; Ruff and compileall pass | | existing DFlash kernel UT: remove v0.28-only kwarg omission | the current Ascend kernel accepts the CP arguments and the only excluded lane was v0.28.0, which this PR replaces | [vllm-ascend #15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) | both, common invocation | Existing NPU test remains enabled with all assertions; Current-head CI pending | | Ascend change | Why / upstream cause | Lane | Verification | |---|---|---|---| | `patch/platform/patch_parallel_config.py`, registration, and global patch documentation | Allow Ascend PCP+DP by removing the generic GPU restriction, following vLLM [#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic `parallel_config.current_platform` lookup and rebuild ParallelConfig → SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing configuration/EPLB UTs and historical PCP+DP NPU execution passed; current-head CI pending. | | `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT | Both supported contracts require `is_drafting`, from [#54282](https://github.com/vllm-project/vllm/pull/54282/files), `fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited optimized kernel and use a common wrapper. | Both | Existing positive drafting assertion retained; current-head CI pending. | | `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache UT | Both support `cache_hit_alignment_tokens`, introduced by [#53598](https://github.com/vllm-project/vllm/pull/53598/files), `2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only write-mask branch. | Both | Existing assertions retained; current-head CI pending. | #### Rebase and retired-fallback evidence | Change | Exact source evidence | Decision | |---|---|---| | `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM [#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95), `d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL before v0.29.0; both exact supported sources lack it. The old conditional came from vllm-ascend [#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0 selector and use the existing exclusion for both supported lanes. This does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy skip. | | `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and `tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM [#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29), `12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export `nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The v0.28.0 fallback originated in vllm-ascend [#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete import-failure/`None` fallbacks and the existing UT's obsolete availability skip; assertions remain unchanged and execute on both lanes. | | `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend [#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3), `799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream revert while resolving the real rebase conflict; do not reintroduce the reverted graph-update behavior. | | `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend [#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2), `d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged 310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation branch. | </details> </details> ### Does this PR introduce _any_ user-facing change? Yes. The supported vLLM release and default container builds move to v0.29.0, while the fixed main remains supported. The vllm-ascend package version does not change. Existing 0.29-only PCP+DP compatibility is restored. ### How was this patch tested? - Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`: [run 35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654) succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected NPU jobs confirm the requested Ascend head, integration base `2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM pins/installations: fixed main `84030bbe3d74d99bad477a3d2e37a973ccd8865c` / `0.1.dev1+g84030bbe3.empty`, release `98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest invocations. Skips/xfails are not passes. The resulting rebased integration commit is not printed and is not inferred. Actual release CPU and failed/cancelled image variants remain gaps. - Current-head CI update (2026-09-20): [main CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955) passed **5157 tests**, with **67 skipped**; actual installed vLLM was `0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head; the log does not print the full resulting integration head/base. Pre-commit and mypy passed. NPU/E2E results are recorded above; actual release CPU remains unverified. - Image build is partially blocked by infrastructure: A5 amd64 [Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005) and [openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068) failed before reading the Dockerfile because BuildKit could not create a snapshot temporary directory (`no space left on device`). No release compatibility code change is justified by this failure; cancelled variants remain unverified. The successful [310P openEuler arm64 build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987) explicitly checked out release `98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`; this is build evidence, not runtime UT coverage. - Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto the baseline above. Previous-head CPU [job 106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458) stopped during Ascend integration rebase after #16803 merged, before any UT ran. Its vLLM checkout was the fixed main and installation reported `0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved; fresh-head CI was retriggered by the push. - Remote official v0.29.0 tag resolved to the exact SHA above; both source pins inspected. - All 95 changed Python files pass Ruff lint, Ruff format and AST parsing; `git diff --check` passes. Dockerfile tag/marker consistency checked across all eight variants. - Full pre-commit invocation: Ruff, codespell, typos, clang-format, markdownlint, actionlint, package-init, forbidden-import and boolean-context checks passed. Bash-dependent hooks cannot run on this Windows host; Python launcher hooks exit 9009. Full CI lint remains pending. - Actual main/release CPU/NPU and image builds are not claimed locally: the required Linux/NPU/container environment is unavailable. New PR E2E and image-build CI are requested. Existing CPU workflow runs fixed main only, so actual release CPU remains a validation gap. - #16393's historical green CI is not a substitute for this new head. No new test functions, workflow changes, golden/threshold changes or additional skips are introduced beyond restoring #16393. Inherited 0.28-only tests/skips are not counted as passes. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…17033) ### What this PR does / why we need it? Fix duplicate MLA implementation weight post-processing on both supported vLLM pins: release `98dff2a81d747d1dba01a47f939f48c3526d4206` (0.29.0) and fixed main `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. Both upstream `MLAAttention.process_weights_after_loading` methods already dispatch to the impl. Ordinary MLA now uses that upstream call exactly once. SFA retains its direct impl-only path because it disposes `kv_b_proj`, which upstream dense packing would read. No version heuristics, persistent one-shot flag, or weight-retention workaround is introduced. Based on `c173a64a44dec4ba97aaba6277b1dfc1562eda19`, identical to the failing [nightly run](https://github.com/vllm-project/vllm-ascend/actions/runs/35529547840/job/106152304668). The earlier release adaptation PR #16393 is already merged. With MLAPO, a KV consumer, DCP disabled, a small token budget and default weight release on supported hardware, the first call frees the source projections and the second accesses `weight.data` on None. Decoder startup fails before performance testing; the later ready=2/6 timeout is a symptom. DFlash and TP4 are not established prerequisites. ### Does this PR introduce _any_ user-facing change? Restores the intended single post-load invocation while preserving MLAPO decoder weight release and SFA special handling. No timeout, deployment config, precision threshold, or memory-retention setting changes. ### How was this patch tested? - Six new regression cases cover MLA/SFA call counts for fp16 and bf16, plus producer/consumer fused processing. They use the actual upstream method and actual Ascend fused transformation/weight release; NPU format conversion and cache cleanup are mocked. - New and existing MLA wrapper tests: **25 passed** on each actual supported vLLM source pin (release 98dff2a and fixed main 84030bb), including the CPU mock fixture mode. - Negative control restoring the original c173a64 constructor: **4 failed, 2 passed**. Ordinary MLA detects duplicate calls; the consumer reproduces `mla_v1.py:1111` / `AttributeError: 'NoneType' object has no attribute 'data'`. SFA passes unchanged. - Current head `38db8c347e85f4856474c248723343d26a32d686`: [GitHub CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35544930945/job/106169524765) **5165 passed, 67 skipped**; pre-commit (including mypy) passed. Local Ruff 0.14 lint/format, Python compilation and whitespace checks passed. - **Real dual-A3 startup and PD smoke test completed** with vLLM 0.29.0 at 98dff2a and this production patch: original 2-prefiller/4-decoder topology, prefill DP2/TP8, decoder DP4/TP4, MLAPO=true, decoder max_num_batched_tokens=256, DCP disabled and default prefill-weight release. All six API servers became healthy. Two proxy chat requests returned HTTP 200 and each generated 32 tokens. No duplicate-postprocessing NoneType failure occurred. Deployment overrides were local model paths, host addresses/ports and compilation cache locations; compiled Ascend artifacts were reused after checking no build-source differences from their source revision to c173a64. - Smoke requests used max_tokens=32 and ended with `finish_reason=length`: this establishes startup and PD execution, not response accuracy or the 64k/1k performance target. Both requests routed through decoder rank 0; all four decoder ranks passed startup. During harness cleanup, sequential rank termination produced DP/Gloo connection-closed/EngineDead shutdown errors; test containers were stopped and all NPU processes released. No clean-shutdown or long-duration stability claim is made. - Latest `/nightly Kimi-K2.6-W4A8-64k-1k-TPOT50-PD` [comment](#17033 (comment)) dispatched [run 35545059509](https://github.com/vllm-project/vllm-ascend/actions/runs/35545059509), currently running. An earlier green run stopped during preparation with no benchmark output and is not counted as a model-test pass. - The initial CI gate required a test-enabling label despite passing CPU UT/pre-commit; `ready-precise` has now been added. Full precision CI, complete nightly performance/accuracy results and real SFA-model startup remain unverified. Full-model fixed-main startup was not run (fixed-main is covered by the tests above). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
…oject#16393) ### What this PR does / why we need it? Upgrade the supported vLLM release dependency from **v0.28.0 to the official v0.29.0**, alongside the fixed vLLM main commit inherited from vllm-project#16216 (already merged). | Reference | Commit | |---|---| | vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` | | PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` | | Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` | | Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a` | | Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` | - Use common implementations where v0.29.0 and the fixed main share KV-cache layouts, Mamba copy/group APIs, PCP handling and speculative-decoding contracts. - Retain explicit `vllm_version_is("0.29.0")` branches for contracts that still differ, including RoPE, scheduler block snapshots, InputBatch, ReplaySSM, KV zeroing and DSpark PP handling. - Remove obsolete v0.28.0 compatibility and adapt existing test fixtures. Version detection uses package versions and the explicit `VLLM_VERSION` override, with local-version suffix handling and UT environment isolation retained; no hard-coded release-SHA inference remains. - Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving other validations and EPLB platform binding. Rebuild dependent Pydantic schemas so nested configuration validation uses the patched validator. The global patch documentation records its rationale and removal criteria. - Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase; use the common exact-release selector and the real function-local DSpark sharing import. This release-only upgrade does not require a new main old-to-new interface scan. The latest rebase incorporates [vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files), which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The now-unnecessary V4.1 drafter import gate and tuple annotation have been removed. Other release adaptations, including vllm-project#15747 Spec+PP handling, remain. Detailed contract evidence is retained below for review. <details> <summary>Per-file adaptations and exact upstream evidence</summary> #### Per-file adaptation ledger Evidence IDs refer to the exact source contract and upstream diff table below. Every row is syntax checked; branch-normalized AST comparison confirms unchanged main function bodies except the version identity helper and the explicitly retired propose argument/type annotations. Both supported versions completed CPU and NPU execution as recorded below. | Ascend file / symbols | Disposition and upstream evidence | Verification | |---|---|---| | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` / `_load_dspark_model_with_target_quant`; `tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased [Ascend vllm-project#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a) (`82df9d871`), including manual PP partition masking. vLLM [#52809 diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share` inside the loader on both supported pins, so retain the earlier release fix: patch `eagle_utils`, never read/patch a nonexistent `dspark_utils._should_share`. v0.29 keeps its global PP guard; main has native PP via [#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks the real function-local import, absent module alias, and restoration on success/failure for both lanes; partition and PP assertions retained. Ruff/syntax pass; actual CPU/NPU pending. | | `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`; `tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29 routing and an obsolete 0.28 dev-build recognition path. For the two supported pins, #50514 exists only on fixed main. Route through `vllm_version_is("0.29.0")`; retain package/local-suffix and explicit environment-override semantics. No release-SHA inference or third release lane. | Existing UTs exercise the real uncached version helper with monkeypatch isolation, local suffix, fixed-main dev string, explicit override, and non-target versions. Isolated routing checks and Ruff/syntax pass; full CPU UT pending. | | `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`, `_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`, `_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436; use shared standardized layouts and retain the v0.29.0 InputBatch gate. The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while removing only legacy v0.28.0 allocation paths. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` | #52839; both lanes use the common PCP import. Rebase keeps current upstream graph code and removes only the v0.28.0 import branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/recompute_scheduler.py`<br>`module imports/dispatch`, `schedule` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`, `AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker` | #52615; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__` | #53614; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module imports/dispatch`, `_get_max_layers_per_page_size`, `_ascend_max_memory_usage_bytes_from_groups`, `_ascend_get_kv_cache_config_from_groups` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` | v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it after KV binding. Evidence: [#52506, `adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c). Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. | Failure reproduced on v0.29.0; Historical release/main NPU validation passed; current-head CI pending | | `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module imports/dispatch` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant` | #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports `_should_share` locally from Eagle utilities; fixed main removes the PP guard. Keep the release `get_pp_group` patch, share through the common Eagle utility, and delete the obsolete v0.28.0 `dspark_utils._should_share` patch. | Exact release failure reproduced; Historical release/main CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`, `__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`, `register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`, `_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`, `_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers | #51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and import the NaN helpers directly because v0.29.0 and fixed main expose the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static checked; historical dual-version CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes `pcp_manager` common to both lanes. The rebase preserves vllm-ascend vllm-project#16409's host-parameter-update revert and removes only the obsolete v0.28.0 capture branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`, `_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/block_table.py`<br>`__init__`, `init_block_table_layout_tensors`, `compute_slot_mappings` | #51718, #51031; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`, `prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`, `execute_model` | #50514, #54436, #52506, #55212, #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch` | #54282; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` | #49811; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module imports/dispatch`, `propose` | #53694, #52188; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`, `wake_up` | #51718, #53508; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | #### Exact upstream evidence | Upstream change | Full commit SHA / direct diff | Actual supported contracts and branch decision | |---|---|---| | #51718 | `8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Both use layers/layer_stride/block_stride/offset, tokens_per_state, CircularBufferSpec and standardized backing; retire shared_by/compress_ratio allocation branches. | | #52839 | `58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8) | Both import PCP operations from vllm.v1.attention.ops.pcp. | | #53896 | `e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec) | Both use Mamba copy-function dictionaries and unwrap UniformTypeKVCacheSpecs. | | #53106 | `1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21) | Both use WeightsMapper instead of AutoWeightsLoader skip_prefixes/skip_substrs. | | #53906 | `98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Only pinned main has the optional MLA storage_block_size dataclass field; release keeps the Ascend derived property, using tokens_per_state. | | #56078 | `719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417) | Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned main uses mrope_num_dims and unified RoPE. | | #52615 | `138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045) | Release uses num_blocks/kv_bytes_per_block; main uses num_chunks/kv_bytes_per_chunk. | | #51358 | `6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release now has boundary_state_offloads and KVConnectorBlockState; remove partial_tail_offloads plumbing. | | #54853 | `0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release constructor takes block_ids snapshots; main takes req_ids and resolve_block_ids. Keep exact release snapshot membership and main lazy-resolution membership. | | #53614 | `144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725) | Only main configures drop_eagle_checkpoint_block for replay-aligned Mamba checkpoints. | | #50514 | `d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Release retains module-level PP/share symbols and Ascend PP workaround; main has the subsequent PP integration. | | #52809 | `91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark module binding to a function-local import from Eagle utilities. The shared Eagle patch remains effective; the old DSpark-module read/write must be removed. | | #54436 | `6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902) | Release InputBatch requires max_seq_len_np; main removed it. | | #52506 | `adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Only main accepts valid_dummy_state_slots/valid_state_slots capture arguments. | | #55212 | `83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Release prepares DCP local sequence lengths before partitioning; main initializes DCP metadata afterwards. | | #53515 | `b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127) | Both accept padded_num_tokens for persistent PCP input buffers. | | #53869 | `b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212) | Both accept pcp_manager during graph capture. | | #51031 | `0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1) | Both distinguish KV and kernel block sizes during DCP slot mapping. | | #54282 | `fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8) | Both gumbel sampling APIs include is_drafting. | | #52188 | `d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0) | Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. | | #53694 | `5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6) | Both propose APIs take DPSyncState; remove the obsolete token-count argument and retain replicated-PCP synchronization. | | #49811 | `01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e) | Both support extract_hidden_states on MRV2; remove old unsupported dispatch/skip. | | #53508 | `479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29) | Both remove post_kv_cache_wake_up; retire release-only call. | | #52494 | `3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1) | Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain release exclusion. | | #52861 | `b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28) | Both include DeepseekV32MTPModel in the two-hidden-state architecture set. | | #54713 | `b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3) | Only main takes replay_boundaries in compressed-prefix hit lookup; preserve release calls without that keyword. | | #42785 | `442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18) | Only main capture callers pass axis_keys; preserve the existing Ascend rejection of nonempty axes. | | #52358 | `8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Both ExecuteModelState have dp_sync; only main has cudagraph_stats. | | #52789 | `9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4) | Release already has mamba_has_prefill_checkpoint_blocks; later main also has fine-grained prefix-cache state. | | #51251 | `7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | v0.29.0 and main expose ec_manager_config; retire the old release-only ScoreEncoder configuration skip. | | #53240 | `b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c) | v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache groups; retire the old release-only replay skip. | | #53853 | `e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both delegate PCP compatibility validation to the PCP manager. | | #53183 | `4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both expose the V1 unsupported-feature helper used by the existing Ascend MRV1 feature filter. | #### Additional inherited contracts | vllm-ascend change | Why it is required | Upstream cause and direct link | Lane | Verification | |---|---|---|---|---| | `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec, list[int]]` from `get_mamba_groups` and both initialize `recoverssm`; keeping the old fallback would preserve an unsupported third contract | [v0.29.0 mamba groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704), [fixed-main mamba groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704), [v0.29.0 RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103), [fixed-main RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103) | both, common implementation | Existing constructor UT now asserts the parent-created RecoverSSM value is retained; Ruff and compileall pass | | `worker/v2/model_runner.py`: always forward `kv_cache_allocation_context` | v0.29.0 and fixed main both accept this keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0 signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538), [fixed-main signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566), [vllm-ascend vllm-project#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) | both, common implementation | Existing UT continues to assert the exact context object reaches the parent; Ruff and compileall pass | | `_310p/worker/v2/model_runner.py`: select the release KV-zeroing contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes `KVCacheGroupSpec.is_eagle_group` but lacks `SpeculativeConfig.use_eagle_block_drop`; fixed main added the method | vLLM [#53388 diff](https://github.com/vllm-project/vllm/pull/53388/files), commit [`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a); [v0.29.0 group field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200), [fixed-main helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876) | release differs from main | Existing two-path KV-zeroing UT retained and renamed for v0.29.0; Ruff and compileall pass | | existing DFlash kernel UT: remove v0.28-only kwarg omission | the current Ascend kernel accepts the CP arguments and the only excluded lane was v0.28.0, which this PR replaces | [vllm-ascend vllm-project#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) | both, common invocation | Existing NPU test remains enabled with all assertions; Current-head CI pending | | Ascend change | Why / upstream cause | Lane | Verification | |---|---|---|---| | `patch/platform/patch_parallel_config.py`, registration, and global patch documentation | Allow Ascend PCP+DP by removing the generic GPU restriction, following vLLM [#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic `parallel_config.current_platform` lookup and rebuild ParallelConfig → SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing configuration/EPLB UTs and historical PCP+DP NPU execution passed; current-head CI pending. | | `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT | Both supported contracts require `is_drafting`, from [#54282](https://github.com/vllm-project/vllm/pull/54282/files), `fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited optimized kernel and use a common wrapper. | Both | Existing positive drafting assertion retained; current-head CI pending. | | `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache UT | Both support `cache_hit_alignment_tokens`, introduced by [#53598](https://github.com/vllm-project/vllm/pull/53598/files), `2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only write-mask branch. | Both | Existing assertions retained; current-head CI pending. | #### Rebase and retired-fallback evidence | Change | Exact source evidence | Decision | |---|---|---| | `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM [#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95), `d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL before v0.29.0; both exact supported sources lack it. The old conditional came from vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0 selector and use the existing exclusion for both supported lanes. This does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy skip. | | `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and `tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM [#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29), `12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export `nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The v0.28.0 fallback originated in vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete import-failure/`None` fallbacks and the existing UT's obsolete availability skip; assertions remain unchanged and execute on both lanes. | | `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend [vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3), `799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream revert while resolving the real rebase conflict; do not reintroduce the reverted graph-update behavior. | | `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend [vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2), `d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged 310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation branch. | </details> ### Does this PR introduce _any_ user-facing change? Yes. The supported release changes from vLLM v0.28.0 to v0.29.0, and Ascend PCP+DP is enabled on v0.29.0. The fixed main commit remains supported. The vllm-ascend package version is unchanged. Obsolete v0.28-only skips are removed where the contracts are now supported. No skip is migrated to v0.29.0, and no golden data or precision threshold is changed. Existing unrelated exclusions remain exclusions, not passes. ### How was this patch tested? Current-head validation: [run 35420046256](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256) completed successfully on `fd876c5269fb6b39afa04e5a52433505e8786e22`: 39 successful jobs and 6 skipped jobs, no failures. - Pre-commit and mypy passed. Fixed-main [CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105836507531): **5078 passed, 67 skipped, 17 warnings**. Actual vLLM checkout `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, installed `0.1.dev1+g84030bbe3.empty`; Ascend checkout is the PR head. CPU `base_sha` is empty; the log says HEAD is up to date, so no unprinted integration SHA is inferred. - All **32 NPU jobs passed**, 16 per lane. Raw-log session totals per lane: **562 passed, 35 skipped, 1 xfailed** (execution counts, not deduplicated cases). Actual checkouts: main `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, release `98dff2a81d747d1dba01a47f939f48c3526d4206`; installed versions `0.1.dev1+g84030bbe3.empty` and `0.29.0+empty`. Each log confirms Ascend head `fd876c5269fb6b39afa04e5a52433505e8786e22`, integration base `e139b7d573d3769fd1407d5027d7d4831f5469f0`, and HEAD is up to date. - PCP+DP and PCP+PP+MTP suite: [main A3 four-card part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105845372541) and [release A3 four-card part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35420046256/job/105845372595) both passed. This does not establish a root cause for historical HCCL failures. - Skipped and expected-failure cases are not passes. **Current-head release CPU verification remains missing**: the existing CPU job runs fixed main only. Historical release CPU results do not fill this gap. No workflow changes or additional skips were introduced. Latest rebase: head `fd876c5269fb6b39afa04e5a52433505e8786e22` onto upstream main `e139b7d573d3769fd1407d5027d7d4831f5469f0`. The old-head green results below are historical; new-head CI has completed successfully; see the verified results below. The fixed vLLM main and release pins remain unchanged. PR-wide Ruff, formatting and Python AST checks passed for 95 Python files; no workflow changes. Rebase reconciliation: - Preserve upstream [vllm-project#16853](https://github.com/vllm-project/vllm-ascend/pull/16853/files), `8f2e3fed73328148ea603ccfe5764542fa7cbbd2`, without migrating its 0.28-only validator/dispatch gates to 0.29. Both supported pins already call the PCP manager for dispatch token counts. Restore the version-helper import needed by the inherited method. Its inherited 0.28-only skipped tests are not counted as passes. - Adapt the existing unsupported-feature UT to preserve the upstream list: both supported pins delegate PCP checks to the manager via vLLM [#53853](https://github.com/vllm-project/vllm/pull/53853/files), `e376d45e82cb7e220da430e3179e81eb0922cf56`; no obsolete 0.28 string filtering is reinstated. - Adapt the existing KVPP allocation-entry UT from [vllm-project#15514](https://github.com/vllm-project/vllm-ascend/pull/15514/files), `b64959a11ea69a51409f64e4e68da8e0924dec42`, to the common allocation entry already used by both supported pins after vLLM #51718. Preserve cache-view assertions and the upstream dtype changes. No test functions or skips added. - vllm-project#16544 remains reverted by vllm-project#16905; its V4.1 compatibility workaround remains removed. Historical validation before this rebase: - **Current head `8c5d785931c495701bf1da8b5bc80b321ebd3cd0`:** [run 35350119603](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603) completed successfully: **40 successful jobs, 6 skipped jobs, no failures**. Pre-commit and mypy passed. Local Ruff lint/format and AST checks passed for all 95 changed Python files. - **Actual main CPU:** [job 105617611031](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105617611031) reports **5074 passed, 37 skipped**, 17 warnings. Logs confirm the exact PR head, vLLM checkout `84030bbe3d74d99bad477a3d2e37a973ccd8865c`, and installed `0.1.dev1+g84030bbe3.empty`. CPU `base_sha` is empty; no integration merge is inferred. - **Actual dual-version NPU:** All 32 NPU job logs were audited, 16 per lane. Fixed main installed `0.1.dev1+g84030bbe3.empty`; official release checkout `98dff2a81d747d1dba01a47f939f48c3526d4206` installed `0.29.0+empty`. Each lane reports **562 passed, 35 skipped, 1 xfailed** across pytest session summaries (execution counts, not deduplicated unique tests). Logs confirm checkout of this PR head and integration base `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`; the rebase step reports HEAD is up to date. - **PCP combinations:** The context-parallel suite containing PCP+DP and PCP+PP+MTP passes on both [main A3 part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105658864889) and [release A3 part4](https://github.com/vllm-project/vllm-ascend/actions/runs/35350119603/job/105658864659). - **Release CPU gap:** Current-head actual release CPU remains missing because the workflow has only a main CPU entry. Version-parameterized UTs and historical release CPU runs do not replace actual release installation. Historical run 35213834751 at head `2560da9c7c2628f1b92b9305a1e80c5856bce74f` passed 4923 tests with 37 skips per lane; its temporary CPU workflow was reverted. - **Backup:** vllm-project#16898 remains unchanged with no test CI triggered. Local branch `codex/backup-16393-before-16905` preserves previous green `7ad4c82ee`; its results are not used as current-head validation. Skipped/xfail cases and jobs are not counted as passes. Earlier intermittent HCCL errors have no proven root cause; no speculative fix is included. This run passes existing CI, but complete dual-version CPU validation remains outstanding. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…#16393) (vllm-project#16949) ### What this PR does / why we need it? Revert vllm-project#16393 at the author's request. This reverses squash commit `af277b2681443089e4d29684b1101707e2eb6426` on upstream main `8038f64129a1f8751ac6684f47c8d058208e4ff3`, restoring the previous fixed-main + v0.28.0 support contract. The revert applies without conflicts. Subsequent upstream patches are retained, including vllm-project#16775 (PCP/DCP slots), vllm-project#16923 (A3 SFA prefill all-to-all), and vllm-project#16834 (PCP hash-routing input IDs). No unrelated fixes are included. | Reverted change | Reason and source | Lane | | --- | --- | --- | | Release marker | Reverse [vllm-project#16393](https://github.com/vllm-project/vllm-ascend/pull/16393/files): v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`) to v0.28.0 (`2cf0a6915ce544dc493a0990f2ea38d81601128a`) | Release | | Version branches and interface adapters across scheduler, KV cache, attention, model runners and speculative decoding | Restore the pre-vllm-project#16393 v0.28.0 contracts by reversing the exact merged diff; keep main implementations behind their original version selection | Main + release | | v0.29.0-only parallel-config patch, registration and documentation | Remove the patch introduced by vllm-project#16393; retain the independent upstream v0.28.0 PCP+DP workaround | Release | | Existing unit/e2e tests and release-only skips | Restore the tests and conditions changed by vllm-project#16393, without introducing additional skip scope beyond the reverted baseline | Main + release | The vLLM main pin remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. No main interface scan is needed for this release-only revert. The only changes beyond the inverse patch are formatting in `tests/ut/test_utils.py` and line-ending normalization in `categorical_sample.py` needed for formatting/whitespace checks. Workflows are unchanged. ### Does this PR introduce _any_ user-facing change? Yes. The supported release returns to vLLM v0.28.0; v0.29.0 support and its specific compatibility changes from vllm-project#16393 are reverted. The vllm-ascend package version and fixed vLLM main pin are unchanged. ### How was this patch tested? - `git revert --no-commit af277b2` applied cleanly on the latest fetched main. - Verified that the complete post-vllm-project#16393 upstream patch can still reverse-apply to the staged result (`git apply --reverse --check --cached`), confirming subsequent changes are preserved. - `git diff --cached --check`, Python AST parsing, Ruff lint and formatting checks passed for all 94 remaining changed Python files. - Ran the equivalent of `format.sh ci`: `pre-commit run --all-files --hook-stage manual`. Ruff, codespell, typos, clang-format, markdownlint, actionlint, package-init, forbidden-import and boolean-context checks passed. Bash-based hooks could not execute on this Windows host; two Python launcher hooks exited 9009. Full pre-commit validation remains for CI. - Actual CPU/NPU tests have not been run locally because the required vLLM/NPU runtime is unavailable. New PR CI is pending; the original upgrade PR's green run is not validation of this revert. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…vllm-project#17004) ### What this PR does / why we need it? Reintroduce the vLLM v0.29.0 release upgrade from vllm-project#16393, reverted by vllm-project#16949, with container defaults aligned to the supported release. Based on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves later merged changes. The image/source mismatch is confirmed in [nightly job 106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020): the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded Ascend failed importing `_get_packed_kv_cache_groups`. All eight root Dockerfiles still defaulted to v0.28.0. The reusable image workflow passes no `VLLM_TAG` override, so those defaults govern release builds. Changing only the release marker does not update those images. - Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. - Release changes from v0.28.0 (`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the remote tag). - Restore vllm-project#16393's source compatibility and existing UT changes by reversing vllm-project#16949, then reconcile current upstream changes. No main interface scan is rerun for this release-only upgrade. - Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P, Ubuntu/openEuler. Preserve existing exact-commit build overrides. No workflow changes. | Current-base adaptation | Exact cause and evidence | Lane | Validation | | --- | --- | --- | --- | | Eight Dockerfile `VLLM_TAG` defaults and release marker | vllm-project#16393 changed the release contract but omitted image defaults; the linked nightly log proves installation of 0.28.0 and missing `_get_packed_kv_cache_groups`. | Release images | All eight defaults match marker; image-build CI requested, pending | | `worker/v2/attn_utils.py::get_kv_cache_spec` and existing `_make_mla_layer` UT fixture | Ascend [vllm-project#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files), `df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field selector. vLLM [#51718](https://github.com/vllm-project/vllm/pull/51718/files), `8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio` with `tokens_per_state`; both exact supported pins use the latter. Use the common field, retaining metadata and cache-view assertions. | Both | Source inspection and static checks passed; actual CI pending | | `attention/attention_v1.py` import conflict | Preserve `attention_transfer_window` from Ascend [vllm-project#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files), `5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from vllm-project#16798; keep the common relocated PCP import from vllm-project#16393. Do not restore the old compute-start import or unused weak-reference import. | Both | Conflict resolved; static checks passed | | `worker/v2/aclgraph_utils.py` import conflict | Preserve `ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend [vllm-project#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files), `34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete 0.28 version selector restored by the revert. | Both | Conflict resolved; static checks passed | | Existing 310P and Mamba model-runner UT imports | Preserve hardware-profile imports and mocks from Ascend [vllm-project#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files), `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused old-release selector import. Hardware capability routing remains unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format passed; fresh CI pending | Newly merged changes were reviewed for version-contract impact: vllm-project#16775/vllm-project#16834 PCP metadata and routing, vllm-project#16923 A3 SFA, vllm-project#16426/vllm-project#16924/vllm-project#16955 custom ops, vllm-project#16081 MTP/SP, vllm-project#16669 xlite, vllm-project#15636/vllm-project#16747 transfer, vllm-project#16913 A5 pages, vllm-project#16798 graph updates, vllm-project#16755 GLM PD, vllm-project#16952 C8 config, vllm-project#16673 operator removal and vllm-project#16320 Kimi-K3 KV pool. Preserve these changes; no additional source-proven version branch was identified beyond the entries above. In particular, vllm-project#16747's `UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size` exist in both exact pins; vllm-project#16320 adds Ascend connector hooks. CI remains necessary to validate runtime interactions. CI/documentation-only PRs are retained unchanged. <details> <summary>Inherited per-file compatibility evidence from vllm-project#16393</summary> The following source-contract ledger is inherited from vllm-project#16393. Any historical verification wording refers only to that earlier PR; it does not certify this new head. New-head validation is listed below. | Reference | Commit | |---|---| | vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` | | PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` | | Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` | | Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a` | | Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` | - Use common implementations where v0.29.0 and the fixed main share KV-cache layouts, Mamba copy/group APIs, PCP handling and speculative-decoding contracts. - Retain explicit `vllm_version_is("0.29.0")` branches for contracts that still differ, including RoPE, scheduler block snapshots, InputBatch, ReplaySSM, KV zeroing and DSpark PP handling. - Remove obsolete v0.28.0 compatibility and adapt existing test fixtures. Version detection uses package versions and the explicit `VLLM_VERSION` override, with local-version suffix handling and UT environment isolation retained; no hard-coded release-SHA inference remains. - Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving other validations and EPLB platform binding. Rebuild dependent Pydantic schemas so nested configuration validation uses the patched validator. The global patch documentation records its rationale and removal criteria. - Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase; use the common exact-release selector and the real function-local DSpark sharing import. This release-only upgrade does not require a new main old-to-new interface scan. The latest rebase incorporates [vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files), which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The now-unnecessary V4.1 drafter import gate and tuple annotation have been removed. Other release adaptations, including vllm-project#15747 Spec+PP handling, remain. Detailed contract evidence is retained below for review. <details> <summary>Per-file adaptations and exact upstream evidence</summary> #### Per-file adaptation ledger Evidence IDs refer to the exact source contract and upstream diff table below. Every row is syntax checked; branch-normalized AST comparison confirms unchanged main function bodies except the version identity helper and the explicitly retired propose argument/type annotations. Both supported versions completed CPU and NPU execution as recorded below. | Ascend file / symbols | Disposition and upstream evidence | Verification | |---|---|---| | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` / `_load_dspark_model_with_target_quant`; `tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased [Ascend vllm-project#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a) (`82df9d871`), including manual PP partition masking. vLLM [#52809 diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share` inside the loader on both supported pins, so retain the earlier release fix: patch `eagle_utils`, never read/patch a nonexistent `dspark_utils._should_share`. v0.29 keeps its global PP guard; main has native PP via [#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks the real function-local import, absent module alias, and restoration on success/failure for both lanes; partition and PP assertions retained. Ruff/syntax pass; actual CPU/NPU pending. | | `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`; `tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29 routing and an obsolete 0.28 dev-build recognition path. For the two supported pins, #50514 exists only on fixed main. Route through `vllm_version_is("0.29.0")`; retain package/local-suffix and explicit environment-override semantics. No release-SHA inference or third release lane. | Existing UTs exercise the real uncached version helper with monkeypatch isolation, local suffix, fixed-main dev string, explicit override, and non-target versions. Isolated routing checks and Ruff/syntax pass; full CPU UT pending. | | `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`, `_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`, `_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436; use shared standardized layouts and retain the v0.29.0 InputBatch gate. The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while removing only legacy v0.28.0 allocation paths. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` | #52839; both lanes use the common PCP import. Rebase keeps current upstream graph code and removes only the v0.28.0 import branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/recompute_scheduler.py`<br>`module imports/dispatch`, `schedule` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`, `AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker` | #52615; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__` | #53614; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module imports/dispatch`, `_get_max_layers_per_page_size`, `_ascend_max_memory_usage_bytes_from_groups`, `_ascend_get_kv_cache_config_from_groups` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` | v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it after KV binding. Evidence: [#52506, `adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c). Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. | Failure reproduced on v0.29.0; Historical release/main NPU validation passed; current-head CI pending | | `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module imports/dispatch` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant` | #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports `_should_share` locally from Eagle utilities; fixed main removes the PP guard. Keep the release `get_pp_group` patch, share through the common Eagle utility, and delete the obsolete v0.28.0 `dspark_utils._should_share` patch. | Exact release failure reproduced; Historical release/main CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`, `__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`, `register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`, `_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`, `_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers | #51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and import the NaN helpers directly because v0.29.0 and fixed main expose the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static checked; historical dual-version CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes `pcp_manager` common to both lanes. The rebase preserves vllm-ascend vllm-project#16409's host-parameter-update revert and removes only the obsolete v0.28.0 capture branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`, `_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/block_table.py`<br>`__init__`, `init_block_table_layout_tensors`, `compute_slot_mappings` | #51718, #51031; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`, `prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`, `execute_model` | #50514, #54436, #52506, #55212, #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch` | #54282; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` | #49811; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module imports/dispatch`, `propose` | #53694, #52188; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`, `wake_up` | #51718, #53508; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | #### Exact upstream evidence | Upstream change | Full commit SHA / direct diff | Actual supported contracts and branch decision | |---|---|---| | #51718 | `8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Both use layers/layer_stride/block_stride/offset, tokens_per_state, CircularBufferSpec and standardized backing; retire shared_by/compress_ratio allocation branches. | | #52839 | `58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8) | Both import PCP operations from vllm.v1.attention.ops.pcp. | | #53896 | `e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec) | Both use Mamba copy-function dictionaries and unwrap UniformTypeKVCacheSpecs. | | #53106 | `1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21) | Both use WeightsMapper instead of AutoWeightsLoader skip_prefixes/skip_substrs. | | #53906 | `98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Only pinned main has the optional MLA storage_block_size dataclass field; release keeps the Ascend derived property, using tokens_per_state. | | #56078 | `719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417) | Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned main uses mrope_num_dims and unified RoPE. | | #52615 | `138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045) | Release uses num_blocks/kv_bytes_per_block; main uses num_chunks/kv_bytes_per_chunk. | | #51358 | `6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release now has boundary_state_offloads and KVConnectorBlockState; remove partial_tail_offloads plumbing. | | #54853 | `0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release constructor takes block_ids snapshots; main takes req_ids and resolve_block_ids. Keep exact release snapshot membership and main lazy-resolution membership. | | #53614 | `144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725) | Only main configures drop_eagle_checkpoint_block for replay-aligned Mamba checkpoints. | | #50514 | `d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Release retains module-level PP/share symbols and Ascend PP workaround; main has the subsequent PP integration. | | #52809 | `91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark module binding to a function-local import from Eagle utilities. The shared Eagle patch remains effective; the old DSpark-module read/write must be removed. | | #54436 | `6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902) | Release InputBatch requires max_seq_len_np; main removed it. | | #52506 | `adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Only main accepts valid_dummy_state_slots/valid_state_slots capture arguments. | | #55212 | `83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Release prepares DCP local sequence lengths before partitioning; main initializes DCP metadata afterwards. | | #53515 | `b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127) | Both accept padded_num_tokens for persistent PCP input buffers. | | #53869 | `b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212) | Both accept pcp_manager during graph capture. | | #51031 | `0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1) | Both distinguish KV and kernel block sizes during DCP slot mapping. | | #54282 | `fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8) | Both gumbel sampling APIs include is_drafting. | | #52188 | `d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0) | Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. | | #53694 | `5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6) | Both propose APIs take DPSyncState; remove the obsolete token-count argument and retain replicated-PCP synchronization. | | #49811 | `01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e) | Both support extract_hidden_states on MRV2; remove old unsupported dispatch/skip. | | #53508 | `479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29) | Both remove post_kv_cache_wake_up; retire release-only call. | | #52494 | `3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1) | Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain release exclusion. | | #52861 | `b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28) | Both include DeepseekV32MTPModel in the two-hidden-state architecture set. | | #54713 | `b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3) | Only main takes replay_boundaries in compressed-prefix hit lookup; preserve release calls without that keyword. | | #42785 | `442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18) | Only main capture callers pass axis_keys; preserve the existing Ascend rejection of nonempty axes. | | #52358 | `8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Both ExecuteModelState have dp_sync; only main has cudagraph_stats. | | #52789 | `9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4) | Release already has mamba_has_prefill_checkpoint_blocks; later main also has fine-grained prefix-cache state. | | #51251 | `7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | v0.29.0 and main expose ec_manager_config; retire the old release-only ScoreEncoder configuration skip. | | #53240 | `b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c) | v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache groups; retire the old release-only replay skip. | | #53853 | `e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both delegate PCP compatibility validation to the PCP manager. | | #53183 | `4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both expose the V1 unsupported-feature helper used by the existing Ascend MRV1 feature filter. | #### Additional inherited contracts | vllm-ascend change | Why it is required | Upstream cause and direct link | Lane | Verification | |---|---|---|---|---| | `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec, list[int]]` from `get_mamba_groups` and both initialize `recoverssm`; keeping the old fallback would preserve an unsupported third contract | [v0.29.0 mamba groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704), [fixed-main mamba groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704), [v0.29.0 RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103), [fixed-main RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103) | both, common implementation | Existing constructor UT now asserts the parent-created RecoverSSM value is retained; Ruff and compileall pass | | `worker/v2/model_runner.py`: always forward `kv_cache_allocation_context` | v0.29.0 and fixed main both accept this keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0 signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538), [fixed-main signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566), [vllm-ascend vllm-project#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) | both, common implementation | Existing UT continues to assert the exact context object reaches the parent; Ruff and compileall pass | | `_310p/worker/v2/model_runner.py`: select the release KV-zeroing contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes `KVCacheGroupSpec.is_eagle_group` but lacks `SpeculativeConfig.use_eagle_block_drop`; fixed main added the method | vLLM [#53388 diff](https://github.com/vllm-project/vllm/pull/53388/files), commit [`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a); [v0.29.0 group field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200), [fixed-main helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876) | release differs from main | Existing two-path KV-zeroing UT retained and renamed for v0.29.0; Ruff and compileall pass | | existing DFlash kernel UT: remove v0.28-only kwarg omission | the current Ascend kernel accepts the CP arguments and the only excluded lane was v0.28.0, which this PR replaces | [vllm-ascend vllm-project#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) | both, common invocation | Existing NPU test remains enabled with all assertions; Current-head CI pending | | Ascend change | Why / upstream cause | Lane | Verification | |---|---|---|---| | `patch/platform/patch_parallel_config.py`, registration, and global patch documentation | Allow Ascend PCP+DP by removing the generic GPU restriction, following vLLM [#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic `parallel_config.current_platform` lookup and rebuild ParallelConfig → SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing configuration/EPLB UTs and historical PCP+DP NPU execution passed; current-head CI pending. | | `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT | Both supported contracts require `is_drafting`, from [#54282](https://github.com/vllm-project/vllm/pull/54282/files), `fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited optimized kernel and use a common wrapper. | Both | Existing positive drafting assertion retained; current-head CI pending. | | `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache UT | Both support `cache_hit_alignment_tokens`, introduced by [#53598](https://github.com/vllm-project/vllm/pull/53598/files), `2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only write-mask branch. | Both | Existing assertions retained; current-head CI pending. | #### Rebase and retired-fallback evidence | Change | Exact source evidence | Decision | |---|---|---| | `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM [#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95), `d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL before v0.29.0; both exact supported sources lack it. The old conditional came from vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0 selector and use the existing exclusion for both supported lanes. This does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy skip. | | `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and `tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM [#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29), `12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export `nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The v0.28.0 fallback originated in vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete import-failure/`None` fallbacks and the existing UT's obsolete availability skip; assertions remain unchanged and execute on both lanes. | | `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend [vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3), `799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream revert while resolving the real rebase conflict; do not reintroduce the reverted graph-update behavior. | | `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend [vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2), `d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged 310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation branch. | </details> </details> ### Does this PR introduce _any_ user-facing change? Yes. The supported vLLM release and default container builds move to v0.29.0, while the fixed main remains supported. The vllm-ascend package version does not change. Existing 0.29-only PCP+DP compatibility is restored. ### How was this patch tested? - Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`: [run 35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654) succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected NPU jobs confirm the requested Ascend head, integration base `2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM pins/installations: fixed main `84030bbe3d74d99bad477a3d2e37a973ccd8865c` / `0.1.dev1+g84030bbe3.empty`, release `98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest invocations. Skips/xfails are not passes. The resulting rebased integration commit is not printed and is not inferred. Actual release CPU and failed/cancelled image variants remain gaps. - Current-head CI update (2026-09-20): [main CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955) passed **5157 tests**, with **67 skipped**; actual installed vLLM was `0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head; the log does not print the full resulting integration head/base. Pre-commit and mypy passed. NPU/E2E results are recorded above; actual release CPU remains unverified. - Image build is partially blocked by infrastructure: A5 amd64 [Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005) and [openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068) failed before reading the Dockerfile because BuildKit could not create a snapshot temporary directory (`no space left on device`). No release compatibility code change is justified by this failure; cancelled variants remain unverified. The successful [310P openEuler arm64 build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987) explicitly checked out release `98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`; this is build evidence, not runtime UT coverage. - Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto the baseline above. Previous-head CPU [job 106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458) stopped during Ascend integration rebase after vllm-project#16803 merged, before any UT ran. Its vLLM checkout was the fixed main and installation reported `0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved; fresh-head CI was retriggered by the push. - Remote official v0.29.0 tag resolved to the exact SHA above; both source pins inspected. - All 95 changed Python files pass Ruff lint, Ruff format and AST parsing; `git diff --check` passes. Dockerfile tag/marker consistency checked across all eight variants. - Full pre-commit invocation: Ruff, codespell, typos, clang-format, markdownlint, actionlint, package-init, forbidden-import and boolean-context checks passed. Bash-dependent hooks cannot run on this Windows host; Python launcher hooks exit 9009. Full CI lint remains pending. - Actual main/release CPU/NPU and image builds are not claimed locally: the required Linux/NPU/container environment is unavailable. New PR E2E and image-build CI are requested. Existing CPU workflow runs fixed main only, so actual release CPU remains a validation gap. - vllm-project#16393's historical green CI is not a substitute for this new head. No new test functions, workflow changes, golden/threshold changes or additional skips are introduced beyond restoring vllm-project#16393. Inherited 0.28-only tests/skips are not counted as passes. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…llm-project#17033) ### What this PR does / why we need it? Fix duplicate MLA implementation weight post-processing on both supported vLLM pins: release `98dff2a81d747d1dba01a47f939f48c3526d4206` (0.29.0) and fixed main `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. Both upstream `MLAAttention.process_weights_after_loading` methods already dispatch to the impl. Ordinary MLA now uses that upstream call exactly once. SFA retains its direct impl-only path because it disposes `kv_b_proj`, which upstream dense packing would read. No version heuristics, persistent one-shot flag, or weight-retention workaround is introduced. Based on `c173a64a44dec4ba97aaba6277b1dfc1562eda19`, identical to the failing [nightly run](https://github.com/vllm-project/vllm-ascend/actions/runs/35529547840/job/106152304668). The earlier release adaptation PR vllm-project#16393 is already merged. With MLAPO, a KV consumer, DCP disabled, a small token budget and default weight release on supported hardware, the first call frees the source projections and the second accesses `weight.data` on None. Decoder startup fails before performance testing; the later ready=2/6 timeout is a symptom. DFlash and TP4 are not established prerequisites. ### Does this PR introduce _any_ user-facing change? Restores the intended single post-load invocation while preserving MLAPO decoder weight release and SFA special handling. No timeout, deployment config, precision threshold, or memory-retention setting changes. ### How was this patch tested? - Six new regression cases cover MLA/SFA call counts for fp16 and bf16, plus producer/consumer fused processing. They use the actual upstream method and actual Ascend fused transformation/weight release; NPU format conversion and cache cleanup are mocked. - New and existing MLA wrapper tests: **25 passed** on each actual supported vLLM source pin (release 98dff2a and fixed main 84030bb), including the CPU mock fixture mode. - Negative control restoring the original c173a64 constructor: **4 failed, 2 passed**. Ordinary MLA detects duplicate calls; the consumer reproduces `mla_v1.py:1111` / `AttributeError: 'NoneType' object has no attribute 'data'`. SFA passes unchanged. - Current head `38db8c347e85f4856474c248723343d26a32d686`: [GitHub CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35544930945/job/106169524765) **5165 passed, 67 skipped**; pre-commit (including mypy) passed. Local Ruff 0.14 lint/format, Python compilation and whitespace checks passed. - **Real dual-A3 startup and PD smoke test completed** with vLLM 0.29.0 at 98dff2a and this production patch: original 2-prefiller/4-decoder topology, prefill DP2/TP8, decoder DP4/TP4, MLAPO=true, decoder max_num_batched_tokens=256, DCP disabled and default prefill-weight release. All six API servers became healthy. Two proxy chat requests returned HTTP 200 and each generated 32 tokens. No duplicate-postprocessing NoneType failure occurred. Deployment overrides were local model paths, host addresses/ports and compilation cache locations; compiled Ascend artifacts were reused after checking no build-source differences from their source revision to c173a64. - Smoke requests used max_tokens=32 and ended with `finish_reason=length`: this establishes startup and PD execution, not response accuracy or the 64k/1k performance target. Both requests routed through decoder rank 0; all four decoder ranks passed startup. During harness cleanup, sequential rank termination produced DP/Gloo connection-closed/EngineDead shutdown errors; test containers were stopped and all NPU processes released. No clean-shutdown or long-duration stability claim is made. - Latest `/nightly Kimi-K2.6-W4A8-64k-1k-TPOT50-PD` [comment](vllm-project#17033 (comment)) dispatched [run 35545059509](https://github.com/vllm-project/vllm-ascend/actions/runs/35545059509), currently running. An earlier green run stopped during preparation with no benchmark output and is not counted as a model-test pass. - The initial CI gate required a test-enabling label despite passing CPU UT/pre-commit; `ready-precise` has now been added. Full precision CI, complete nightly performance/accuracy results and real SFA-model startup remain unverified. Full-model fixed-main startup was not run (fixed-main is covered by the tests above). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
…#16393) (vllm-project#16949) Revert vllm-project#16393 at the author's request. This reverses squash commit `af277b2681443089e4d29684b1101707e2eb6426` on upstream main `8038f64129a1f8751ac6684f47c8d058208e4ff3`, restoring the previous fixed-main + v0.28.0 support contract. The revert applies without conflicts. Subsequent upstream patches are retained, including vllm-project#16775 (PCP/DCP slots), vllm-project#16923 (A3 SFA prefill all-to-all), and vllm-project#16834 (PCP hash-routing input IDs). No unrelated fixes are included. | Reverted change | Reason and source | Lane | | --- | --- | --- | | Release marker | Reverse [vllm-project#16393](https://github.com/vllm-project/vllm-ascend/pull/16393/files): v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`) to v0.28.0 (`2cf0a6915ce544dc493a0990f2ea38d81601128a`) | Release | | Version branches and interface adapters across scheduler, KV cache, attention, model runners and speculative decoding | Restore the pre-vllm-project#16393 v0.28.0 contracts by reversing the exact merged diff; keep main implementations behind their original version selection | Main + release | | v0.29.0-only parallel-config patch, registration and documentation | Remove the patch introduced by vllm-project#16393; retain the independent upstream v0.28.0 PCP+DP workaround | Release | | Existing unit/e2e tests and release-only skips | Restore the tests and conditions changed by vllm-project#16393, without introducing additional skip scope beyond the reverted baseline | Main + release | The vLLM main pin remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. No main interface scan is needed for this release-only revert. The only changes beyond the inverse patch are formatting in `tests/ut/test_utils.py` and line-ending normalization in `categorical_sample.py` needed for formatting/whitespace checks. Workflows are unchanged. Yes. The supported release returns to vLLM v0.28.0; v0.29.0 support and its specific compatibility changes from vllm-project#16393 are reverted. The vllm-ascend package version and fixed vLLM main pin are unchanged. - `git revert --no-commit af277b2` applied cleanly on the latest fetched main. - Verified that the complete post-vllm-project#16393 upstream patch can still reverse-apply to the staged result (`git apply --reverse --check --cached`), confirming subsequent changes are preserved. - `git diff --cached --check`, Python AST parsing, Ruff lint and formatting checks passed for all 94 remaining changed Python files. - Ran the equivalent of `format.sh ci`: `pre-commit run --all-files --hook-stage manual`. Ruff, codespell, typos, clang-format, markdownlint, actionlint, package-init, forbidden-import and boolean-context checks passed. Bash-based hooks could not execute on this Windows host; two Python launcher hooks exited 9009. Full pre-commit validation remains for CI. - Actual CPU/NPU tests have not been run locally because the required vLLM/NPU runtime is unavailable. New PR CI is pending; the original upgrade PR's green run is not validation of this revert. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
…vllm-project#17004) Reintroduce the vLLM v0.29.0 release upgrade from vllm-project#16393, reverted by on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves later merged changes. The image/source mismatch is confirmed in [nightly job 106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020): the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded Ascend failed importing `_get_packed_kv_cache_groups`. All eight root Dockerfiles still defaulted to v0.28.0. The reusable image workflow passes no `VLLM_TAG` override, so those defaults govern release builds. Changing only the release marker does not update those images. - Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`. - Release changes from v0.28.0 (`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the remote tag). - Restore vllm-project#16393's source compatibility and existing UT changes by reversing vllm-project#16949, then reconcile current upstream changes. No main interface scan is rerun for this release-only upgrade. - Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P, Ubuntu/openEuler. Preserve existing exact-commit build overrides. No workflow changes. | Current-base adaptation | Exact cause and evidence | Lane | Validation | | --- | --- | --- | --- | | Eight Dockerfile `VLLM_TAG` defaults and release marker | vllm-project#16393 changed the release contract but omitted image defaults; the linked nightly log proves installation of 0.28.0 and missing `_get_packed_kv_cache_groups`. | Release images | All eight defaults match marker; image-build CI requested, pending | | `worker/v2/attn_utils.py::get_kv_cache_spec` and existing `_make_mla_layer` UT fixture | Ascend [vllm-project#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files), `df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field selector. vLLM [#51718](https://github.com/vllm-project/vllm/pull/51718/files), `8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio` with `tokens_per_state`; both exact supported pins use the latter. Use the common field, retaining metadata and cache-view assertions. | Both | Source inspection and static checks passed; actual CI pending | | `attention/attention_v1.py` import conflict | Preserve `attention_transfer_window` from Ascend [vllm-project#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files), `5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from the old compute-start import or unused weak-reference import. | Both | Conflict resolved; static checks passed | | `worker/v2/aclgraph_utils.py` import conflict | Preserve `ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend [vllm-project#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files), `34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete 0.28 version selector restored by the revert. | Both | Conflict resolved; static checks passed | | Existing 310P and Mamba model-runner UT imports | Preserve hardware-profile imports and mocks from Ascend [vllm-project#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files), `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused old-release selector import. Hardware capability routing remains unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format passed; fresh CI pending | Newly merged changes were reviewed for version-contract impact: GLM PD, vllm-project#16952 C8 config, vllm-project#16673 operator removal and vllm-project#16320 Kimi-K3 KV pool. Preserve these changes; no additional source-proven version branch was identified beyond the entries above. In particular, vllm-project#16747's `UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size` exist in both exact pins; vllm-project#16320 adds Ascend connector hooks. CI remains necessary to validate runtime interactions. CI/documentation-only PRs are retained unchanged. <details> <summary>Inherited per-file compatibility evidence from vllm-project#16393</summary> The following source-contract ledger is inherited from vllm-project#16393. Any historical verification wording refers only to that earlier PR; it does not certify this new head. New-head validation is listed below. | Reference | Commit | |---|---| | vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` | | PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` | | Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` | | Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a` | | Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` | - Use common implementations where v0.29.0 and the fixed main share KV-cache layouts, Mamba copy/group APIs, PCP handling and speculative-decoding contracts. - Retain explicit `vllm_version_is("0.29.0")` branches for contracts that still differ, including RoPE, scheduler block snapshots, InputBatch, ReplaySSM, KV zeroing and DSpark PP handling. - Remove obsolete v0.28.0 compatibility and adapt existing test fixtures. Version detection uses package versions and the explicit `VLLM_VERSION` override, with local-version suffix handling and UT environment isolation retained; no hard-coded release-SHA inference remains. - Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving other validations and EPLB platform binding. Rebuild dependent Pydantic schemas so nested configuration validation uses the patched validator. The global patch documentation records its rationale and removal criteria. - Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase; use the common exact-release selector and the real function-local DSpark sharing import. This release-only upgrade does not require a new main old-to-new interface scan. The latest rebase incorporates [vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files), which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The now-unnecessary V4.1 drafter import gate and tuple annotation have been removed. Other release adaptations, including vllm-project#15747 Spec+PP handling, remain. Detailed contract evidence is retained below for review. <details> <summary>Per-file adaptations and exact upstream evidence</summary> Evidence IDs refer to the exact source contract and upstream diff table below. Every row is syntax checked; branch-normalized AST comparison confirms unchanged main function bodies except the version identity helper and the explicitly retired propose argument/type annotations. Both supported versions completed CPU and NPU execution as recorded below. | Ascend file / symbols | Disposition and upstream evidence | Verification | |---|---|---| | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` / `_load_dspark_model_with_target_quant`; `tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased [Ascend (`82df9d871`), including manual PP partition masking. vLLM [#52809 diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share` inside the loader on both supported pins, so retain the earlier release fix: patch `eagle_utils`, never read/patch a nonexistent `dspark_utils._should_share`. v0.29 keeps its global PP guard; main has native PP via [#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) (`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks the real function-local import, absent module alias, and restoration on success/failure for both lanes; partition and PP assertions retained. Ruff/syntax pass; actual CPU/NPU pending. | | `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`; `tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29 routing and an obsolete 0.28 dev-build recognition path. For the two supported pins, #50514 exists only on fixed main. Route through `vllm_version_is("0.29.0")`; retain package/local-suffix and explicit environment-override semantics. No release-SHA inference or third release lane. | Existing UTs exercise the real uncached version helper with monkeypatch isolation, local suffix, fixed-main dev string, explicit override, and non-target versions. Isolated routing checks and Ruff/syntax pass; full CPU UT pending. | | `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`, `_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass` | #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`, `_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436; use shared standardized layouts and retain the v0.29.0 InputBatch gate. The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while removing only legacy v0.28.0 allocation paths. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` | upstream graph code and removes only the v0.28.0 import branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module imports/dispatch` | #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/recompute_scheduler.py`<br>`module imports/dispatch`, `schedule` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`, `AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker` | #52615; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module imports/dispatch` | #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__` | #53614; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module imports/dispatch`, `_get_max_layers_per_page_size`, `_ascend_max_memory_usage_bytes_from_groups`, `_ascend_get_kv_cache_config_from_groups` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` | v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it after KV binding. Evidence: [#52506, `adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c). Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. | Failure reproduced on v0.29.0; Historical release/main NPU validation passed; current-head CI pending | | `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module imports/dispatch` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant` | #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports `_should_share` locally from Eagle utilities; fixed main removes the PP guard. Keep the release `get_pp_group` patch, share through the common Eagle utility, and delete the obsolete v0.28.0 `dspark_utils._should_share` patch. | Exact release failure reproduced; Historical release/main CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`, `__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`, `register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`, `_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`, `_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers | import the NaN helpers directly because v0.29.0 and fixed main expose the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static checked; historical dual-version CPU/NPU validation passed; current-head CI pending | | `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes `pcp_manager` common to both lanes. The rebase preserves vllm-ascend v0.28.0 capture branch. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`, `_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/block_table.py`<br>`__init__`, `init_block_table_layout_tensors`, `compute_slot_mappings` | #51718, gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`, `prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`, `execute_model` | #50514, #54436, #52506, #55212, #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch` | #54282; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` | gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module imports/dispatch`, `propose` | #53694, #52188; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose` | #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | `vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`, `wake_up` | #51718, #53508; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending | | Upstream change | Full commit SHA / direct diff | Actual supported contracts and branch decision | |---|---|---| | #51718 | `8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Both use layers/layer_stride/block_stride/offset, tokens_per_state, CircularBufferSpec and standardized backing; retire shared_by/compress_ratio allocation branches. | | #52839 | `58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8) | Both import PCP operations from vllm.v1.attention.ops.pcp. | | #53896 | `e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec) | Both use Mamba copy-function dictionaries and unwrap UniformTypeKVCacheSpecs. | | #53106 | `1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21) | Both use WeightsMapper instead of AutoWeightsLoader skip_prefixes/skip_substrs. | | #53906 | `98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb) | Only pinned main has the optional MLA storage_block_size dataclass field; release keeps the Ascend derived property, using tokens_per_state. | | #56078 | `719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417) | Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned main uses mrope_num_dims and unified RoPE. | | #52615 | `138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045) | Release uses num_blocks/kv_bytes_per_block; main uses num_chunks/kv_bytes_per_chunk. | | #51358 | `6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release now has boundary_state_offloads and KVConnectorBlockState; remove partial_tail_offloads plumbing. | | #54853 | `0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c) | Release constructor takes block_ids snapshots; main takes req_ids and resolve_block_ids. Keep exact release snapshot membership and main lazy-resolution membership. | | #53614 | `144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725) | Only main configures drop_eagle_checkpoint_block for replay-aligned Mamba checkpoints. | | #50514 | `d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Release retains module-level PP/share symbols and Ascend PP workaround; main has the subsequent PP integration. | | #52809 | `91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393) | Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark module binding to a function-local import from Eagle utilities. The shared Eagle patch remains effective; the old DSpark-module read/write must be removed. | | #54436 | `6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902) | Release InputBatch requires max_seq_len_np; main removed it. | | #52506 | `adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Only main accepts valid_dummy_state_slots/valid_state_slots capture arguments. | | #55212 | `83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Release prepares DCP local sequence lengths before partitioning; main initializes DCP metadata afterwards. | | #53515 | `b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127) | Both accept padded_num_tokens for persistent PCP input buffers. | | #53869 | `b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212) | Both accept pcp_manager during graph capture. | | #51031 | `0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1) | Both distinguish KV and kernel block sizes during DCP slot mapping. | | #54282 | `fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8) | Both gumbel sampling APIs include is_drafting. | | #52188 | `d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0) | Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. | | #53694 | `5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6) | Both propose APIs take DPSyncState; remove the obsolete token-count argument and retain replicated-PCP synchronization. | | #49811 | `01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e) | Both support extract_hidden_states on MRV2; remove old unsupported dispatch/skip. | | #53508 | `479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29) | Both remove post_kv_cache_wake_up; retire release-only call. | | #52494 | `3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1) | Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain release exclusion. | | #52861 | `b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28) | Both include DeepseekV32MTPModel in the two-hidden-state architecture set. | | #54713 | `b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3) | Only main takes replay_boundaries in compressed-prefix hit lookup; preserve release calls without that keyword. | | #42785 | `442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18) | Only main capture callers pass axis_keys; preserve the existing Ascend rejection of nonempty axes. | | #52358 | `8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0) | Both ExecuteModelState have dp_sync; only main has cudagraph_stats. | | #52789 | `9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4) | Release already has mamba_has_prefill_checkpoint_blocks; later main also has fine-grained prefix-cache state. | | #51251 | `7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | v0.29.0 and main expose ec_manager_config; retire the old release-only ScoreEncoder configuration skip. | | #53240 | `b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c) | v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache groups; retire the old release-only replay skip. | | #53853 | `e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both delegate PCP compatibility validation to the PCP manager. | | #53183 | `4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650) | Both expose the V1 unsupported-feature helper used by the existing Ascend MRV1 feature filter. | | vllm-ascend change | Why it is required | Upstream cause and direct link | Lane | Verification | |---|---|---|---|---| | `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec, list[int]]` from `get_mamba_groups` and both initialize `recoverssm`; keeping the old fallback would preserve an unsupported third contract | [v0.29.0 mamba groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704), [fixed-main mamba groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704), [v0.29.0 RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103), [fixed-main RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103) | both, common implementation | Existing constructor UT now asserts the parent-created RecoverSSM value is retained; Ruff and compileall pass | | `worker/v2/model_runner.py`: always forward `kv_cache_allocation_context` | v0.29.0 and fixed main both accept this keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0 signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538), [fixed-main signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566), [vllm-ascend both, common implementation | Existing UT continues to assert the exact context object reaches the parent; Ruff and compileall pass | | `_310p/worker/v2/model_runner.py`: select the release KV-zeroing contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes `KVCacheGroupSpec.is_eagle_group` but lacks `SpeculativeConfig.use_eagle_block_drop`; fixed main added the method | vLLM [#53388 diff](https://github.com/vllm-project/vllm/pull/53388/files), commit [`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a); [v0.29.0 group field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200), [fixed-main helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876) | release differs from main | Existing two-path KV-zeroing UT retained and renamed for v0.29.0; Ruff and compileall pass | | existing DFlash kernel UT: remove v0.28-only kwarg omission | the current Ascend kernel accepts the CP arguments and the only excluded lane was v0.28.0, which this PR replaces | [vllm-ascend both, common invocation | Existing NPU test remains enabled with all assertions; Current-head CI pending | | Ascend change | Why / upstream cause | Lane | Verification | |---|---|---|---| | `patch/platform/patch_parallel_config.py`, registration, and global patch documentation | Allow Ascend PCP+DP by removing the generic GPU restriction, following vLLM [#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic `parallel_config.current_platform` lookup and rebuild ParallelConfig → SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing configuration/EPLB UTs and historical PCP+DP NPU execution passed; current-head CI pending. | | `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT | Both supported contracts require `is_drafting`, from [#54282](https://github.com/vllm-project/vllm/pull/54282/files), `fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited optimized kernel and use a common wrapper. | Both | Existing positive drafting assertion retained; current-head CI pending. | | `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache UT | Both support `cache_hit_alignment_tokens`, introduced by [#53598](https://github.com/vllm-project/vllm/pull/53598/files), `2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only write-mask branch. | Both | Existing assertions retained; current-head CI pending. | | Change | Exact source evidence | Decision | |---|---|---| | `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM [#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95), `d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL before v0.29.0; both exact supported sources lack it. The old conditional came from vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0 selector and use the existing exclusion for both supported lanes. This does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy skip. | | `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and `tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM [#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29), `12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export `nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The v0.28.0 fallback originated in vllm-ascend [vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08), `e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete import-failure/`None` fallbacks and the existing UT's obsolete availability skip; assertions remain unchanged and execute on both lanes. | | `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend [vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3), `799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream revert while resolving the real rebase conflict; do not reintroduce the reverted graph-update behavior. | | `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend [vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2), `d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged 310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation branch. | </details> </details> Yes. The supported vLLM release and default container builds move to v0.29.0, while the fixed main remains supported. The vllm-ascend package version does not change. Existing 0.29-only PCP+DP compatibility is restored. - Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`: [run 35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654) succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected NPU jobs confirm the requested Ascend head, integration base `2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM pins/installations: fixed main `84030bbe3d74d99bad477a3d2e37a973ccd8865c` / `0.1.dev1+g84030bbe3.empty`, release `98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest invocations. Skips/xfails are not passes. The resulting rebased integration commit is not printed and is not inferred. Actual release CPU and failed/cancelled image variants remain gaps. - Current-head CI update (2026-09-20): [main CPU UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955) passed **5157 tests**, with **67 skipped**; actual installed vLLM was `0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head; the log does not print the full resulting integration head/base. Pre-commit and mypy passed. NPU/E2E results are recorded above; actual release CPU remains unverified. - Image build is partially blocked by infrastructure: A5 amd64 [Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005) and [openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068) failed before reading the Dockerfile because BuildKit could not create a snapshot temporary directory (`no space left on device`). No release compatibility code change is justified by this failure; cancelled variants remain unverified. The successful [310P openEuler arm64 build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987) explicitly checked out release `98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`; this is build evidence, not runtime UT coverage. - Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto the baseline above. Previous-head CPU [job 106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458) stopped during Ascend integration rebase after vllm-project#16803 merged, before any UT ran. Its vLLM checkout was the fixed main and installation reported `0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved; fresh-head CI was retriggered by the push. - Remote official v0.29.0 tag resolved to the exact SHA above; both source pins inspected. - All 95 changed Python files pass Ruff lint, Ruff format and AST parsing; `git diff --check` passes. Dockerfile tag/marker consistency checked across all eight variants. - Full pre-commit invocation: Ruff, codespell, typos, clang-format, markdownlint, actionlint, package-init, forbidden-import and boolean-context checks passed. Bash-dependent hooks cannot run on this Windows host; Python launcher hooks exit 9009. Full CI lint remains pending. - Actual main/release CPU/NPU and image builds are not claimed locally: the required Linux/NPU/container environment is unavailable. New PR E2E and image-build CI are requested. Existing CPU workflow runs fixed main only, so actual release CPU remains a validation gap. - vllm-project#16393's historical green CI is not a substitute for this new head. No new test functions, workflow changes, golden/threshold changes or additional skips are introduced beyond restoring vllm-project#16393. Inherited 0.28-only tests/skips are not counted as passes. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
What this PR does / why we need it?
Upgrade the supported vLLM release dependency from v0.28.0 to the official v0.29.0, alongside the fixed vLLM main commit inherited from #16216 (already merged).
| Reference | Commit |
|---|---|
| vllm-ascend rebase base |
9dc6704559ebe2809b5cb7b7fc44184bd1338b3c|| PR head |
8c5d785931c495701bf1da8b5bc80b321ebd3cd0|| Fixed vLLM main |
84030bbe3d74d99bad477a3d2e37a973ccd8865c|| Previous release: v0.28.0 |
2cf0a6915ce544dc493a0990f2ea38d81601128a|| Target release: v0.29.0 |
98dff2a81d747d1dba01a47f939f48c3526d4206|Use common implementations where v0.29.0 and the fixed main share KV-cache layouts, Mamba copy/group APIs, PCP handling and speculative-decoding contracts.
Retain explicit
vllm_version_is("0.29.0")branches for contracts that still differ, including RoPE, scheduler block snapshots, InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.Remove obsolete v0.28.0 compatibility and adapt existing test fixtures. Version detection uses package versions and the explicit
VLLM_VERSIONoverride, with local-version suffix handling and UT environment isolation retained; no hard-coded release-SHA inference remains.Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving other validations and EPLB platform binding. Rebuild dependent Pydantic schemas so nested configuration validation uses the patched validator. The global patch documentation records its rationale and removal criteria.
Preserve [Feature][MRV2] Support GLM5.2 DSpark with PP #15747's Spec+PP protocol/partition handling after rebase; use the common exact-release selector and the real function-local DSpark sharing import.
This release-only upgrade does not require a new main old-to-new interface scan. The latest rebase incorporates #16905, which reverted #16544 at
9dc6704559ebe2809b5cb7b7fc44184bd1338b3c. The now-unnecessary V4.1 drafter import gate and tuple annotation have been removed. Other release adaptations, including #15747 Spec+PP handling, remain. Detailed contract evidence is retained below for review.Per-file adaptations and exact upstream evidence
Per-file adaptation ledger
Evidence IDs refer to the exact source contract and upstream diff table below. Every row is syntax checked; branch-normalized AST comparison confirms unchanged main function bodies except the version identity helper and the explicitly retired propose argument/type annotations. Both supported versions completed CPU and NPU execution as recorded below.
| Ascend file / symbols | Disposition and upstream evidence | Verification |
|---|---|---|
|
vllm_ascend/patch/worker/patch_v2/patch_dspark.py/_load_dspark_model_with_target_quant;tests/ut/patch/worker/test_patch_dspark_pp.py| Preserve rebased Ascend #15747 (82df9d871), including manual PP partition masking. vLLM #52809 diff (91a893de64722019ea2faf852e06cabe143b3490) moved_should_shareinside the loader on both supported pins, so retain the earlier release fix: patcheagle_utils, never read/patch a nonexistentdspark_utils._should_share. v0.29 keeps its global PP guard; main has native PP via #50514 (d87a440f88e28e5b37f9b1e22ce214d0426d5352). | Existing UT now checks the real function-local import, absent module alias, and restoration on success/failure for both lanes; partition and PP assertions retained. Ruff/syntax pass; actual CPU/NPU pending. ||
vllm_ascend/worker/v2/pp_utils.py/use_legacy_spec_pp;tests/ut/worker/v2/test_pp_utils.py| #15747 added broad 0.28/0.29 routing and an obsolete 0.28 dev-build recognition path. For the two supported pins, #50514 exists only on fixed main. Route throughvllm_version_is("0.29.0"); retain package/local-suffix and explicit environment-override semantics. No release-SHA inference or third release lane. | Existing UTs exercise the real uncached version helper with monkeypatch isolation, local suffix, fixed-main dev string, explicit override, and non-target versions. Isolated routing checks and Ruff/syntax pass; full CPU UT pending. ||
vllm_ascend/_310p/model_runner_310p.py_prepare_inputs,_allocate_kv_cache_tensors| #51718, #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/_310p/spec_decode/llm_base_proposer_310.pyset_inputs_first_pass| #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/_310p/worker/v2/model_runner.pyinitialize_kv_cache,_allocate_kv_cache_tensors,_prepare_inputs_310p| #51718/#54436; use shared standardized layouts and retain the v0.29.0 InputBatch gate. The rebase preserves vllm-ascend #16043's MTP copy tracking while removing only legacy v0.28.0 allocation paths. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/_310p/worker/v2/rope.pyget_310p_rope_state| #56078; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/attention/attention_v1.pymodule imports/dispatch| #52839; both lanes use the common PCP import. Rebase keeps current upstream graph code and removes only the v0.28.0 import branch. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/attention/context_parallel/sfa_cp.pymodule imports/dispatch| #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/attention/indexer.pymodule imports/dispatch| #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/attention/mla_v1.pymodule imports/dispatch| #52839; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/core/dyntra_lb_scheduler.pymodule imports/dispatch| #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/core/recompute_scheduler.pymodule imports/dispatch,schedule| #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/core/scheduler_profiling_chunk.pymodule imports/dispatch| #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/core/kv_cache_interface.pyget_kv_cache_compression_ratio,AscendMLAAttentionSpec,get_storage_block_size| #51718, #53906; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py_merge_specs| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py_derive_cpu_config| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.pycreate_worker| #52615; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/models/deepseek_mtp.pyload_weights| #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/models/deepseek_v4/indexer.pyget_kv_cache_spec| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/models/glm5next/cache_config.pymake_tensor| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/models/glm5next/kv_cache.pyget_kv_cache_spec| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/models/kimi_k3_dspark.pyload_weights| #53106; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/models/layer/attention/layer.pyget_kv_cache_spec| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/platform/patch_balance_schedule.pymodule imports/dispatch| #51358, #54853; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/platform/patch_kv_cache_coordinator.py__init__| #53614; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/platform/patch_kv_cache_utils.pymodule imports/dispatch,_get_max_layers_per_page_size,_ascend_max_memory_usage_bytes_from_groups,_ascend_get_kv_cache_config_from_groups| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/platform/patch_use_v2_model_runner.pymodule imports/dispatch,_patched_get_unsupported_features| #53853, #53183; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/worker/patch_bind_kv_cache.pybind_kv_cache| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache| v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it after KV binding. Evidence: #52506,adebc41b7e9f1085d3f73434e23beb76883b9eb4, worker-utils diff. Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. | Failure reproduced on v0.29.0; Historical release/main NPU validation passed; current-head CI pending ||
vllm_ascend/patch/worker/patch_mamba_utils.pymodule imports/dispatch,_get_state_copy_funcs_for_layer| #53896; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/worker/patch_v2/patch_attn_utils.pymodule imports/dispatch| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/patch/worker/patch_v2/patch_dspark.py_load_dspark_model_with_target_quant| #50514, #52809; v0.29.0 keepsget_pp_groupmodule-bound but imports_should_sharelocally from Eagle utilities; fixed main removes the PP guard. Keep the releaseget_pp_grouppatch, share through the common Eagle utility, and delete the obsolete v0.28.0dspark_utils._should_sharepatch. | Exact release failure reproduced; Historical release/main CPU/NPU validation passed; current-head CI pending ||
vllm_ascend/spec_decode/llm_base_proposer.pymodel_returns_tuple,__init__,set_inputs_first_pass| #56078, #52861; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/utils.pyget_kv_cache_tensor_layers,register_ascend_customop,vllm_version_is| #52839, #51718, #52494; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/model_runner_v1.py_get_ascend_mamba_state_copy_funcs,_allocate_kv_cache_tensors,_prepare_inputs,_dummy_run,_reshape_kv_cache_tensors,get_kv_cache_spec, NaN helpers | #51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and import the NaN helpers directly because v0.29.0 and fixed main expose the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static checked; historical dual-version CPU/NPU validation passed; current-head CI pending ||
vllm_ascend/worker/v2/aclgraph_utils.pycapture| #53869 makespcp_managercommon to both lanes. The rebase preserves vllm-ascend #16409's host-parameter-update revert and removes only the obsolete v0.28.0 capture branch. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/attn_utils.pyget_kv_cache_spec,_allocate_kv_cache,_reshape_kv_cache_v2| #51718; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/block_table.py__init__,init_block_table_layout_tensors,compute_slot_mappings| #51718, #51031; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/model_runner.pymodule imports/dispatch,prepare_inputs,__init__,sample_tokens,prepare_dummy_attn,execute_model| #50514, #54436, #52506, #55212, #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/pcp_manager.pypartition_batch| #53515; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/sample/gumbel.pymodule imports/dispatch| #54282; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/spec_decode/__init__.pyinit_speculator| #49811; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.pypropose| #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/spec_decode/dflash/speculator.pymodule imports/dispatch,propose| #53694, #52188; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/v2/spec_decode/dspark/speculator.pypropose| #53694; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending ||
vllm_ascend/worker/worker.py_scale_kv_cache_memory_for_multi_group,wake_up| #51718, #53508; common contracts merged, differing release contracts explicitly gated as detailed below. | Static checked; prior-head dual-version results below; rebased CI pending |Exact upstream evidence
| Upstream change | Full commit SHA / direct diff | Actual supported contracts and branch decision |
|---|---|---|
| #51718 |
8bdc70ec7b379279ec0152343239c2d50aced687vllm/v1/kv_cache_interface.py | Both use layers/layer_stride/block_stride/offset, tokens_per_state, CircularBufferSpec and standardized backing; retire shared_by/compress_ratio allocation branches. |
| #52839 |
58e5ee0158b6a264c3506f00480e108a34b33ee3vllm/v1/attention/ops/pcp.py | Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
e126687a9a828d513c01a07cd69f025f27d63280vllm/v1/worker/mamba_utils.py | Both use Mamba copy-function dictionaries and unwrap UniformTypeKVCacheSpecs. |
| #53106 |
1fe3a1571ac67581478a11743e55a306de1d136fvllm/model_executor/models/utils.py | Both use WeightsMapper instead of AutoWeightsLoader skip_prefixes/skip_substrs. |
| #53906 |
98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5vllm/v1/kv_cache_interface.py | Only pinned main has the optional MLA storage_block_size dataclass field; release keeps the Ascend derived property, using tokens_per_state. |
| #56078 |
719284fe158f1be8a9dd92953295fc9d49015730vllm/config/model.py | Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned main uses mrope_num_dims and unified RoPE. |
| #52615 |
138d137b5b955a5ebee98dfc946ecdd65d7b87cevllm/v1/kv_offload/cpu/spec.py | Release uses num_blocks/kv_bytes_per_block; main uses num_chunks/kv_bytes_per_chunk. |
| #51358 |
6b110badbb22d3f66c7218b71138f13b7a6b3419vllm/v1/core/sched/output.py | Release now has boundary_state_offloads and KVConnectorBlockState; remove partial_tail_offloads plumbing. |
| #54853 |
0b066293f3c738a0cbd3a087bf893f2f4dcd61f2vllm/v1/core/sched/output.py | Release constructor takes block_ids snapshots; main takes req_ids and resolve_block_ids. Keep exact release snapshot membership and main lazy-resolution membership. |
| #53614 |
144e79c8106da23141ac010394b782f730cc7fe8vllm/v1/core/kv_cache_coordinator.py | Only main configures drop_eagle_checkpoint_block for replay-aligned Mamba checkpoints. |
| #50514 |
d87a440f88e28e5b37f9b1e22ce214d0426d5352vllm/v1/worker/gpu/spec_decode/dspark/utils.py | Release retains module-level PP/share symbols and Ascend PP workaround; main has the subsequent PP integration. |
| #52809 |
91a893de64722019ea2faf852e06cabe143b3490vllm/v1/worker/gpu/spec_decode/dspark/utils.py | Between v0.28.0 and v0.29.0,
_should_sharemoved from a DSpark module binding to a function-local import from Eagle utilities. The shared Eagle patch remains effective; the old DSpark-module read/write must be removed. || #54436 |
6bafc049aae6c26e210162630914ee9177a4b586vllm/v1/worker/gpu/input_batch.py | Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
adebc41b7e9f1085d3f73434e23beb76883b9eb4vllm/v1/worker/gpu/model_runner.py | Only main accepts valid_dummy_state_slots/valid_state_slots capture arguments. |
| #55212 |
83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3vllm/v1/worker/gpu/model_runner.py | Release prepares DCP local sequence lengths before partitioning; main initializes DCP metadata afterwards. |
| #53515 |
b1fbbc2ade51e3826bc92e4733c9c692ee21d42dvllm/v1/worker/gpu/pcp_manager.py | Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
b3af042abd8fe5a297ec3ec72db276fd661a67b3vllm/v1/worker/gpu/cudagraph_utils.py | Both accept pcp_manager during graph capture. |
| #51031 |
0ecc284790e5403f74b899524ef82ecb69f83cb3vllm/v1/worker/gpu/block_table.py | Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
fe755c88995ad468882517b6c4bdd60138d46a3avllm/v1/worker/gpu/sample/gumbel.py | Both gumbel sampling APIs include is_drafting. |
| #52188 |
d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9vllm/v1/worker/gpu/spec_decode/dflash/speculator.py | Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py | Both propose APIs take DPSyncState; remove the obsolete token-count argument and retain replicated-PCP synchronization. |
| #49811 |
01af92e175407231b1433b0aef01a1b9c983d955vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py | Both support extract_hidden_states on MRV2; remove old unsupported dispatch/skip. |
| #53508 |
479eeb32d2b432dbb4e442fbd3f94ca2eca35d67vllm/v1/worker/gpu_model_runner.py | Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fccvllm/models/kimi_k3/amd/mla.py | Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain release exclusion. |
| #52861 |
b09bd69b5bf14911abef9a0e8e493b83c8a38fa6vllm/v1/spec_decode/llm_base_proposer.py | Both include DeepseekV32MTPModel in the two-hidden-state architecture set. |
| #54713 |
b28c3e1568bfae930f61d4b24940e47528c85d4avllm/v1/core/single_type_kv_cache_manager.py | Only main takes replay_boundaries in compressed-prefix hit lookup; preserve release calls without that keyword. |
| #42785 |
442d36031ce710aa2a777c353becb660c57ab2bdvllm/v1/worker/encoder_cudagraph.py | Only main capture callers pass axis_keys; preserve the existing Ascend rejection of nonempty axes. |
| #52358 |
8f816a3f665489d7f0d222115d4f72ebab01076bvllm/v1/worker/gpu/model_runner.py | Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
9eb9d9d3953959695108600c8ed33d36bc6a1e5fvllm/v1/core/sched/scheduler.py | Release already has mamba_has_prefill_checkpoint_blocks; later main also has fine-grained prefix-cache state. |
| #51251 |
7bbbf7c8e5040f7ebd374e8cbc657e01af1136ddvllm/config/vllm.py | v0.29.0 and main expose ec_manager_config; retire the old release-only ScoreEncoder configuration skip. |
| #53240 |
b2db227a7c4c5e55f85524a094b607e5d27408b4vllm/model_executor/layers/fused_moe/routed_experts_capturer.py | v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache groups; retire the old release-only replay skip. |
| #53853 |
e376d45e82cb7e220da430e3179e81eb0922cf56vllm/config/vllm.py | Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
4aab2b0ebed20343efe543c633f71b3c1336d5b8vllm/config/vllm.py | Both expose the V1 unsupported-feature helper used by the existing Ascend MRV1 feature filter. |
Additional inherited contracts
| vllm-ascend change | Why it is required | Upstream cause and direct link | Lane | Verification |
|---|---|---|---|---|
|
_310p/worker/v2/model_state.py: remove tuple-return and RecoverSSM v0.28 fallbacks | v0.29.0 and fixed main both returndict[MambaSpec, list[int]]fromget_mamba_groupsand both initializerecoverssm; keeping the old fallback would preserve an unsupported third contract | v0.29.0 mamba groups, fixed-main mamba groups, v0.29.0 RecoverSSM, fixed-main RecoverSSM | both, common implementation | Existing constructor UT now asserts the parent-created RecoverSSM value is retained; Ruff and compileall pass ||
worker/v2/model_runner.py: always forwardkv_cache_allocation_context| v0.29.0 and fixed main both accept this keyword; the new-base gate was specifically for v0.28.0 | v0.29.0 signature, fixed-main signature, vllm-ascend #16791 | both, common implementation | Existing UT continues to assert the exact context object reaches the parent; Ruff and compileall pass ||
_310p/worker/v2/model_runner.py: select the release KV-zeroing contract withvllm_version_is("0.29.0")| v0.29.0 exposesKVCacheGroupSpec.is_eagle_groupbut lacksSpeculativeConfig.use_eagle_block_drop; fixed main added the method | vLLM #53388 diff, commit481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81; v0.29.0 group field, fixed-main helper | release differs from main | Existing two-path KV-zeroing UT retained and renamed for v0.29.0; Ruff and compileall pass || existing DFlash kernel UT: remove v0.28-only kwarg omission | the current Ascend kernel accepts the CP arguments and the only excluded lane was v0.28.0, which this PR replaces | vllm-ascend #15098 | both, common invocation | Existing NPU test remains enabled with all assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |
|---|---|---|---|
|
patch/platform/patch_parallel_config.py, registration, and global patch documentation | Allow Ascend PCP+DP by removing the generic GPU restriction, following vLLM #54523, commit7c2f1ff4958eaf0818405e9192c71608fe4a16b1. Preserve dynamicparallel_config.current_platformlookup and rebuild ParallelConfig → SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing configuration/EPLB UTs and historical PCP+DP NPU execution passed; current-head CI pending. ||
ops/triton/v2/sample/categorical_sample.py, existing Gumbel UT | Both supported contracts requireis_drafting, from #54282,fe755c88995ad468882517b6c4bdd60138d46a3a; preserve the inherited optimized kernel and use a common wrapper. | Both | Existing positive drafting assertion retained; current-head CI pending. ||
patch/platform/patch_kv_cache_coordinator.py, existing prefix-cache UT | Both supportcache_hit_alignment_tokens, introduced by #53598,2ba984a5d06db414f3b2474fe9338faf6cd80a1c; retire the v0.28-only write-mask branch. | Both | Existing assertions retained; current-head CI pending. |Rebase and retired-fallback evidence
| Change | Exact source evidence | Decision |
|---|---|---|
|
tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]| vLLM #53272,d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5removes native Hunyuan V1/VL before v0.29.0; both exact supported sources lack it. The old conditional came from vllm-ascend #14898,e5118d151314ae18e56c0b63aa8dd00d294adc22. | Remove the v0.28.0 selector and use the existing exclusion for both supported lanes. This does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy skip. ||
vllm_ascend/worker/model_runner_v1.pyNaN helper imports andtests/ut/worker/test_model_runner_v1_nan_detection.py| vLLM #50323,12292d94b25869be2af6b6d4f8eea6c2445e935f; both exact sources exportnans_to_dict,gpu_sync_allowed, andraise_if_nan_logits. The v0.28.0 fallback originated in vllm-ascend #14898,e5118d151314ae18e56c0b63aa8dd00d294adc22. | Delete import-failure/Nonefallbacks and the existing UT's obsolete availability skip; assertions remain unchanged and execute on both lanes. ||
vllm_ascend/worker/v2/aclgraph_utils.py| vllm-ascend #16409,799801feef347469d5e9b39374b9210e0d4f7431. | Preserve the upstream revert while resolving the real rebase conflict; do not reintroduce the reverted graph-update behavior. ||
vllm_ascend/_310p/worker/v2/model_runner.py| vllm-ascend #16043,d4d2957e5208c2f464d4625c05920bd29ea233cb. | Preserve the newly merged 310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation branch. |Does this PR introduce any user-facing change?
Yes. The supported release changes from vLLM v0.28.0 to v0.29.0, and Ascend PCP+DP is enabled on v0.29.0. The fixed main commit remains supported. The vllm-ascend package version is unchanged.
Obsolete v0.28-only skips are removed where the contracts are now supported. No skip is migrated to v0.29.0, and no golden data or precision threshold is changed. Existing unrelated exclusions remain exclusions, not passes.
How was this patch tested?
Current-head validation: run 35420046256 completed successfully on
fd876c5269fb6b39afa04e5a52433505e8786e22: 39 successful jobs and 6 skipped jobs, no failures.84030bbe3d74d99bad477a3d2e37a973ccd8865c, installed0.1.dev1+g84030bbe3.empty; Ascend checkout is the PR head. CPUbase_shais empty; the log says HEAD is up to date, so no unprinted integration SHA is inferred.84030bbe3d74d99bad477a3d2e37a973ccd8865c, release98dff2a81d747d1dba01a47f939f48c3526d4206; installed versions0.1.dev1+g84030bbe3.emptyand0.29.0+empty. Each log confirms Ascend headfd876c5269fb6b39afa04e5a52433505e8786e22, integration basee139b7d573d3769fd1407d5027d7d4831f5469f0, and HEAD is up to date.Latest rebase: head
fd876c5269fb6b39afa04e5a52433505e8786e22onto upstream maine139b7d573d3769fd1407d5027d7d4831f5469f0. The old-head green results below are historical; new-head CI has completed successfully; see the verified results below. The fixed vLLM main and release pins remain unchanged. PR-wide Ruff, formatting and Python AST checks passed for 95 Python files; no workflow changes.Rebase reconciliation:
Preserve upstream #16853,
8f2e3fed73328148ea603ccfe5764542fa7cbbd2, without migrating its 0.28-only validator/dispatch gates to 0.29. Both supported pins already call the PCP manager for dispatch token counts. Restore the version-helper import needed by the inherited method. Its inherited 0.28-only skipped tests are not counted as passes.Adapt the existing unsupported-feature UT to preserve the upstream list: both supported pins delegate PCP checks to the manager via vLLM #53853,
e376d45e82cb7e220da430e3179e81eb0922cf56; no obsolete 0.28 string filtering is reinstated.Adapt the existing KVPP allocation-entry UT from #15514,
b64959a11ea69a51409f64e4e68da8e0924dec42, to the common allocation entry already used by both supported pins after vLLM #51718. Preserve cache-view assertions and the upstream dtype changes. No test functions or skips added.[Feature] Add DeepSeek V4.1 framework support and Engram host offload #16544 remains reverted by Revert "[Feature] Add DeepSeek V4.1 framework support and Engram host offload (#16544)" #16905; its V4.1 compatibility workaround remains removed.
Historical validation before this rebase:
Current head
8c5d785931c495701bf1da8b5bc80b321ebd3cd0: run 35350119603 completed successfully: 40 successful jobs, 6 skipped jobs, no failures. Pre-commit and mypy passed. Local Ruff lint/format and AST checks passed for all 95 changed Python files.Actual main CPU: job 105617611031 reports 5074 passed, 37 skipped, 17 warnings. Logs confirm the exact PR head, vLLM checkout
84030bbe3d74d99bad477a3d2e37a973ccd8865c, and installed0.1.dev1+g84030bbe3.empty. CPUbase_shais empty; no integration merge is inferred.Actual dual-version NPU: All 32 NPU job logs were audited, 16 per lane. Fixed main installed
0.1.dev1+g84030bbe3.empty; official release checkout98dff2a81d747d1dba01a47f939f48c3526d4206installed0.29.0+empty. Each lane reports 562 passed, 35 skipped, 1 xfailed across pytest session summaries (execution counts, not deduplicated unique tests). Logs confirm checkout of this PR head and integration base9dc6704559ebe2809b5cb7b7fc44184bd1338b3c; the rebase step reports HEAD is up to date.PCP combinations: The context-parallel suite containing PCP+DP and PCP+PP+MTP passes on both main A3 part4 and release A3 part4.
Release CPU gap: Current-head actual release CPU remains missing because the workflow has only a main CPU entry. Version-parameterized UTs and historical release CPU runs do not replace actual release installation. Historical run 35213834751 at head
2560da9c7c2628f1b92b9305a1e80c5856bce74fpassed 4923 tests with 37 skips per lane; its temporary CPU workflow was reverted.Backup: [Misc] Backup green v0.29.0 upgrade snapshot from #16393 #16898 remains unchanged with no test CI triggered. Local branch
codex/backup-16393-before-16905preserves previous green7ad4c82ee; its results are not used as current-head validation.Skipped/xfail cases and jobs are not counted as passes. Earlier intermittent HCCL errors have no proven root cause; no speculative fix is included. This run passes existing CI, but complete dual-version CPU validation remains outstanding.