Skip to content

[Feature] Enable GLM-5.3-Flash PD disaggregation on Model Runner V2 - #16755

Merged
weijinqian0 merged 1 commit into
vllm-project:mainfrom
sunbaosong:main
Sep 20, 2026
Merged

weijinqian0 merged 1 commit into
vllm-project:mainfrom
sunbaosong:main

Conversation

@sunbaosong

@sunbaosong sunbaosong commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Enables GLM-Next (GLM-5.3-Flash) prefill/decode disaggregation on Model Runner V2 with MooncakeConnectorV2, for both equal TP (P=TP4 / D=TP4) and unequal TP (P=TP8 / D=DP2xTP4).

GLM-Next could not run PD on MRV2 because of three structural gaps:

  1. Granularity mismatch. MRV2 hands metadata builders kernel-granularity specs (block_size == the C128 SFA kernel size), while the scheduler, the common block tables, and the GLM-Next cache contract all use the logical block from cache_config.block_size. Sizing persistent buffers from the spec overflows them by the ceil-rounding slack (required=1044 vs capacity=1040 on the first forward).
  2. GLM-Next layout. The v2 MLA spec re-wrap dropped the model_version / compress_ratio markers the cache-group classification needs; the compressed indexer small page (compress_ratio > 1) is a single-tensor overlay on a shared padded slot, not an MLA latent/rope K/V pair; NoPE MLA layers legitimately carry an empty rope/V component.
  3. Unequal TP. Page-size unification is TP-dependent for hybrid models (the Mamba state page scales with heads/TP), so a TP8 producer and a TP4 consumer publish different unified block sizes for byte-identical replicated caches — kernel blocks can no longer be paired index-wise.

Changes:

  • v2 attn_utils: propagate model_version / indexes_kv_by_block_stride / compress_ratio / non_causal_multi_token_decode through the MLA spec re-wrap; reshape the compressed indexer small-page layer (compress_ratio > 1) as a single strided tensor overlaid from byte zero of the shared padded slot.
  • KPool indexer/tail backends: accept upstream's cache_dtype_str probe kwarg; take the logical block size from cache_config.block_size.
  • SFA metadata state: size the persistent block table buffer from the cache config's logical block expansion instead of the spec's.
  • Mooncake V2 connector: skip zero-size cache components when collecting register ranges (their degenerate storage corrupts the range arithmetic) and skip zero-length transfer entries the engine rejects with E19999.
  • Unequal-TP: transfer replicated MLA/indexer caches as GCD-sized token-segment byte runs (_compute_cross_tp_runs); transfer the indexer tail whole-block instead of through the HND head-sharded path; raise with the layer name on kernel-block-size mismatch.

Does this PR introduce any user-facing change?

Yes — it enables a new deployment topology: GLM-Next hybrid models can now run PD disaggregation on MRV2 with MooncakeConnectorV2 (pull mode), including a TP8 producer feeding a DP2xTP4 consumer. Requires VLLM_USE_V2_MODEL_RUNNER=1 and a kv-transfer-config with MooncakeConnectorV2. No behavior change for existing models and configurations: the new code paths are gated (the strided reshape only triggers for MLA-family specs with compress_ratio > 1; the cross-TP run path only activates when local and remote kernel block sizes differ; the equal-TP block pairing path is unchanged).

How was this patch tested?

  • Unit tests: the mooncake_v2 suite passes 150/150, covering cross-TP run splitting/coalescing/computed-offset/bounds, tail replica/whole-block mapping, bucket routing, and per-component address scaling. The full UT sweep over tests/ut/{models,attention,worker,core,kv_offload} produces a failure set identical to upstream main (pre-existing failures only, none introduced by this PR).

  • E2E, 10-layer GLM-5.3-Flash debug model: 1P1D TP4/TP4 (producer cards 0-3, consumer cards 4-7, eager), pull-mode KV transfer over the mooncake engine relayed through a load-balance proxy: first-token fidelity 18/18, outputs match the single-instance baseline to the model's own determinism, real per-rank transfer logs, zero recompute fallback.

  • E2E, full w8a8 GLM-5.3-Flash: P=TP8 producer, D=DP2xTP4 consumer (FULL_DECODE_ONLY graph mode): zero pull errors, ~18.5 MB transferred per request, TextVQA accuracy consistent with the colocated baseline.

  • vLLM main: vllm-project/vllm@84030bb

@sunbaosong sunbaosong changed the title [Feature] Enable GLM-Next PD disaggregation on Model Runner V2 with M… [Feature] Enable GLM-5.3-Flash PD disaggregation on Model Runner V2 Sep 17, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for GLM-5.3-Flash prefill/decode disaggregation on Model Runner V2 using the MooncakeConnectorV2 engine. It enables efficient KV cache transfers across both equal and unequal tensor parallelism (TP) topologies by implementing robust cross-TP run logic and updating metadata propagation and cache spec handling.

Highlights

  • GLM-Next PD Disaggregation Support: Enabled prefill/decode disaggregation on Model Runner V2 using MooncakeConnectorV2 for both equal and unequal tensor parallelism (TP) topologies.
  • Cross-TP Transfer Logic: Implemented GCD-based token-segment transfer logic to handle replicated MLA and indexer caches across mismatched TP topologies.
  • Metadata and Spec Updates: Updated attn_utils, KPool indexer, and SFA metadata to propagate model versions and utilize correct cache configuration block sizes.
  • Connector Robustness: Added handling for zero-size cache components and zero-length transfer entries in the Mooncake V2 connector to improve stability.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Support cross-TP Mooncake transfers and indexer tail replication

Suggested PR Summary:

### What this PR does / why we need it?
This PR introduces support for cross-TP Mooncake transfers across unequal block grids (e.g., TP4 consumer vs TP8 producer) by splitting and coalescing token-unit runs on GCD boundaries. It also adds support for replicating the GLM-Next indexer tail spec (`AscendIndexerKPoolTailSpec`) across mismatched TP deployments. Additionally, it handles compressed indexer cache reshaping using `torch.as_strided` and updates the relevant unit tests.

Feedback from the review highlights two main improvements:
1. Update the return type annotation of `_build_transfer_block_buckets` in `pull_worker.py` to reflect the new 3-tuple return value.
2. Use `getattr(spec, "non_causal_multi_token_decode", False)` instead of direct attribute access in `attn_utils.py` to prevent potential `AttributeError` across different vLLM versions.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with new and updated unit tests in `tests/ut/kv_offload/mooncake_v2/test_pull_worker.py` covering cross-TP runs, tail spec replication, and transfer bucket routing.

Comment thread vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/pull_worker.py Outdated
Comment thread vllm_ascend/worker/v2/attn_utils.py
@sunbaosong
sunbaosong marked this pull request as ready for review September 17, 2026 06:10
@sunbaosong
sunbaosong force-pushed the main branch 3 times, most recently from 40ec335 to a871b30 Compare September 17, 2026 06:34
@nwpu-zxr

Copy link
Copy Markdown
Collaborator

I think using a large block size such as 160 or 288 is not well, it's better that separate the padding size to each small blocks with 32 kernel block size indexer cache.

@sunbaosong
sunbaosong force-pushed the main branch 4 times, most recently from afc0dfb to 5a6bdda Compare September 18, 2026 02:10
weijinqian0
weijinqian0 previously approved these changes Sep 18, 2026
…ooncakeConnectorV2

Support GLM-5.3-Flash prefill/decode disaggregation on MRV2 with the
MooncakeConnectorV2 pull engine for equal and unequal TP topologies
(TP4/TP4 and P=TP8 -> D=DP2xTP4), transferring every cache at
whole-block granularity with zero changes to the pull path:

- v2 attn_utils: propagate model_version / indexes_kv_by_block_stride /
  compression through the MLA spec re-wrap.
- Indexer kernel-block layout: lay the compressed indexer cache out as
  [num_blocks x scale, 32, H] natural kernel blocks instead of
  [num_blocks, 288/160, H] pages inflated by the TP-dependent unified
  page, so dim0 blocks are byte-identical across engines with different
  TP. The unified small slot keeps its size: the indexer takes a
  contiguous prefix and the per-request tail rings a contiguous suffix
  at natural 4KB granularity, each region bounded to half the slot.
- kpool builder: report the 32-row kernel blocks and pass the common
  expanded block table through as a view; slot mapping and the
  read/write kernels are unchanged (flat-row addressing is invariant).
- Classify the indexer tail as TP-replicated: its block shape leads
  with the fixed K/gate pair (2), which head-count inference read as a
  per-rank KV head count and rejected with "[8, 16]" across unequal TP.
- KPool/SFA/tail backends: accept upstream's cache_dtype_str probe
  kwarg; take the logical block size from cache_config.block_size; size
  the persistent block-table buffer from the cache config's expansion.
- Mooncake V2: skip zero-size cache components in register ranges and
  zero-length transfer entries.

Verified on GLM-5.3-Flash w8a8 (P=TP8 eager, D=DP2xTP4 decode graphs):
zero pull errors, decode output matches single-engine baselines,
colocated throughput unchanged (2002 vs 1913 tok/s on 8k-in/1k-out at
10 concurrency), and TextVQA accuracy is identical between PD and
colocated (32% both, authoritative rescore).

Co-authored-by: bubaishenhua112-netizen <bubaishenhua112@gmail.com>
Signed-off-by: sunbaosong <13793883820@163.com>
lijiahang226 added a commit to lijiahang226/vllm-ascend that referenced this pull request Sep 19, 2026
Preserve cache capabilities and reuse page-strided storage for shared MLA, KDA and KPool caches. Expose complete shared pages to copy-on-write and align speculative graph capture with upstream lifecycle hooks. Keep tail history across draft rejection, compatible Mooncake transfer handling and inherited video processing.

Refs vllm-project#15665. Refs vllm-project#16755.

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
lijiahang226 added a commit to lijiahang226/vllm-ascend that referenced this pull request Sep 19, 2026
Preserve cache capabilities and reuse page-strided storage for shared MLA, KDA and KPool caches. Expose complete shared pages to copy-on-write and align speculative graph capture with upstream lifecycle hooks. Keep tail history across draft rejection, compatible Mooncake transfer handling and inherited video processing.

Refs vllm-project#15665. Refs vllm-project#16755.

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
@weijinqian0
weijinqian0 merged commit df3e755 into vllm-project:main Sep 20, 2026
31 checks passed
ningjingbengxiaohai pushed a commit that referenced this pull request Sep 20, 2026
…#17004)

### What this PR does / why we need it?



Reintroduce the vLLM v0.29.0 release upgrade from #16393, reverted by
#16949, with container defaults aligned to the supported release. Based
on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves
later merged changes.


The image/source mismatch is confirmed in [nightly job
106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020):
the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded
Ascend failed importing `_get_packed_kv_cache_groups`. All eight root
Dockerfiles still defaulted to v0.28.0. The reusable image workflow
passes no `VLLM_TAG` override, so those defaults govern release builds.
Changing only the release marker does not update those images.


- Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

- Release changes from v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0
(`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the
remote tag).
- Restore #16393's source compatibility and existing UT changes by
reversing #16949, then reconcile current upstream changes. No main
interface scan is rerun for this release-only upgrade.
- Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P,
Ubuntu/openEuler. Preserve existing exact-commit build overrides. No
workflow changes.


| Current-base adaptation | Exact cause and evidence | Lane | Validation
|
| --- | --- | --- | --- |

| Eight Dockerfile `VLLM_TAG` defaults and release marker | #16393
changed the release contract but omitted image defaults; the linked
nightly log proves installation of 0.28.0 and missing
`_get_packed_kv_cache_groups`. | Release images | All eight defaults
match marker; image-build CI requested, pending |
| `worker/v2/attn_utils.py::get_kv_cache_spec` and existing
`_make_mla_layer` UT fixture | Ascend
[#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files),
`df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field
selector. vLLM
[#51718](https://github.com/vllm-project/vllm/pull/51718/files),
`8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio`
with `tokens_per_state`; both exact supported pins use the latter. Use
the common field, retaining metadata and cache-view assertions. | Both |
Source inspection and static checks passed; actual CI pending |
| `attention/attention_v1.py` import conflict | Preserve
`attention_transfer_window` from Ascend
[#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files),
`5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from
#16798; keep the common relocated PCP import from #16393. Do not restore
the old compute-start import or unused weak-reference import. | Both |
Conflict resolved; static checks passed |
| `worker/v2/aclgraph_utils.py` import conflict | Preserve
`ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend
[#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files),
`34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete
0.28 version selector restored by the revert. | Both | Conflict
resolved; static checks passed |


| Existing 310P and Mamba model-runner UT imports | Preserve
hardware-profile imports and mocks from Ascend
[#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files),
`d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused
old-release selector import. Hardware capability routing remains
unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format
passed; fresh CI pending |

Newly merged changes were reviewed for version-contract impact:
#16775/#16834 PCP metadata and routing, #16923 A3 SFA,
#16426/#16924/#16955 custom ops, #16081 MTP/SP, #16669 xlite,
#15636/#16747 transfer, #16913 A5 pages, #16798 graph updates, #16755
GLM PD, #16952 C8 config, #16673 operator removal and #16320 Kimi-K3 KV
pool. Preserve these changes; no additional source-proven version branch
was identified beyond the entries above. In particular, #16747's
`UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size`
exist in both exact pins; #16320 adds Ascend connector hooks. CI remains
necessary to validate runtime interactions. CI/documentation-only PRs
are retained unchanged.


<details>

<summary>Inherited per-file compatibility evidence from #16393</summary>


The following source-contract ledger is inherited from #16393. Any
historical verification wording refers only to that earlier PR; it does
not certify this new head. New-head validation is listed below.


| Reference | Commit |





|---|---|





| vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` |
| PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` |





| Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` |





| Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a`
|
| Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` |






- Use common implementations where v0.29.0 and the fixed main share
KV-cache layouts, Mamba copy/group APIs, PCP handling and
speculative-decoding contracts.
- Retain explicit `vllm_version_is("0.29.0")` branches for contracts
that still differ, including RoPE, scheduler block snapshots,
InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.
- Remove obsolete v0.28.0 compatibility and adapt existing test
fixtures. Version detection uses package versions and the explicit
`VLLM_VERSION` override, with local-version suffix handling and UT
environment isolation retained; no hard-coded release-SHA inference
remains.
- Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving
other validations and EPLB platform binding. Rebuild dependent Pydantic
schemas so nested configuration validation uses the patched validator.
The global patch documentation records its rationale and removal
criteria.






- Preserve #15747's Spec+PP protocol/partition handling after rebase;
use the common exact-release selector and the real function-local DSpark
sharing import.






This release-only upgrade does not require a new main old-to-new
interface scan. The latest rebase incorporates
[#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files),
which reverted #16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The
now-unnecessary V4.1 drafter import gate and tuple annotation have been
removed. Other release adaptations, including #15747 Spec+PP handling,
remain. Detailed contract evidence is retained below for review.






<details>





<summary>Per-file adaptations and exact upstream evidence</summary>






#### Per-file adaptation ledger











Evidence IDs refer to the exact source contract and upstream diff table
below. Every row is syntax checked; branch-normalized AST comparison
confirms unchanged main function bodies except the version identity
helper and the explicitly retired propose argument/type annotations.
Both supported versions completed CPU and NPU execution as recorded
below.






| Ascend file / symbols | Disposition and upstream evidence |
Verification |
|---|---|---|





| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` /
`_load_dspark_model_with_target_quant`;
`tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased
[Ascend
#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a)
(`82df9d871`), including manual PP partition masking. vLLM [#52809
diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share`
inside the loader on both supported pins, so retain the earlier release
fix: patch `eagle_utils`, never read/patch a nonexistent
`dspark_utils._should_share`. v0.29 keeps its global PP guard; main has
native PP via
[#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks
the real function-local import, absent module alias, and restoration on
success/failure for both lanes; partition and PP assertions retained.
Ruff/syntax pass; actual CPU/NPU pending. |
| `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`;
`tests/ut/worker/v2/test_pp_utils.py` | #15747 added broad 0.28/0.29
routing and an obsolete 0.28 dev-build recognition path. For the two
supported pins, #50514 exists only on fixed main. Route through
`vllm_version_is("0.29.0")`; retain package/local-suffix and explicit
environment-override semantics. No release-SHA inference or third
release lane. | Existing UTs exercise the real uncached version helper
with monkeypatch isolation, local suffix, fixed-main dev string,
explicit override, and non-target versions. Isolated routing checks and
Ruff/syntax pass; full CPU UT pending. |
| `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`,
`_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass`
| #56078; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`,
`_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436;
use shared standardized layouts and retain the v0.29.0 InputBatch gate.
The rebase preserves vllm-ascend #16043's MTP copy tracking while
removing only legacy v0.28.0 allocation paths. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` |
#56078; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` |
#52839; both lanes use the common PCP import. Rebase keeps current
upstream graph code and removes only the v0.28.0 import branch. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module
imports/dispatch` | #52839; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch`
| #51358, #54853; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/core/recompute_scheduler.py`<br>`module
imports/dispatch`, `schedule` | #51358, #54853; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`,
`AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker`
| #52615; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__`
| #53614; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module
imports/dispatch`, `_get_max_layers_per_page_size`,
`_ascend_max_memory_usage_bytes_from_groups`,
`_ascend_get_kv_cache_config_from_groups` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module
imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` |
v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it
after KV binding. Evidence: [#52506,
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils
diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c).
Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. |
Failure reproduced on v0.29.0; Historical release/main NPU validation
passed; current-head CI pending |
| `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module
imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module
imports/dispatch` | #51718; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant`
| #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports
`_should_share` locally from Eagle utilities; fixed main removes the PP
guard. Keep the release `get_pp_group` patch, share through the common
Eagle utility, and delete the obsolete v0.28.0
`dspark_utils._should_share` patch. | Exact release failure reproduced;
Historical release/main CPU/NPU validation passed; current-head CI
pending |
|
`vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`,
`__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`,
`register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`,
`_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`,
`_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers |
#51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and
import the NaN helpers directly because v0.29.0 and fixed main expose
the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static
checked; historical dual-version CPU/NPU validation passed; current-head
CI pending |
| `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes
`pcp_manager` common to both lanes. The rebase preserves vllm-ascend
#16409's host-parameter-update revert and removes only the obsolete
v0.28.0 capture branch. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`,
`_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/worker/v2/block_table.py`<br>`__init__`,
`init_block_table_layout_tensors`, `compute_slot_mappings` | #51718,
#51031; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`,
`prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`,
`execute_model` | #50514, #54436, #52506, #55212, #53515; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch`
| #54282; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` |
#49811; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
|
`vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module
imports/dispatch`, `propose` | #53694, #52188; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`,
`wake_up` | #51718, #53508; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |






#### Exact upstream evidence











| Upstream change | Full commit SHA / direct diff | Actual supported
contracts and branch decision |
|---|---|---|





| #51718 |
`8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Both use layers/layer_stride/block_stride/offset, tokens_per_state,
CircularBufferSpec and standardized backing; retire
shared_by/compress_ratio allocation branches. |
| #52839 |
`58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8)
| Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
`e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec)
| Both use Mamba copy-function dictionaries and unwrap
UniformTypeKVCacheSpecs. |
| #53106 |
`1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21)
| Both use WeightsMapper instead of AutoWeightsLoader
skip_prefixes/skip_substrs. |
| #53906 |
`98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Only pinned main has the optional MLA storage_block_size dataclass
field; release keeps the Ascend derived property, using
tokens_per_state. |
| #56078 |
`719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417)
| Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned
main uses mrope_num_dims and unified RoPE. |
| #52615 |
`138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045)
| Release uses num_blocks/kv_bytes_per_block; main uses
num_chunks/kv_bytes_per_chunk. |
| #51358 |
`6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release now has boundary_state_offloads and KVConnectorBlockState;
remove partial_tail_offloads plumbing. |
| #54853 |
`0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release constructor takes block_ids snapshots; main takes req_ids and
resolve_block_ids. Keep exact release snapshot membership and main
lazy-resolution membership. |
| #53614 |
`144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725)
| Only main configures drop_eagle_checkpoint_block for replay-aligned
Mamba checkpoints. |
| #50514 |
`d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Release retains module-level PP/share symbols and Ascend PP
workaround; main has the subsequent PP integration. |
| #52809 |
`91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark
module binding to a function-local import from Eagle utilities. The
shared Eagle patch remains effective; the old DSpark-module read/write
must be removed. |
| #54436 |
`6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902)
| Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Only main accepts valid_dummy_state_slots/valid_state_slots capture
arguments. |
| #55212 |
`83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Release prepares DCP local sequence lengths before partitioning; main
initializes DCP metadata afterwards. |
| #53515 |
`b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127)
| Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
`b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212)
| Both accept pcp_manager during graph capture. |
| #51031 |
`0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1)
| Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
`fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8)
| Both gumbel sampling APIs include is_drafting. |
| #52188 |
`d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0)
| Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
`5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6)
| Both propose APIs take DPSyncState; remove the obsolete token-count
argument and retain replicated-PCP synchronization. |
| #49811 |
`01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e)
| Both support extract_hidden_states on MRV2; remove old unsupported
dispatch/skip. |
| #53508 |
`479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29)
| Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
`3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1)
| Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain
release exclusion. |
| #52861 |
`b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28)
| Both include DeepseekV32MTPModel in the two-hidden-state architecture
set. |
| #54713 |
`b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3)
| Only main takes replay_boundaries in compressed-prefix hit lookup;
preserve release calls without that keyword. |
| #42785 |
`442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18)
| Only main capture callers pass axis_keys; preserve the existing Ascend
rejection of nonempty axes. |
| #52358 |
`8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
`9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4)
| Release already has mamba_has_prefill_checkpoint_blocks; later main
also has fine-grained prefix-cache state. |
| #51251 |
`7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| v0.29.0 and main expose ec_manager_config; retire the old release-only
ScoreEncoder configuration skip. |
| #53240 |
`b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c)
| v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache
groups; retire the old release-only replay skip. |
| #53853 |
`e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
`4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both expose the V1 unsupported-feature helper used by the existing
Ascend MRV1 feature filter. |






#### Additional inherited contracts











| vllm-ascend change | Why it is required | Upstream cause and direct
link | Lane | Verification |
|---|---|---|---|---|





| `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM
v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec,
list[int]]` from `get_mamba_groups` and both initialize `recoverssm`;
keeping the old fallback would preserve an unsupported third contract |
[v0.29.0 mamba
groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704),
[fixed-main mamba
groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704),
[v0.29.0
RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103),
[fixed-main
RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103)
| both, common implementation | Existing constructor UT now asserts the
parent-created RecoverSSM value is retained; Ruff and compileall pass |
| `worker/v2/model_runner.py`: always forward
`kv_cache_allocation_context` | v0.29.0 and fixed main both accept this
keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0
signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538),
[fixed-main
signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566),
[vllm-ascend
#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) |
both, common implementation | Existing UT continues to assert the exact
context object reaches the parent; Ruff and compileall pass |
| `_310p/worker/v2/model_runner.py`: select the release KV-zeroing
contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes
`KVCacheGroupSpec.is_eagle_group` but lacks
`SpeculativeConfig.use_eagle_block_drop`; fixed main added the method |
vLLM [#53388
diff](https://github.com/vllm-project/vllm/pull/53388/files), commit
[`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a);
[v0.29.0 group
field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200),
[fixed-main
helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876)
| release differs from main | Existing two-path KV-zeroing UT retained
and renamed for v0.29.0; Ruff and compileall pass |
| existing DFlash kernel UT: remove v0.28-only kwarg omission | the
current Ascend kernel accepts the CP arguments and the only excluded
lane was v0.28.0, which this PR replaces | [vllm-ascend
#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) |
both, common invocation | Existing NPU test remains enabled with all
assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |





|---|---|---|---|





| `patch/platform/patch_parallel_config.py`, registration, and global
patch documentation | Allow Ascend PCP+DP by removing the generic GPU
restriction, following vLLM
[#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit
`7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic
`parallel_config.current_platform` lookup and rebuild ParallelConfig →
SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing
configuration/EPLB UTs and historical PCP+DP NPU execution passed;
current-head CI pending. |
| `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT |
Both supported contracts require `is_drafting`, from
[#54282](https://github.com/vllm-project/vllm/pull/54282/files),
`fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited
optimized kernel and use a common wrapper. | Both | Existing positive
drafting assertion retained; current-head CI pending. |
| `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache
UT | Both support `cache_hit_alignment_tokens`, introduced by
[#53598](https://github.com/vllm-project/vllm/pull/53598/files),
`2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only
write-mask branch. | Both | Existing assertions retained; current-head
CI pending. |






#### Rebase and retired-fallback evidence











| Change | Exact source evidence | Decision |





|---|---|---|





| `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM
[#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95),
`d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL
before v0.29.0; both exact supported sources lack it. The old
conditional came from vllm-ascend
[#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0
selector and use the existing exclusion for both supported lanes. This
does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy
skip. |
| `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and
`tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM
[#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29),
`12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export
`nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The
v0.28.0 fallback originated in vllm-ascend
[#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete
import-failure/`None` fallbacks and the existing UT's obsolete
availability skip; assertions remain unchanged and execute on both
lanes. |
| `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend
[#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3),
`799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream
revert while resolving the real rebase conflict; do not reintroduce the
reverted graph-update behavior. |
| `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend
[#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2),
`d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged
310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation
branch. |






</details>













</details>



### Does this PR introduce _any_ user-facing change?



Yes. The supported vLLM release and default container builds move to
v0.29.0, while the fixed main remains supported. The vllm-ascend package
version does not change. Existing 0.29-only PCP+DP compatibility is
restored.


### How was this patch tested?

- Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`:
[run
35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654)
succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected
NPU jobs confirm the requested Ascend head, integration base
`2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM
pins/installations: fixed main
`84030bbe3d74d99bad477a3d2e37a973ccd8865c` /
`0.1.dev1+g84030bbe3.empty`, release
`98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane
totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest
invocations. Skips/xfails are not passes. The resulting rebased
integration commit is not printed and is not inferred. Actual release
CPU and failed/cancelled image variants remain gaps.


- Current-head CI update (2026-09-20): [main CPU
UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955)
passed **5157 tests**, with **67 skipped**; actual installed vLLM was
`0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head;
the log does not print the full resulting integration head/base.
Pre-commit and mypy passed. NPU/E2E results are recorded above; actual
release CPU remains unverified.
- Image build is partially blocked by infrastructure: A5 amd64
[Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005)
and
[openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068)
failed before reading the Dockerfile because BuildKit could not create a
snapshot temporary directory (`no space left on device`). No release
compatibility code change is justified by this failure; cancelled
variants remain unverified. The successful [310P openEuler arm64
build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987)
explicitly checked out release
`98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`;
this is build evidence, not runtime UT coverage.


- Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto
the baseline above. Previous-head CPU [job
106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458)
stopped during Ascend integration rebase after #16803 merged, before any
UT ran. Its vLLM checkout was the fixed main and installation reported
`0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved;
fresh-head CI was retriggered by the push.


- Remote official v0.29.0 tag resolved to the exact SHA above; both
source pins inspected.
- All 95 changed Python files pass Ruff lint, Ruff format and AST
parsing; `git diff --check` passes. Dockerfile tag/marker consistency
checked across all eight variants.
- Full pre-commit invocation: Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-dependent hooks cannot run on this
Windows host; Python launcher hooks exit 9009. Full CI lint remains
pending.
- Actual main/release CPU/NPU and image builds are not claimed locally:
the required Linux/NPU/container environment is unavailable. New PR E2E
and image-build CI are requested. Existing CPU workflow runs fixed main
only, so actual release CPU remains a validation gap.
- #16393's historical green CI is not a substitute for this new head. No
new test functions, workflow changes, golden/threshold changes or
additional skips are introduced beyond restoring #16393. Inherited
0.28-only tests/skips are not counted as passes.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
weijinqian0 pushed a commit that referenced this pull request Sep 22, 2026
…16910)

### What this PR does / why we need it?

Brings GLM-5.3-Flash prefetch/decode disaggregation to the default model
runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing the
kernel-block layout prefract landed for MRV2 (#16755). Without this, PD
disaggregation only works with VLLM_USE_V2_MODEL_RUNNER=1; the default
runner rejects the layout at KV init and the cross-TP handshake fails.

Three adaptation points:
- V1 reshape: two small-slot view branches mirrored from
_reshape_kv_cache_v2 – the per-request kpool tail ring packs a
contiguous suffix of the shared small slot, and the compressed indexer
packs natural kernel blocks as a contiguous prefix, each bounded to half
the slot.
- GLM-Next cache layout: size the shared small slot by the unified page.
_align_glm5_next_cache_specs derived the small page from the small
candidates only, so with MRV1's unpadded worker specs the slot collapsed
to the indexer's own page and the half-slot bound rejected the layout at
KV init. MRV2 workers pre-pad their specs, so their behavior is
unchanged.
- V1 reshape: expose the NOPE main as natural 128-row kernel blocks
instead of one page-sized block per scheduler block, so dim0 blocks stay
byte-identical across engines with different TP. The page-sized form
made the MooncakeConnectorV2 handshake reject cross-TP transfers with
different local and remote kernel block size 1152 | 640.

Depends on #16755.

### Does this PR introduce any user-facing change?

No new flags, config options, or API changes. For users running
GLM-5.3-Flash with PD disaggregation on the default model runner,
cross-TP deployments (e.g. P=TP8 + D=DP2xTP4) now work out of the box.
All other models and colocated workloads are unaffected – branch
selection is identical to before (only GLM-Next publishes the stride
marker).

How was this patch tested?

On GLM-5.3-Flash w8a8 real hardware (TP8 single instance, P=TP8 + D=TP8,
and P=TP8 → D=DP2xTP4):
- Indexer and tail views bit-identical with the MRV2 reshape.
- Equal-TP full-verify passes (Serial + concurrent, all match; per-rank
ms-level pulls).
- Unequal-TP chain: handshake clean, three factual questions all
answered correctly, zero pull/transfer errors.
- TextVQA authoritative rescore (per-row gold alignment, post-
extraction): PD 92.0% identical to colocated 90.0% in the same batch.


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: sunbaosong <13793883820@163.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
…llm-project#16755)

### What this PR does / why we need it?

Enables GLM-Next (GLM-5.3-Flash) prefill/decode disaggregation on Model
Runner V2 with `MooncakeConnectorV2`, for both equal TP (P=TP4 / D=TP4)
and unequal TP (P=TP8 / D=DP2xTP4).

GLM-Next could not run PD on MRV2 because of three structural gaps:

1. **Granularity mismatch.** MRV2 hands metadata builders
kernel-granularity specs (`block_size` == the C128 SFA kernel size),
while the scheduler, the common block tables, and the GLM-Next cache
contract all use the logical block from `cache_config.block_size`.
Sizing persistent buffers from the spec overflows them by the
ceil-rounding slack (`required=1044` vs `capacity=1040` on the first
forward).
2. **GLM-Next layout.** The v2 MLA spec re-wrap dropped the
`model_version` / `compress_ratio` markers the cache-group
classification needs; the compressed indexer small page (compress_ratio
> 1) is a single-tensor overlay on a shared padded slot, not an MLA
latent/rope K/V pair; NoPE MLA layers legitimately carry an empty rope/V
component.
3. **Unequal TP.** Page-size unification is TP-dependent for hybrid
models (the Mamba state page scales with heads/TP), so a TP8 producer
and a TP4 consumer publish different unified block sizes for
byte-identical replicated caches — kernel blocks can no longer be paired
index-wise.

Changes:

- v2 attn_utils: propagate `model_version` /
`indexes_kv_by_block_stride` / `compress_ratio` /
`non_causal_multi_token_decode` through the MLA spec re-wrap; reshape
the compressed indexer small-page layer (compress_ratio > 1) as a single
strided tensor overlaid from byte zero of the shared padded slot.
- KPool indexer/tail backends: accept upstream's `cache_dtype_str` probe
kwarg; take the logical block size from `cache_config.block_size`.
- SFA metadata state: size the persistent block table buffer from the
cache config's logical block expansion instead of the spec's.
- Mooncake V2 connector: skip zero-size cache components when collecting
register ranges (their degenerate storage corrupts the range arithmetic)
and skip zero-length transfer entries the engine rejects with E19999.
- Unequal-TP: transfer replicated MLA/indexer caches as GCD-sized
token-segment byte runs (`_compute_cross_tp_runs`); transfer the indexer
tail whole-block instead of through the HND head-sharded path; raise
with the layer name on kernel-block-size mismatch.

### Does this PR introduce _any_ user-facing change?

Yes — it enables a new deployment topology: GLM-Next hybrid models can
now run PD disaggregation on MRV2 with `MooncakeConnectorV2` (pull
mode), including a TP8 producer feeding a DP2xTP4 consumer. Requires
`VLLM_USE_V2_MODEL_RUNNER=1` and a kv-transfer-config with
`MooncakeConnectorV2`. No behavior change for existing models and
configurations: the new code paths are gated (the strided reshape only
triggers for MLA-family specs with compress_ratio > 1; the cross-TP run
path only activates when local and remote kernel block sizes differ; the
equal-TP block pairing path is unchanged).

### How was this patch tested?

- **Unit tests:** the `mooncake_v2` suite passes 150/150, covering
cross-TP run splitting/coalescing/computed-offset/bounds, tail
replica/whole-block mapping, bucket routing, and per-component address
scaling. The full UT sweep over
`tests/ut/{models,attention,worker,core,kv_offload}` produces a failure
set identical to upstream main (pre-existing failures only, none
introduced by this PR).
- **E2E, 10-layer GLM-5.3-Flash debug model:** 1P1D TP4/TP4 (producer
cards 0-3, consumer cards 4-7, eager), pull-mode KV transfer over the
mooncake engine relayed through a load-balance proxy: first-token
fidelity 18/18, outputs match the single-instance baseline to the
model's own determinism, real per-rank transfer logs, zero recompute
fallback.
- **E2E, full w8a8 GLM-5.3-Flash:** P=TP8 producer, D=DP2xTP4 consumer
(FULL_DECODE_ONLY graph mode): zero pull errors, ~18.5 MB transferred
per request, TextVQA accuracy consistent with the colocated baseline.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: sunbaosong <13793883820@163.com>
Co-authored-by: bubaishenhua112-netizen <bubaishenhua112@gmail.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
…vllm-project#17004)

### What this PR does / why we need it?



Reintroduce the vLLM v0.29.0 release upgrade from vllm-project#16393, reverted by
vllm-project#16949, with container defaults aligned to the supported release. Based
on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves
later merged changes.


The image/source mismatch is confirmed in [nightly job
106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020):
the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded
Ascend failed importing `_get_packed_kv_cache_groups`. All eight root
Dockerfiles still defaulted to v0.28.0. The reusable image workflow
passes no `VLLM_TAG` override, so those defaults govern release builds.
Changing only the release marker does not update those images.


- Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

- Release changes from v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0
(`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the
remote tag).
- Restore vllm-project#16393's source compatibility and existing UT changes by
reversing vllm-project#16949, then reconcile current upstream changes. No main
interface scan is rerun for this release-only upgrade.
- Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P,
Ubuntu/openEuler. Preserve existing exact-commit build overrides. No
workflow changes.


| Current-base adaptation | Exact cause and evidence | Lane | Validation
|
| --- | --- | --- | --- |

| Eight Dockerfile `VLLM_TAG` defaults and release marker | vllm-project#16393
changed the release contract but omitted image defaults; the linked
nightly log proves installation of 0.28.0 and missing
`_get_packed_kv_cache_groups`. | Release images | All eight defaults
match marker; image-build CI requested, pending |
| `worker/v2/attn_utils.py::get_kv_cache_spec` and existing
`_make_mla_layer` UT fixture | Ascend
[vllm-project#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files),
`df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field
selector. vLLM
[#51718](https://github.com/vllm-project/vllm/pull/51718/files),
`8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio`
with `tokens_per_state`; both exact supported pins use the latter. Use
the common field, retaining metadata and cache-view assertions. | Both |
Source inspection and static checks passed; actual CI pending |
| `attention/attention_v1.py` import conflict | Preserve
`attention_transfer_window` from Ascend
[vllm-project#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files),
`5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from
vllm-project#16798; keep the common relocated PCP import from vllm-project#16393. Do not restore
the old compute-start import or unused weak-reference import. | Both |
Conflict resolved; static checks passed |
| `worker/v2/aclgraph_utils.py` import conflict | Preserve
`ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend
[vllm-project#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files),
`34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete
0.28 version selector restored by the revert. | Both | Conflict
resolved; static checks passed |


| Existing 310P and Mamba model-runner UT imports | Preserve
hardware-profile imports and mocks from Ascend
[vllm-project#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files),
`d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused
old-release selector import. Hardware capability routing remains
unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format
passed; fresh CI pending |

Newly merged changes were reviewed for version-contract impact:
vllm-project#16775/vllm-project#16834 PCP metadata and routing, vllm-project#16923 A3 SFA,
vllm-project#16426/vllm-project#16924/vllm-project#16955 custom ops, vllm-project#16081 MTP/SP, vllm-project#16669 xlite,
vllm-project#15636/vllm-project#16747 transfer, vllm-project#16913 A5 pages, vllm-project#16798 graph updates, vllm-project#16755
GLM PD, vllm-project#16952 C8 config, vllm-project#16673 operator removal and vllm-project#16320 Kimi-K3 KV
pool. Preserve these changes; no additional source-proven version branch
was identified beyond the entries above. In particular, vllm-project#16747's
`UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size`
exist in both exact pins; vllm-project#16320 adds Ascend connector hooks. CI remains
necessary to validate runtime interactions. CI/documentation-only PRs
are retained unchanged.


<details>

<summary>Inherited per-file compatibility evidence from vllm-project#16393</summary>


The following source-contract ledger is inherited from vllm-project#16393. Any
historical verification wording refers only to that earlier PR; it does
not certify this new head. New-head validation is listed below.


| Reference | Commit |





|---|---|





| vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` |
| PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` |





| Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` |





| Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a`
|
| Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` |






- Use common implementations where v0.29.0 and the fixed main share
KV-cache layouts, Mamba copy/group APIs, PCP handling and
speculative-decoding contracts.
- Retain explicit `vllm_version_is("0.29.0")` branches for contracts
that still differ, including RoPE, scheduler block snapshots,
InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.
- Remove obsolete v0.28.0 compatibility and adapt existing test
fixtures. Version detection uses package versions and the explicit
`VLLM_VERSION` override, with local-version suffix handling and UT
environment isolation retained; no hard-coded release-SHA inference
remains.
- Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving
other validations and EPLB platform binding. Rebuild dependent Pydantic
schemas so nested configuration validation uses the patched validator.
The global patch documentation records its rationale and removal
criteria.






- Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase;
use the common exact-release selector and the real function-local DSpark
sharing import.






This release-only upgrade does not require a new main old-to-new
interface scan. The latest rebase incorporates
[vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files),
which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The
now-unnecessary V4.1 drafter import gate and tuple annotation have been
removed. Other release adaptations, including vllm-project#15747 Spec+PP handling,
remain. Detailed contract evidence is retained below for review.






<details>





<summary>Per-file adaptations and exact upstream evidence</summary>






#### Per-file adaptation ledger











Evidence IDs refer to the exact source contract and upstream diff table
below. Every row is syntax checked; branch-normalized AST comparison
confirms unchanged main function bodies except the version identity
helper and the explicitly retired propose argument/type annotations.
Both supported versions completed CPU and NPU execution as recorded
below.






| Ascend file / symbols | Disposition and upstream evidence |
Verification |
|---|---|---|





| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` /
`_load_dspark_model_with_target_quant`;
`tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased
[Ascend
vllm-project#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a)
(`82df9d871`), including manual PP partition masking. vLLM [#52809
diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share`
inside the loader on both supported pins, so retain the earlier release
fix: patch `eagle_utils`, never read/patch a nonexistent
`dspark_utils._should_share`. v0.29 keeps its global PP guard; main has
native PP via
[#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks
the real function-local import, absent module alias, and restoration on
success/failure for both lanes; partition and PP assertions retained.
Ruff/syntax pass; actual CPU/NPU pending. |
| `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`;
`tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29
routing and an obsolete 0.28 dev-build recognition path. For the two
supported pins, #50514 exists only on fixed main. Route through
`vllm_version_is("0.29.0")`; retain package/local-suffix and explicit
environment-override semantics. No release-SHA inference or third
release lane. | Existing UTs exercise the real uncached version helper
with monkeypatch isolation, local suffix, fixed-main dev string,
explicit override, and non-target versions. Isolated routing checks and
Ruff/syntax pass; full CPU UT pending. |
| `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`,
`_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass`
| #56078; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`,
`_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436;
use shared standardized layouts and retain the v0.29.0 InputBatch gate.
The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while
removing only legacy v0.28.0 allocation paths. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` |
#56078; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` |
#52839; both lanes use the common PCP import. Rebase keeps current
upstream graph code and removes only the v0.28.0 import branch. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module
imports/dispatch` | #52839; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch`
| #51358, #54853; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/core/recompute_scheduler.py`<br>`module
imports/dispatch`, `schedule` | #51358, #54853; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`,
`AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker`
| #52615; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__`
| #53614; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module
imports/dispatch`, `_get_max_layers_per_page_size`,
`_ascend_max_memory_usage_bytes_from_groups`,
`_ascend_get_kv_cache_config_from_groups` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module
imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` |
v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it
after KV binding. Evidence: [#52506,
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils
diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c).
Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. |
Failure reproduced on v0.29.0; Historical release/main NPU validation
passed; current-head CI pending |
| `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module
imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module
imports/dispatch` | #51718; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant`
| #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports
`_should_share` locally from Eagle utilities; fixed main removes the PP
guard. Keep the release `get_pp_group` patch, share through the common
Eagle utility, and delete the obsolete v0.28.0
`dspark_utils._should_share` patch. | Exact release failure reproduced;
Historical release/main CPU/NPU validation passed; current-head CI
pending |
|
`vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`,
`__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`,
`register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`,
`_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`,
`_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers |
#51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and
import the NaN helpers directly because v0.29.0 and fixed main expose
the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static
checked; historical dual-version CPU/NPU validation passed; current-head
CI pending |
| `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes
`pcp_manager` common to both lanes. The rebase preserves vllm-ascend
vllm-project#16409's host-parameter-update revert and removes only the obsolete
v0.28.0 capture branch. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`,
`_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/worker/v2/block_table.py`<br>`__init__`,
`init_block_table_layout_tensors`, `compute_slot_mappings` | #51718,
#51031; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`,
`prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`,
`execute_model` | #50514, #54436, #52506, #55212, #53515; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch`
| #54282; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` |
#49811; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
|
`vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module
imports/dispatch`, `propose` | #53694, #52188; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`,
`wake_up` | #51718, #53508; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |






#### Exact upstream evidence











| Upstream change | Full commit SHA / direct diff | Actual supported
contracts and branch decision |
|---|---|---|





| #51718 |
`8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Both use layers/layer_stride/block_stride/offset, tokens_per_state,
CircularBufferSpec and standardized backing; retire
shared_by/compress_ratio allocation branches. |
| #52839 |
`58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8)
| Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
`e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec)
| Both use Mamba copy-function dictionaries and unwrap
UniformTypeKVCacheSpecs. |
| #53106 |
`1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21)
| Both use WeightsMapper instead of AutoWeightsLoader
skip_prefixes/skip_substrs. |
| #53906 |
`98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Only pinned main has the optional MLA storage_block_size dataclass
field; release keeps the Ascend derived property, using
tokens_per_state. |
| #56078 |
`719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417)
| Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned
main uses mrope_num_dims and unified RoPE. |
| #52615 |
`138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045)
| Release uses num_blocks/kv_bytes_per_block; main uses
num_chunks/kv_bytes_per_chunk. |
| #51358 |
`6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release now has boundary_state_offloads and KVConnectorBlockState;
remove partial_tail_offloads plumbing. |
| #54853 |
`0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release constructor takes block_ids snapshots; main takes req_ids and
resolve_block_ids. Keep exact release snapshot membership and main
lazy-resolution membership. |
| #53614 |
`144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725)
| Only main configures drop_eagle_checkpoint_block for replay-aligned
Mamba checkpoints. |
| #50514 |
`d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Release retains module-level PP/share symbols and Ascend PP
workaround; main has the subsequent PP integration. |
| #52809 |
`91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark
module binding to a function-local import from Eagle utilities. The
shared Eagle patch remains effective; the old DSpark-module read/write
must be removed. |
| #54436 |
`6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902)
| Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Only main accepts valid_dummy_state_slots/valid_state_slots capture
arguments. |
| #55212 |
`83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Release prepares DCP local sequence lengths before partitioning; main
initializes DCP metadata afterwards. |
| #53515 |
`b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127)
| Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
`b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212)
| Both accept pcp_manager during graph capture. |
| #51031 |
`0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1)
| Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
`fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8)
| Both gumbel sampling APIs include is_drafting. |
| #52188 |
`d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0)
| Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
`5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6)
| Both propose APIs take DPSyncState; remove the obsolete token-count
argument and retain replicated-PCP synchronization. |
| #49811 |
`01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e)
| Both support extract_hidden_states on MRV2; remove old unsupported
dispatch/skip. |
| #53508 |
`479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29)
| Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
`3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1)
| Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain
release exclusion. |
| #52861 |
`b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28)
| Both include DeepseekV32MTPModel in the two-hidden-state architecture
set. |
| #54713 |
`b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3)
| Only main takes replay_boundaries in compressed-prefix hit lookup;
preserve release calls without that keyword. |
| #42785 |
`442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18)
| Only main capture callers pass axis_keys; preserve the existing Ascend
rejection of nonempty axes. |
| #52358 |
`8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
`9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4)
| Release already has mamba_has_prefill_checkpoint_blocks; later main
also has fine-grained prefix-cache state. |
| #51251 |
`7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| v0.29.0 and main expose ec_manager_config; retire the old release-only
ScoreEncoder configuration skip. |
| #53240 |
`b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c)
| v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache
groups; retire the old release-only replay skip. |
| #53853 |
`e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
`4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both expose the V1 unsupported-feature helper used by the existing
Ascend MRV1 feature filter. |






#### Additional inherited contracts











| vllm-ascend change | Why it is required | Upstream cause and direct
link | Lane | Verification |
|---|---|---|---|---|





| `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM
v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec,
list[int]]` from `get_mamba_groups` and both initialize `recoverssm`;
keeping the old fallback would preserve an unsupported third contract |
[v0.29.0 mamba
groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704),
[fixed-main mamba
groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704),
[v0.29.0
RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103),
[fixed-main
RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103)
| both, common implementation | Existing constructor UT now asserts the
parent-created RecoverSSM value is retained; Ruff and compileall pass |
| `worker/v2/model_runner.py`: always forward
`kv_cache_allocation_context` | v0.29.0 and fixed main both accept this
keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0
signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538),
[fixed-main
signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566),
[vllm-ascend
vllm-project#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) |
both, common implementation | Existing UT continues to assert the exact
context object reaches the parent; Ruff and compileall pass |
| `_310p/worker/v2/model_runner.py`: select the release KV-zeroing
contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes
`KVCacheGroupSpec.is_eagle_group` but lacks
`SpeculativeConfig.use_eagle_block_drop`; fixed main added the method |
vLLM [#53388
diff](https://github.com/vllm-project/vllm/pull/53388/files), commit
[`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a);
[v0.29.0 group
field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200),
[fixed-main
helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876)
| release differs from main | Existing two-path KV-zeroing UT retained
and renamed for v0.29.0; Ruff and compileall pass |
| existing DFlash kernel UT: remove v0.28-only kwarg omission | the
current Ascend kernel accepts the CP arguments and the only excluded
lane was v0.28.0, which this PR replaces | [vllm-ascend
vllm-project#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) |
both, common invocation | Existing NPU test remains enabled with all
assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |





|---|---|---|---|





| `patch/platform/patch_parallel_config.py`, registration, and global
patch documentation | Allow Ascend PCP+DP by removing the generic GPU
restriction, following vLLM
[#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit
`7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic
`parallel_config.current_platform` lookup and rebuild ParallelConfig →
SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing
configuration/EPLB UTs and historical PCP+DP NPU execution passed;
current-head CI pending. |
| `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT |
Both supported contracts require `is_drafting`, from
[#54282](https://github.com/vllm-project/vllm/pull/54282/files),
`fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited
optimized kernel and use a common wrapper. | Both | Existing positive
drafting assertion retained; current-head CI pending. |
| `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache
UT | Both support `cache_hit_alignment_tokens`, introduced by
[#53598](https://github.com/vllm-project/vllm/pull/53598/files),
`2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only
write-mask branch. | Both | Existing assertions retained; current-head
CI pending. |






#### Rebase and retired-fallback evidence











| Change | Exact source evidence | Decision |





|---|---|---|





| `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM
[#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95),
`d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL
before v0.29.0; both exact supported sources lack it. The old
conditional came from vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0
selector and use the existing exclusion for both supported lanes. This
does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy
skip. |
| `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and
`tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM
[#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29),
`12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export
`nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The
v0.28.0 fallback originated in vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete
import-failure/`None` fallbacks and the existing UT's obsolete
availability skip; assertions remain unchanged and execute on both
lanes. |
| `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend
[vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3),
`799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream
revert while resolving the real rebase conflict; do not reintroduce the
reverted graph-update behavior. |
| `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend
[vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2),
`d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged
310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation
branch. |






</details>













</details>



### Does this PR introduce _any_ user-facing change?



Yes. The supported vLLM release and default container builds move to
v0.29.0, while the fixed main remains supported. The vllm-ascend package
version does not change. Existing 0.29-only PCP+DP compatibility is
restored.


### How was this patch tested?

- Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`:
[run
35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654)
succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected
NPU jobs confirm the requested Ascend head, integration base
`2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM
pins/installations: fixed main
`84030bbe3d74d99bad477a3d2e37a973ccd8865c` /
`0.1.dev1+g84030bbe3.empty`, release
`98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane
totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest
invocations. Skips/xfails are not passes. The resulting rebased
integration commit is not printed and is not inferred. Actual release
CPU and failed/cancelled image variants remain gaps.


- Current-head CI update (2026-09-20): [main CPU
UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955)
passed **5157 tests**, with **67 skipped**; actual installed vLLM was
`0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head;
the log does not print the full resulting integration head/base.
Pre-commit and mypy passed. NPU/E2E results are recorded above; actual
release CPU remains unverified.
- Image build is partially blocked by infrastructure: A5 amd64
[Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005)
and
[openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068)
failed before reading the Dockerfile because BuildKit could not create a
snapshot temporary directory (`no space left on device`). No release
compatibility code change is justified by this failure; cancelled
variants remain unverified. The successful [310P openEuler arm64
build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987)
explicitly checked out release
`98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`;
this is build evidence, not runtime UT coverage.


- Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto
the baseline above. Previous-head CPU [job
106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458)
stopped during Ascend integration rebase after vllm-project#16803 merged, before any
UT ran. Its vLLM checkout was the fixed main and installation reported
`0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved;
fresh-head CI was retriggered by the push.


- Remote official v0.29.0 tag resolved to the exact SHA above; both
source pins inspected.
- All 95 changed Python files pass Ruff lint, Ruff format and AST
parsing; `git diff --check` passes. Dockerfile tag/marker consistency
checked across all eight variants.
- Full pre-commit invocation: Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-dependent hooks cannot run on this
Windows host; Python launcher hooks exit 9009. Full CI lint remains
pending.
- Actual main/release CPU/NPU and image builds are not claimed locally:
the required Linux/NPU/container environment is unavailable. New PR E2E
and image-build CI are requested. Existing CPU workflow runs fixed main
only, so actual release CPU remains a validation gap.
- vllm-project#16393's historical green CI is not a substitute for this new head. No
new test functions, workflow changes, golden/threshold changes or
additional skips are introduced beyond restoring vllm-project#16393. Inherited
0.28-only tests/skips are not counted as passes.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
…llm-project#16755)

### What this PR does / why we need it?

Enables GLM-Next (GLM-5.3-Flash) prefill/decode disaggregation on Model
Runner V2 with `MooncakeConnectorV2`, for both equal TP (P=TP4 / D=TP4)
and unequal TP (P=TP8 / D=DP2xTP4).

GLM-Next could not run PD on MRV2 because of three structural gaps:

1. **Granularity mismatch.** MRV2 hands metadata builders
kernel-granularity specs (`block_size` == the C128 SFA kernel size),
while the scheduler, the common block tables, and the GLM-Next cache
contract all use the logical block from `cache_config.block_size`.
Sizing persistent buffers from the spec overflows them by the
ceil-rounding slack (`required=1044` vs `capacity=1040` on the first
forward).
2. **GLM-Next layout.** The v2 MLA spec re-wrap dropped the
`model_version` / `compress_ratio` markers the cache-group
classification needs; the compressed indexer small page (compress_ratio
> 1) is a single-tensor overlay on a shared padded slot, not an MLA
latent/rope K/V pair; NoPE MLA layers legitimately carry an empty rope/V
component.
3. **Unequal TP.** Page-size unification is TP-dependent for hybrid
models (the Mamba state page scales with heads/TP), so a TP8 producer
and a TP4 consumer publish different unified block sizes for
byte-identical replicated caches — kernel blocks can no longer be paired
index-wise.

Changes:

- v2 attn_utils: propagate `model_version` /
`indexes_kv_by_block_stride` / `compress_ratio` /
`non_causal_multi_token_decode` through the MLA spec re-wrap; reshape
the compressed indexer small-page layer (compress_ratio > 1) as a single
strided tensor overlaid from byte zero of the shared padded slot.
- KPool indexer/tail backends: accept upstream's `cache_dtype_str` probe
kwarg; take the logical block size from `cache_config.block_size`.
- SFA metadata state: size the persistent block table buffer from the
cache config's logical block expansion instead of the spec's.
- Mooncake V2 connector: skip zero-size cache components when collecting
register ranges (their degenerate storage corrupts the range arithmetic)
and skip zero-length transfer entries the engine rejects with E19999.
- Unequal-TP: transfer replicated MLA/indexer caches as GCD-sized
token-segment byte runs (`_compute_cross_tp_runs`); transfer the indexer
tail whole-block instead of through the HND head-sharded path; raise
with the layer name on kernel-block-size mismatch.

### Does this PR introduce _any_ user-facing change?

Yes — it enables a new deployment topology: GLM-Next hybrid models can
now run PD disaggregation on MRV2 with `MooncakeConnectorV2` (pull
mode), including a TP8 producer feeding a DP2xTP4 consumer. Requires
`VLLM_USE_V2_MODEL_RUNNER=1` and a kv-transfer-config with
`MooncakeConnectorV2`. No behavior change for existing models and
configurations: the new code paths are gated (the strided reshape only
triggers for MLA-family specs with compress_ratio > 1; the cross-TP run
path only activates when local and remote kernel block sizes differ; the
equal-TP block pairing path is unchanged).

### How was this patch tested?

- **Unit tests:** the `mooncake_v2` suite passes 150/150, covering
cross-TP run splitting/coalescing/computed-offset/bounds, tail
replica/whole-block mapping, bucket routing, and per-component address
scaling. The full UT sweep over
`tests/ut/{models,attention,worker,core,kv_offload}` produces a failure
set identical to upstream main (pre-existing failures only, none
introduced by this PR).
- **E2E, 10-layer GLM-5.3-Flash debug model:** 1P1D TP4/TP4 (producer
cards 0-3, consumer cards 4-7, eager), pull-mode KV transfer over the
mooncake engine relayed through a load-balance proxy: first-token
fidelity 18/18, outputs match the single-instance baseline to the
model's own determinism, real per-rank transfer logs, zero recompute
fallback.
- **E2E, full w8a8 GLM-5.3-Flash:** P=TP8 producer, D=DP2xTP4 consumer
(FULL_DECODE_ONLY graph mode): zero pull errors, ~18.5 MB transferred
per request, TextVQA accuracy consistent with the colocated baseline.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: sunbaosong <13793883820@163.com>
Co-authored-by: bubaishenhua112-netizen <bubaishenhua112@gmail.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
…vllm-project#17004)

Reintroduce the vLLM v0.29.0 release upgrade from vllm-project#16393, reverted by
on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves
later merged changes.

The image/source mismatch is confirmed in [nightly job
106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020):
the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded
Ascend failed importing `_get_packed_kv_cache_groups`. All eight root
Dockerfiles still defaulted to v0.28.0. The reusable image workflow
passes no `VLLM_TAG` override, so those defaults govern release builds.
Changing only the release marker does not update those images.

- Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`.
- Release changes from v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0
(`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the
remote tag).
- Restore vllm-project#16393's source compatibility and existing UT changes by
reversing vllm-project#16949, then reconcile current upstream changes. No main
interface scan is rerun for this release-only upgrade.
- Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P,
Ubuntu/openEuler. Preserve existing exact-commit build overrides. No
workflow changes.

| Current-base adaptation | Exact cause and evidence | Lane | Validation
|
| --- | --- | --- | --- |
| Eight Dockerfile `VLLM_TAG` defaults and release marker | vllm-project#16393
changed the release contract but omitted image defaults; the linked
nightly log proves installation of 0.28.0 and missing
`_get_packed_kv_cache_groups`. | Release images | All eight defaults
match marker; image-build CI requested, pending |
| `worker/v2/attn_utils.py::get_kv_cache_spec` and existing
`_make_mla_layer` UT fixture | Ascend
[vllm-project#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files),
`df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field
selector. vLLM
[#51718](https://github.com/vllm-project/vllm/pull/51718/files),
`8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio`
with `tokens_per_state`; both exact supported pins use the latter. Use
the common field, retaining metadata and cache-view assertions. | Both |
Source inspection and static checks passed; actual CI pending |
| `attention/attention_v1.py` import conflict | Preserve
`attention_transfer_window` from Ascend
[vllm-project#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files),
`5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from
the old compute-start import or unused weak-reference import. | Both |
Conflict resolved; static checks passed |
| `worker/v2/aclgraph_utils.py` import conflict | Preserve
`ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend
[vllm-project#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files),
`34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete
0.28 version selector restored by the revert. | Both | Conflict
resolved; static checks passed |

| Existing 310P and Mamba model-runner UT imports | Preserve
hardware-profile imports and mocks from Ascend
[vllm-project#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files),
`d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused
old-release selector import. Hardware capability routing remains
unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format
passed; fresh CI pending |

Newly merged changes were reviewed for version-contract impact:
GLM PD, vllm-project#16952 C8 config, vllm-project#16673 operator removal and vllm-project#16320 Kimi-K3 KV
pool. Preserve these changes; no additional source-proven version branch
was identified beyond the entries above. In particular, vllm-project#16747's
`UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size`
exist in both exact pins; vllm-project#16320 adds Ascend connector hooks. CI remains
necessary to validate runtime interactions. CI/documentation-only PRs
are retained unchanged.

<details>
<summary>Inherited per-file compatibility evidence from vllm-project#16393</summary>

The following source-contract ledger is inherited from vllm-project#16393. Any
historical verification wording refers only to that earlier PR; it does
not certify this new head. New-head validation is listed below.

| Reference | Commit |
|---|---|
| vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` |
| PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` |
| Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` |
| Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a`
|
| Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` |

- Use common implementations where v0.29.0 and the fixed main share
KV-cache layouts, Mamba copy/group APIs, PCP handling and
speculative-decoding contracts.
- Retain explicit `vllm_version_is("0.29.0")` branches for contracts
that still differ, including RoPE, scheduler block snapshots,
InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.
- Remove obsolete v0.28.0 compatibility and adapt existing test
fixtures. Version detection uses package versions and the explicit
`VLLM_VERSION` override, with local-version suffix handling and UT
environment isolation retained; no hard-coded release-SHA inference
remains.
- Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving
other validations and EPLB platform binding. Rebuild dependent Pydantic
schemas so nested configuration validation uses the patched validator.
The global patch documentation records its rationale and removal
criteria.

- Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase;
use the common exact-release selector and the real function-local DSpark
sharing import.

This release-only upgrade does not require a new main old-to-new
interface scan. The latest rebase incorporates
[vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files),
which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The
now-unnecessary V4.1 drafter import gate and tuple annotation have been
removed. Other release adaptations, including vllm-project#15747 Spec+PP handling,
remain. Detailed contract evidence is retained below for review.

<details>
<summary>Per-file adaptations and exact upstream evidence</summary>

Evidence IDs refer to the exact source contract and upstream diff table
below. Every row is syntax checked; branch-normalized AST comparison
confirms unchanged main function bodies except the version identity
helper and the explicitly retired propose argument/type annotations.
Both supported versions completed CPU and NPU execution as recorded
below.

| Ascend file / symbols | Disposition and upstream evidence |
Verification |
|---|---|---|
| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` /
`_load_dspark_model_with_target_quant`;
`tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased
[Ascend
(`82df9d871`), including manual PP partition masking. vLLM [#52809
diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share`
inside the loader on both supported pins, so retain the earlier release
fix: patch `eagle_utils`, never read/patch a nonexistent
`dspark_utils._should_share`. v0.29 keeps its global PP guard; main has
native PP via
[#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks
the real function-local import, absent module alias, and restoration on
success/failure for both lanes; partition and PP assertions retained.
Ruff/syntax pass; actual CPU/NPU pending. |
| `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`;
`tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29
routing and an obsolete 0.28 dev-build recognition path. For the two
supported pins, #50514 exists only on fixed main. Route through
`vllm_version_is("0.29.0")`; retain package/local-suffix and explicit
environment-override semantics. No release-SHA inference or third
release lane. | Existing UTs exercise the real uncached version helper
with monkeypatch isolation, local suffix, fixed-main dev string,
explicit override, and non-target versions. Isolated routing checks and
Ruff/syntax pass; full CPU UT pending. |
| `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`,
`_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass`
| #56078; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`,
`_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436;
use shared standardized layouts and retain the v0.29.0 InputBatch gate.
The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while
removing only legacy v0.28.0 allocation paths. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` |
upstream graph code and removes only the v0.28.0 import branch. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module
imports/dispatch` | #52839; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch`
| #51358, #54853; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/core/recompute_scheduler.py`<br>`module
imports/dispatch`, `schedule` | #51358, #54853; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`,
`AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker`
| #52615; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__`
| #53614; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module
imports/dispatch`, `_get_max_layers_per_page_size`,
`_ascend_max_memory_usage_bytes_from_groups`,
`_ascend_get_kv_cache_config_from_groups` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module
imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` |
v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it
after KV binding. Evidence: [#52506,
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils
diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c).
Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. |
Failure reproduced on v0.29.0; Historical release/main NPU validation
passed; current-head CI pending |
| `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module
imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module
imports/dispatch` | #51718; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant`
| #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports
`_should_share` locally from Eagle utilities; fixed main removes the PP
guard. Keep the release `get_pp_group` patch, share through the common
Eagle utility, and delete the obsolete v0.28.0
`dspark_utils._should_share` patch. | Exact release failure reproduced;
Historical release/main CPU/NPU validation passed; current-head CI
pending |
|
`vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`,
`__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`,
`register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`,
`_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`,
`_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers |
import the NaN helpers directly because v0.29.0 and fixed main expose
the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static
checked; historical dual-version CPU/NPU validation passed; current-head
CI pending |
| `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes
`pcp_manager` common to both lanes. The rebase preserves vllm-ascend
v0.28.0 capture branch. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`,
`_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/worker/v2/block_table.py`<br>`__init__`,
`init_block_table_layout_tensors`, `compute_slot_mappings` | #51718,
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`,
`prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`,
`execute_model` | #50514, #54436, #52506, #55212, #53515; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch`
| #54282; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` |
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
|
`vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module
imports/dispatch`, `propose` | #53694, #52188; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`,
`wake_up` | #51718, #53508; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |

| Upstream change | Full commit SHA / direct diff | Actual supported
contracts and branch decision |
|---|---|---|
| #51718 |
`8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Both use layers/layer_stride/block_stride/offset, tokens_per_state,
CircularBufferSpec and standardized backing; retire
shared_by/compress_ratio allocation branches. |
| #52839 |
`58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8)
| Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
`e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec)
| Both use Mamba copy-function dictionaries and unwrap
UniformTypeKVCacheSpecs. |
| #53106 |
`1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21)
| Both use WeightsMapper instead of AutoWeightsLoader
skip_prefixes/skip_substrs. |
| #53906 |
`98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Only pinned main has the optional MLA storage_block_size dataclass
field; release keeps the Ascend derived property, using
tokens_per_state. |
| #56078 |
`719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417)
| Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned
main uses mrope_num_dims and unified RoPE. |
| #52615 |
`138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045)
| Release uses num_blocks/kv_bytes_per_block; main uses
num_chunks/kv_bytes_per_chunk. |
| #51358 |
`6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release now has boundary_state_offloads and KVConnectorBlockState;
remove partial_tail_offloads plumbing. |
| #54853 |
`0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release constructor takes block_ids snapshots; main takes req_ids and
resolve_block_ids. Keep exact release snapshot membership and main
lazy-resolution membership. |
| #53614 |
`144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725)
| Only main configures drop_eagle_checkpoint_block for replay-aligned
Mamba checkpoints. |
| #50514 |
`d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Release retains module-level PP/share symbols and Ascend PP
workaround; main has the subsequent PP integration. |
| #52809 |
`91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark
module binding to a function-local import from Eagle utilities. The
shared Eagle patch remains effective; the old DSpark-module read/write
must be removed. |
| #54436 |
`6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902)
| Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Only main accepts valid_dummy_state_slots/valid_state_slots capture
arguments. |
| #55212 |
`83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Release prepares DCP local sequence lengths before partitioning; main
initializes DCP metadata afterwards. |
| #53515 |
`b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127)
| Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
`b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212)
| Both accept pcp_manager during graph capture. |
| #51031 |
`0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1)
| Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
`fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8)
| Both gumbel sampling APIs include is_drafting. |
| #52188 |
`d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0)
| Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
`5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6)
| Both propose APIs take DPSyncState; remove the obsolete token-count
argument and retain replicated-PCP synchronization. |
| #49811 |
`01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e)
| Both support extract_hidden_states on MRV2; remove old unsupported
dispatch/skip. |
| #53508 |
`479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29)
| Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
`3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1)
| Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain
release exclusion. |
| #52861 |
`b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28)
| Both include DeepseekV32MTPModel in the two-hidden-state architecture
set. |
| #54713 |
`b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3)
| Only main takes replay_boundaries in compressed-prefix hit lookup;
preserve release calls without that keyword. |
| #42785 |
`442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18)
| Only main capture callers pass axis_keys; preserve the existing Ascend
rejection of nonempty axes. |
| #52358 |
`8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
`9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4)
| Release already has mamba_has_prefill_checkpoint_blocks; later main
also has fine-grained prefix-cache state. |
| #51251 |
`7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| v0.29.0 and main expose ec_manager_config; retire the old release-only
ScoreEncoder configuration skip. |
| #53240 |
`b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c)
| v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache
groups; retire the old release-only replay skip. |
| #53853 |
`e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
`4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both expose the V1 unsupported-feature helper used by the existing
Ascend MRV1 feature filter. |

| vllm-ascend change | Why it is required | Upstream cause and direct
link | Lane | Verification |
|---|---|---|---|---|
| `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM
v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec,
list[int]]` from `get_mamba_groups` and both initialize `recoverssm`;
keeping the old fallback would preserve an unsupported third contract |
[v0.29.0 mamba
groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704),
[fixed-main mamba
groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704),
[v0.29.0
RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103),
[fixed-main
RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103)
| both, common implementation | Existing constructor UT now asserts the
parent-created RecoverSSM value is retained; Ruff and compileall pass |
| `worker/v2/model_runner.py`: always forward
`kv_cache_allocation_context` | v0.29.0 and fixed main both accept this
keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0
signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538),
[fixed-main
signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566),
[vllm-ascend
both, common implementation | Existing UT continues to assert the exact
context object reaches the parent; Ruff and compileall pass |
| `_310p/worker/v2/model_runner.py`: select the release KV-zeroing
contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes
`KVCacheGroupSpec.is_eagle_group` but lacks
`SpeculativeConfig.use_eagle_block_drop`; fixed main added the method |
vLLM [#53388
diff](https://github.com/vllm-project/vllm/pull/53388/files), commit
[`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a);
[v0.29.0 group
field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200),
[fixed-main
helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876)
| release differs from main | Existing two-path KV-zeroing UT retained
and renamed for v0.29.0; Ruff and compileall pass |
| existing DFlash kernel UT: remove v0.28-only kwarg omission | the
current Ascend kernel accepts the CP arguments and the only excluded
lane was v0.28.0, which this PR replaces | [vllm-ascend
both, common invocation | Existing NPU test remains enabled with all
assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |
|---|---|---|---|
| `patch/platform/patch_parallel_config.py`, registration, and global
patch documentation | Allow Ascend PCP+DP by removing the generic GPU
restriction, following vLLM
[#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit
`7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic
`parallel_config.current_platform` lookup and rebuild ParallelConfig →
SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing
configuration/EPLB UTs and historical PCP+DP NPU execution passed;
current-head CI pending. |
| `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT |
Both supported contracts require `is_drafting`, from
[#54282](https://github.com/vllm-project/vllm/pull/54282/files),
`fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited
optimized kernel and use a common wrapper. | Both | Existing positive
drafting assertion retained; current-head CI pending. |
| `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache
UT | Both support `cache_hit_alignment_tokens`, introduced by
[#53598](https://github.com/vllm-project/vllm/pull/53598/files),
`2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only
write-mask branch. | Both | Existing assertions retained; current-head
CI pending. |

| Change | Exact source evidence | Decision |
|---|---|---|
| `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM
[#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95),
`d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL
before v0.29.0; both exact supported sources lack it. The old
conditional came from vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0
selector and use the existing exclusion for both supported lanes. This
does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy
skip. |
| `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and
`tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM
[#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29),
`12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export
`nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The
v0.28.0 fallback originated in vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete
import-failure/`None` fallbacks and the existing UT's obsolete
availability skip; assertions remain unchanged and execute on both
lanes. |
| `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend
[vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3),
`799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream
revert while resolving the real rebase conflict; do not reintroduce the
reverted graph-update behavior. |
| `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend
[vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2),
`d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged
310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation
branch. |

</details>

</details>

Yes. The supported vLLM release and default container builds move to
v0.29.0, while the fixed main remains supported. The vllm-ascend package
version does not change. Existing 0.29-only PCP+DP compatibility is
restored.

- Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`:
[run
35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654)
succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected
NPU jobs confirm the requested Ascend head, integration base
`2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM
pins/installations: fixed main
`84030bbe3d74d99bad477a3d2e37a973ccd8865c` /
`0.1.dev1+g84030bbe3.empty`, release
`98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane
totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest
invocations. Skips/xfails are not passes. The resulting rebased
integration commit is not printed and is not inferred. Actual release
CPU and failed/cancelled image variants remain gaps.

- Current-head CI update (2026-09-20): [main CPU
UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955)
passed **5157 tests**, with **67 skipped**; actual installed vLLM was
`0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head;
the log does not print the full resulting integration head/base.
Pre-commit and mypy passed. NPU/E2E results are recorded above; actual
release CPU remains unverified.
- Image build is partially blocked by infrastructure: A5 amd64
[Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005)
and
[openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068)
failed before reading the Dockerfile because BuildKit could not create a
snapshot temporary directory (`no space left on device`). No release
compatibility code change is justified by this failure; cancelled
variants remain unverified. The successful [310P openEuler arm64
build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987)
explicitly checked out release
`98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`;
this is build evidence, not runtime UT coverage.

- Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto
the baseline above. Previous-head CPU [job
106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458)
stopped during Ascend integration rebase after vllm-project#16803 merged, before any
UT ran. Its vLLM checkout was the fixed main and installation reported
`0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved;
fresh-head CI was retriggered by the push.

- Remote official v0.29.0 tag resolved to the exact SHA above; both
source pins inspected.
- All 95 changed Python files pass Ruff lint, Ruff format and AST
parsing; `git diff --check` passes. Dockerfile tag/marker consistency
checked across all eight variants.
- Full pre-commit invocation: Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-dependent hooks cannot run on this
Windows host; Python launcher hooks exit 9009. Full CI lint remains
pending.
- Actual main/release CPU/NPU and image builds are not claimed locally:
the required Linux/NPU/container environment is unavailable. New PR E2E
and image-build CI are requested. Existing CPU workflow runs fixed main
only, so actual release CPU remains a validation gap.
- vllm-project#16393's historical green CI is not a substitute for this new head. No
new test functions, workflow changes, golden/threshold changes or
additional skips are introduced beyond restoring vllm-project#16393. Inherited
0.28-only tests/skips are not counted as passes.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…llm-project#16755)

### What this PR does / why we need it?

Enables GLM-Next (GLM-5.3-Flash) prefill/decode disaggregation on Model
Runner V2 with `MooncakeConnectorV2`, for both equal TP (P=TP4 / D=TP4)
and unequal TP (P=TP8 / D=DP2xTP4).

GLM-Next could not run PD on MRV2 because of three structural gaps:

1. **Granularity mismatch.** MRV2 hands metadata builders
kernel-granularity specs (`block_size` == the C128 SFA kernel size),
while the scheduler, the common block tables, and the GLM-Next cache
contract all use the logical block from `cache_config.block_size`.
Sizing persistent buffers from the spec overflows them by the
ceil-rounding slack (`required=1044` vs `capacity=1040` on the first
forward).
2. **GLM-Next layout.** The v2 MLA spec re-wrap dropped the
`model_version` / `compress_ratio` markers the cache-group
classification needs; the compressed indexer small page (compress_ratio
> 1) is a single-tensor overlay on a shared padded slot, not an MLA
latent/rope K/V pair; NoPE MLA layers legitimately carry an empty rope/V
component.
3. **Unequal TP.** Page-size unification is TP-dependent for hybrid
models (the Mamba state page scales with heads/TP), so a TP8 producer
and a TP4 consumer publish different unified block sizes for
byte-identical replicated caches — kernel blocks can no longer be paired
index-wise.

Changes:

- v2 attn_utils: propagate `model_version` /
`indexes_kv_by_block_stride` / `compress_ratio` /
`non_causal_multi_token_decode` through the MLA spec re-wrap; reshape
the compressed indexer small-page layer (compress_ratio > 1) as a single
strided tensor overlaid from byte zero of the shared padded slot.
- KPool indexer/tail backends: accept upstream's `cache_dtype_str` probe
kwarg; take the logical block size from `cache_config.block_size`.
- SFA metadata state: size the persistent block table buffer from the
cache config's logical block expansion instead of the spec's.
- Mooncake V2 connector: skip zero-size cache components when collecting
register ranges (their degenerate storage corrupts the range arithmetic)
and skip zero-length transfer entries the engine rejects with E19999.
- Unequal-TP: transfer replicated MLA/indexer caches as GCD-sized
token-segment byte runs (`_compute_cross_tp_runs`); transfer the indexer
tail whole-block instead of through the HND head-sharded path; raise
with the layer name on kernel-block-size mismatch.

### Does this PR introduce _any_ user-facing change?

Yes — it enables a new deployment topology: GLM-Next hybrid models can
now run PD disaggregation on MRV2 with `MooncakeConnectorV2` (pull
mode), including a TP8 producer feeding a DP2xTP4 consumer. Requires
`VLLM_USE_V2_MODEL_RUNNER=1` and a kv-transfer-config with
`MooncakeConnectorV2`. No behavior change for existing models and
configurations: the new code paths are gated (the strided reshape only
triggers for MLA-family specs with compress_ratio > 1; the cross-TP run
path only activates when local and remote kernel block sizes differ; the
equal-TP block pairing path is unchanged).

### How was this patch tested?

- **Unit tests:** the `mooncake_v2` suite passes 150/150, covering
cross-TP run splitting/coalescing/computed-offset/bounds, tail
replica/whole-block mapping, bucket routing, and per-component address
scaling. The full UT sweep over
`tests/ut/{models,attention,worker,core,kv_offload}` produces a failure
set identical to upstream main (pre-existing failures only, none
introduced by this PR).
- **E2E, 10-layer GLM-5.3-Flash debug model:** 1P1D TP4/TP4 (producer
cards 0-3, consumer cards 4-7, eager), pull-mode KV transfer over the
mooncake engine relayed through a load-balance proxy: first-token
fidelity 18/18, outputs match the single-instance baseline to the
model's own determinism, real per-rank transfer logs, zero recompute
fallback.
- **E2E, full w8a8 GLM-5.3-Flash:** P=TP8 producer, D=DP2xTP4 consumer
(FULL_DECODE_ONLY graph mode): zero pull errors, ~18.5 MB transferred
per request, TextVQA accuracy consistent with the colocated baseline.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: sunbaosong <13793883820@163.com>
Co-authored-by: bubaishenhua112-netizen <bubaishenhua112@gmail.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…vllm-project#17004)

### What this PR does / why we need it?



Reintroduce the vLLM v0.29.0 release upgrade from vllm-project#16393, reverted by
vllm-project#16949, with container defaults aligned to the supported release. Based
on upstream main `d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; preserves
later merged changes.


The image/source mismatch is confirmed in [nightly job
106004497020](https://github.com/vllm-project/vllm-ascend/actions/runs/35482881089/job/106004497020):
the image installed `vllm 0.28.0+empty` (tag `v0.28.0`) while upgraded
Ascend failed importing `_get_packed_kv_cache_groups`. All eight root
Dockerfiles still defaulted to v0.28.0. The reusable image workflow
passes no `VLLM_TAG` override, so those defaults govern release builds.
Changing only the release marker does not update those images.


- Fixed vLLM main remains `84030bbe3d74d99bad477a3d2e37a973ccd8865c`.

- Release changes from v0.28.0
(`2cf0a6915ce544dc493a0990f2ea38d81601128a`) to official v0.29.0
(`98dff2a81d747d1dba01a47f939f48c3526d4206`, rechecked against the
remote tag).
- Restore vllm-project#16393's source compatibility and existing UT changes by
reversing vllm-project#16949, then reconcile current upstream changes. No main
interface scan is rerun for this release-only upgrade.
- Update `VLLM_TAG` in all eight Dockerfiles: default/A3/A5/310P,
Ubuntu/openEuler. Preserve existing exact-commit build overrides. No
workflow changes.


| Current-base adaptation | Exact cause and evidence | Lane | Validation
|
| --- | --- | --- | --- |

| Eight Dockerfile `VLLM_TAG` defaults and release marker | vllm-project#16393
changed the release contract but omitted image defaults; the linked
nightly log proves installation of 0.28.0 and missing
`_get_packed_kv_cache_groups`. | Release images | All eight defaults
match marker; image-build CI requested, pending |
| `worker/v2/attn_utils.py::get_kv_cache_spec` and existing
`_make_mla_layer` UT fixture | Ascend
[vllm-project#16755](https://github.com/vllm-project/vllm-ascend/pull/16755/files),
`df3e755e98fba8c6a18f200c645e0c2050469bc3`, added an old-release field
selector. vLLM
[#51718](https://github.com/vllm-project/vllm/pull/51718/files),
`8bdc70ec7b379279ec0152343239c2d50aced687`, replaced `compress_ratio`
with `tokens_per_state`; both exact supported pins use the latter. Use
the common field, retaining metadata and cache-view assertions. | Both |
Source inspection and static checks passed; actual CI pending |
| `attention/attention_v1.py` import conflict | Preserve
`attention_transfer_window` from Ascend
[vllm-project#15636](https://github.com/vllm-project/vllm-ascend/pull/15636/files),
`5c80630f28f8529aa82716e58b981a78819ec429`, and graph changes from
vllm-project#16798; keep the common relocated PCP import from vllm-project#16393. Do not restore
the old compute-start import or unused weak-reference import. | Both |
Conflict resolved; static checks passed |
| `worker/v2/aclgraph_utils.py` import conflict | Preserve
`ContextSource`, `UpdatableGraph`, and `use_updatable_graph` from Ascend
[vllm-project#16798](https://github.com/vllm-project/vllm-ascend/pull/16798/files),
`34bb51f93724c565362f5108f5226303e1b56cad`, while removing the obsolete
0.28 version selector restored by the revert. | Both | Conflict
resolved; static checks passed |


| Existing 310P and Mamba model-runner UT imports | Preserve
hardware-profile imports and mocks from Ascend
[vllm-project#16803](https://github.com/vllm-project/vllm-ascend/pull/16803/files),
`d3f6b4b59ab55dc42af1f30f425c893b95af5ee2`; remove only the unused
old-release selector import. Hardware capability routing remains
unchanged. | Both | Real rebase conflicts resolved; AST/Ruff/format
passed; fresh CI pending |

Newly merged changes were reviewed for version-contract impact:
vllm-project#16775/vllm-project#16834 PCP metadata and routing, vllm-project#16923 A3 SFA,
vllm-project#16426/vllm-project#16924/vllm-project#16955 custom ops, vllm-project#16081 MTP/SP, vllm-project#16669 xlite,
vllm-project#15636/vllm-project#16747 transfer, vllm-project#16913 A5 pages, vllm-project#16798 graph updates, vllm-project#16755
GLM PD, vllm-project#16952 C8 config, vllm-project#16673 operator removal and vllm-project#16320 Kimi-K3 KV
pool. Preserve these changes; no additional source-proven version branch
was identified beyond the entries above. In particular, vllm-project#16747's
`UniformTypeKVCacheSpecs.kv_cache_specs` and per-layer `block_size`
exist in both exact pins; vllm-project#16320 adds Ascend connector hooks. CI remains
necessary to validate runtime interactions. CI/documentation-only PRs
are retained unchanged.


<details>

<summary>Inherited per-file compatibility evidence from vllm-project#16393</summary>


The following source-contract ledger is inherited from vllm-project#16393. Any
historical verification wording refers only to that earlier PR; it does
not certify this new head. New-head validation is listed below.


| Reference | Commit |





|---|---|





| vllm-ascend rebase base | `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c` |
| PR head | `8c5d785931c495701bf1da8b5bc80b321ebd3cd0` |





| Fixed vLLM main | `84030bbe3d74d99bad477a3d2e37a973ccd8865c` |





| Previous release: v0.28.0 | `2cf0a6915ce544dc493a0990f2ea38d81601128a`
|
| Target release: v0.29.0 | `98dff2a81d747d1dba01a47f939f48c3526d4206` |






- Use common implementations where v0.29.0 and the fixed main share
KV-cache layouts, Mamba copy/group APIs, PCP handling and
speculative-decoding contracts.
- Retain explicit `vllm_version_is("0.29.0")` branches for contracts
that still differ, including RoPE, scheduler block snapshots,
InputBatch, ReplaySSM, KV zeroing and DSpark PP handling.
- Remove obsolete v0.28.0 compatibility and adapt existing test
fixtures. Version detection uses package versions and the explicit
`VLLM_VERSION` override, with local-version suffix handling and UT
environment isolation retained; no hard-coded release-SHA inference
remains.
- Apply the Ascend PCP+DP validator patch only to v0.29.0, preserving
other validations and EPLB platform binding. Rebuild dependent Pydantic
schemas so nested configuration validation uses the patched validator.
The global patch documentation records its rationale and removal
criteria.






- Preserve vllm-project#15747's Spec+PP protocol/partition handling after rebase;
use the common exact-release selector and the real function-local DSpark
sharing import.






This release-only upgrade does not require a new main old-to-new
interface scan. The latest rebase incorporates
[vllm-project#16905](https://github.com/vllm-project/vllm-ascend/pull/16905/files),
which reverted vllm-project#16544 at `9dc6704559ebe2809b5cb7b7fc44184bd1338b3c`. The
now-unnecessary V4.1 drafter import gate and tuple annotation have been
removed. Other release adaptations, including vllm-project#15747 Spec+PP handling,
remain. Detailed contract evidence is retained below for review.






<details>





<summary>Per-file adaptations and exact upstream evidence</summary>






#### Per-file adaptation ledger











Evidence IDs refer to the exact source contract and upstream diff table
below. Every row is syntax checked; branch-normalized AST comparison
confirms unchanged main function bodies except the version identity
helper and the explicitly retired propose argument/type annotations.
Both supported versions completed CPU and NPU execution as recorded
below.






| Ascend file / symbols | Disposition and upstream evidence |
Verification |
|---|---|---|





| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` /
`_load_dspark_model_with_target_quant`;
`tests/ut/patch/worker/test_patch_dspark_pp.py` | Preserve rebased
[Ascend
vllm-project#15747](https://github.com/vllm-project/vllm-ascend/pull/15747/files#diff-a170a42fe1e9275c999642b05c4a437d0c9104dd8ea19151796142909b00ab2a)
(`82df9d871`), including manual PP partition masking. vLLM [#52809
diff](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`91a893de64722019ea2faf852e06cabe143b3490`) moved `_should_share`
inside the loader on both supported pins, so retain the earlier release
fix: patch `eagle_utils`, never read/patch a nonexistent
`dspark_utils._should_share`. v0.29 keeps its global PP guard; main has
native PP via
[#50514](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
(`d87a440f88e28e5b37f9b1e22ce214d0426d5352`). | Existing UT now checks
the real function-local import, absent module alias, and restoration on
success/failure for both lanes; partition and PP assertions retained.
Ruff/syntax pass; actual CPU/NPU pending. |
| `vllm_ascend/worker/v2/pp_utils.py` / `use_legacy_spec_pp`;
`tests/ut/worker/v2/test_pp_utils.py` | vllm-project#15747 added broad 0.28/0.29
routing and an obsolete 0.28 dev-build recognition path. For the two
supported pins, #50514 exists only on fixed main. Route through
`vllm_version_is("0.29.0")`; retain package/local-suffix and explicit
environment-override semantics. No release-SHA inference or third
release lane. | Existing UTs exercise the real uncached version helper
with monkeypatch isolation, local suffix, fixed-main dev string,
explicit override, and non-target versions. Isolated routing checks and
Ruff/syntax pass; full CPU UT pending. |
| `vllm_ascend/_310p/model_runner_310p.py`<br>`_prepare_inputs`,
`_allocate_kv_cache_tensors` | #51718, #56078; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py`<br>`set_inputs_first_pass`
| #56078; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/_310p/worker/v2/model_runner.py`<br>`initialize_kv_cache`,
`_allocate_kv_cache_tensors`, `_prepare_inputs_310p` | #51718/#54436;
use shared standardized layouts and retain the v0.29.0 InputBatch gate.
The rebase preserves vllm-ascend vllm-project#16043's MTP copy tracking while
removing only legacy v0.28.0 allocation paths. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/_310p/worker/v2/rope.py`<br>`get_310p_rope_state` |
#56078; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/attention_v1.py`<br>`module imports/dispatch` |
#52839; both lanes use the common PCP import. Rebase keeps current
upstream graph code and removes only the v0.28.0 import branch. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/context_parallel/sfa_cp.py`<br>`module
imports/dispatch` | #52839; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/attention/indexer.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/attention/mla_v1.py`<br>`module imports/dispatch` |
#52839; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/core/dyntra_lb_scheduler.py`<br>`module imports/dispatch`
| #51358, #54853; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/core/recompute_scheduler.py`<br>`module
imports/dispatch`, `schedule` | #51358, #54853; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/core/scheduler_profiling_chunk.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/core/kv_cache_interface.py`<br>`get_kv_cache_compression_ratio`,
`AscendMLAAttentionSpec`, `get_storage_block_size` | #51718, #53906;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/layerwise_cache_layout.py`<br>`_merge_specs`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/preempt_offload/manager.py`<br>`_derive_cpu_config`
| #51718; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`<br>`create_worker`
| #52615; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/models/deepseek_mtp.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/deepseek_v4/indexer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/cache_config.py`<br>`make_tensor` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/glm5next/kv_cache.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/models/kimi_k3_dspark.py`<br>`load_weights` | #53106;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/models/layer/attention/layer.py`<br>`get_kv_cache_spec` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_balance_schedule.py`<br>`module
imports/dispatch` | #51358, #54853; common contracts merged, differing
release contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/platform/patch_kv_cache_coordinator.py`<br>`__init__`
| #53614; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/patch/platform/patch_kv_cache_utils.py`<br>`module
imports/dispatch`, `_get_max_layers_per_page_size`,
`_ascend_max_memory_usage_bytes_from_groups`,
`_ascend_get_kv_cache_config_from_groups` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py`<br>`module
imports/dispatch`, `_patched_get_unsupported_features` | #53853, #53183;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py`<br>`bind_kv_cache` |
#51718; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py::bind_kv_cache` |
v0.29.0 has no ReplaySSM ring-tracker helper; fixed main requires it
after KV binding. Evidence: [#52506,
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`, worker-utils
diff](https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c).
Keep an explicit v0.29.0 exclusion and preserve the fixed-main call. |
Failure reproduced on v0.29.0; Historical release/main NPU validation
passed; current-head CI pending |
| `vllm_ascend/patch/worker/patch_mamba_utils.py`<br>`module
imports/dispatch`, `_get_state_copy_funcs_for_layer` | #53896; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py`<br>`module
imports/dispatch` | #51718; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |
|
`vllm_ascend/patch/worker/patch_v2/patch_dspark.py`<br>`_load_dspark_model_with_target_quant`
| #50514, #52809; v0.29.0 keeps `get_pp_group` module-bound but imports
`_should_share` locally from Eagle utilities; fixed main removes the PP
guard. Keep the release `get_pp_group` patch, share through the common
Eagle utility, and delete the obsolete v0.28.0
`dspark_utils._should_share` patch. | Exact release failure reproduced;
Historical release/main CPU/NPU validation passed; current-head CI
pending |
|
`vllm_ascend/spec_decode/llm_base_proposer.py`<br>`model_returns_tuple`,
`__init__`, `set_inputs_first_pass` | #56078, #52861; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/utils.py`<br>`get_kv_cache_tensor_layers`,
`register_ascend_customop`, `vllm_version_is` | #52839, #51718, #52494;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
|
`vllm_ascend/worker/model_runner_v1.py`<br>`_get_ascend_mamba_state_copy_funcs`,
`_allocate_kv_cache_tensors`, `_prepare_inputs`, `_dummy_run`,
`_reshape_kv_cache_tensors`, `get_kv_cache_spec`, NaN helpers |
#51718/#53896/#56078 plus #50323; remove v0.28.0 layouts/copy APIs and
import the NaN helpers directly because v0.29.0 and fixed main expose
the same contract. XD-RoPE remains explicitly gated to v0.29.0. | Static
checked; historical dual-version CPU/NPU validation passed; current-head
CI pending |
| `vllm_ascend/worker/v2/aclgraph_utils.py`<br>`capture` | #53869 makes
`pcp_manager` common to both lanes. The rebase preserves vllm-ascend
vllm-project#16409's host-parameter-update revert and removes only the obsolete
v0.28.0 capture branch. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/attn_utils.py`<br>`get_kv_cache_spec`,
`_allocate_kv_cache`, `_reshape_kv_cache_v2` | #51718; common contracts
merged, differing release contracts explicitly gated as detailed below.
| Static checked; prior-head dual-version results below; rebased CI
pending |
| `vllm_ascend/worker/v2/block_table.py`<br>`__init__`,
`init_block_table_layout_tensors`, `compute_slot_mappings` | #51718,
#51031; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
| `vllm_ascend/worker/v2/model_runner.py`<br>`module imports/dispatch`,
`prepare_inputs`, `__init__`, `sample_tokens`, `prepare_dummy_attn`,
`execute_model` | #50514, #54436, #52506, #55212, #53515; common
contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/pcp_manager.py`<br>`partition_batch` | #53515;
common contracts merged, differing release contracts explicitly gated as
detailed below. | Static checked; prior-head dual-version results below;
rebased CI pending |
| `vllm_ascend/worker/v2/sample/gumbel.py`<br>`module imports/dispatch`
| #54282; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/__init__.py`<br>`init_speculator` |
#49811; common contracts merged, differing release contracts explicitly
gated as detailed below. | Static checked; prior-head dual-version
results below; rebased CI pending |
|
`vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`module
imports/dispatch`, `propose` | #53694, #52188; common contracts merged,
differing release contracts explicitly gated as detailed below. | Static
checked; prior-head dual-version results below; rebased CI pending |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`<br>`propose`
| #53694; common contracts merged, differing release contracts
explicitly gated as detailed below. | Static checked; prior-head
dual-version results below; rebased CI pending |
|
`vllm_ascend/worker/worker.py`<br>`_scale_kv_cache_memory_for_multi_group`,
`wake_up` | #51718, #53508; common contracts merged, differing release
contracts explicitly gated as detailed below. | Static checked;
prior-head dual-version results below; rebased CI pending |






#### Exact upstream evidence











| Upstream change | Full commit SHA / direct diff | Actual supported
contracts and branch decision |
|---|---|---|





| #51718 |
`8bdc70ec7b379279ec0152343239c2d50aced687`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Both use layers/layer_stride/block_stride/offset, tokens_per_state,
CircularBufferSpec and standardized backing; retire
shared_by/compress_ratio allocation branches. |
| #52839 |
`58e5ee0158b6a264c3506f00480e108a34b33ee3`<br>[vllm/v1/attention/ops/pcp.py](https://github.com/vllm-project/vllm/pull/52839/files#diff-23d7e7ab01a7f44414796e1e1b08cfa5277a05f54143ee2dd3a1041ca47772a8)
| Both import PCP operations from vllm.v1.attention.ops.pcp. |
| #53896 |
`e126687a9a828d513c01a07cd69f025f27d63280`<br>[vllm/v1/worker/mamba_utils.py](https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec)
| Both use Mamba copy-function dictionaries and unwrap
UniformTypeKVCacheSpecs. |
| #53106 |
`1fe3a1571ac67581478a11743e55a306de1d136f`<br>[vllm/model_executor/models/utils.py](https://github.com/vllm-project/vllm/pull/53106/files#diff-b0ba1095e9881e5c87e33dfd20958d1e1ceafe8a4433aa692f468e61be130b21)
| Both use WeightsMapper instead of AutoWeightsLoader
skip_prefixes/skip_substrs. |
| #53906 |
`98ed0856f31fa3aaf5e27464e2b4ef5a8ee6b2f5`<br>[vllm/v1/kv_cache_interface.py](https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb)
| Only pinned main has the optional MLA storage_block_size dataclass
field; release keeps the Ascend derived property, using
tokens_per_state. |
| #56078 |
`719284fe158f1be8a9dd92953295fc9d49015730`<br>[vllm/config/model.py](https://github.com/vllm-project/vllm/pull/56078/files#diff-998c640befaf137b9af825f29f4e6e47d273caab1fd04093c97df24b18f5c417)
| Release retains uses_xdrope_dim and three M-RoPE dimensions; pinned
main uses mrope_num_dims and unified RoPE. |
| #52615 |
`138d137b5b955a5ebee98dfc946ecdd65d7b87ce`<br>[vllm/v1/kv_offload/cpu/spec.py](https://github.com/vllm-project/vllm/pull/52615/files#diff-7d073c26c9b22f74e9ff9e4c8da733d1807233edb83439fc64fbd48c43ae6045)
| Release uses num_blocks/kv_bytes_per_block; main uses
num_chunks/kv_bytes_per_chunk. |
| #51358 |
`6b110badbb22d3f66c7218b71138f13b7a6b3419`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/51358/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release now has boundary_state_offloads and KVConnectorBlockState;
remove partial_tail_offloads plumbing. |
| #54853 |
`0b066293f3c738a0cbd3a087bf893f2f4dcd61f2`<br>[vllm/v1/core/sched/output.py](https://github.com/vllm-project/vllm/pull/54853/files#diff-cafd89ce8a698a56acb24ada62831cbc7a980782f78a52d1742ba238031f296c)
| Release constructor takes block_ids snapshots; main takes req_ids and
resolve_block_ids. Keep exact release snapshot membership and main
lazy-resolution membership. |
| #53614 |
`144e79c8106da23141ac010394b782f730cc7fe8`<br>[vllm/v1/core/kv_cache_coordinator.py](https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725)
| Only main configures drop_eagle_checkpoint_block for replay-aligned
Mamba checkpoints. |
| #50514 |
`d87a440f88e28e5b37f9b1e22ce214d0426d5352`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Release retains module-level PP/share symbols and Ascend PP
workaround; main has the subsequent PP integration. |
| #52809 |
`91a893de64722019ea2faf852e06cabe143b3490`<br>[vllm/v1/worker/gpu/spec_decode/dspark/utils.py](https://github.com/vllm-project/vllm/pull/52809/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393)
| Between v0.28.0 and v0.29.0, `_should_share` moved from a DSpark
module binding to a function-local import from Eagle utilities. The
shared Eagle patch remains effective; the old DSpark-module read/write
must be removed. |
| #54436 |
`6bafc049aae6c26e210162630914ee9177a4b586`<br>[vllm/v1/worker/gpu/input_batch.py](https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902)
| Release InputBatch requires max_seq_len_np; main removed it. |
| #52506 |
`adebc41b7e9f1085d3f73434e23beb76883b9eb4`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Only main accepts valid_dummy_state_slots/valid_state_slots capture
arguments. |
| #55212 |
`83990f5bcc0b5eb08b2f1fd1b109fa8a31b122f3`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/55212/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Release prepares DCP local sequence lengths before partitioning; main
initializes DCP metadata afterwards. |
| #53515 |
`b1fbbc2ade51e3826bc92e4733c9c692ee21d42d`<br>[vllm/v1/worker/gpu/pcp_manager.py](https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127)
| Both accept padded_num_tokens for persistent PCP input buffers. |
| #53869 |
`b3af042abd8fe5a297ec3ec72db276fd661a67b3`<br>[vllm/v1/worker/gpu/cudagraph_utils.py](https://github.com/vllm-project/vllm/pull/53869/files#diff-fc699ff1c69fd17adb60116b3a7a20aa88b5187205f589952d0f3792475ec212)
| Both accept pcp_manager during graph capture. |
| #51031 |
`0ecc284790e5403f74b899524ef82ecb69f83cb3`<br>[vllm/v1/worker/gpu/block_table.py](https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1)
| Both distinguish KV and kernel block sizes during DCP slot mapping. |
| #54282 |
`fe755c88995ad468882517b6c4bdd60138d46a3a`<br>[vllm/v1/worker/gpu/sample/gumbel.py](https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8)
| Both gumbel sampling APIs include is_drafting. |
| #52188 |
`d1e3eee6fb8ed3623241ef5c8e3ac533f775bff9`<br>[vllm/v1/worker/gpu/spec_decode/dflash/speculator.py](https://github.com/vllm-project/vllm/pull/52188/files#diff-0220b514682125dea26026db2d7caa1a0d9c772d84452e9e124434d1e11be5f0)
| Both DFlash input kernels take cp_rank/CP_SIZE/CP_INTERLEAVE. |
| #53694 |
`5acc1c4e4b8730298cbff7a4a7c68c814dc24fd7`<br>[vllm/v1/worker/gpu/spec_decode/autoregressive/speculator.py](https://github.com/vllm-project/vllm/pull/53694/files#diff-3e7e2ba21d64e309503b7c5fe537363a7669fa1baf84d3c2e4891d9aed64fbe6)
| Both propose APIs take DPSyncState; remove the obsolete token-count
argument and retain replicated-PCP synchronization. |
| #49811 |
`01af92e175407231b1433b0aef01a1b9c983d955`<br>[vllm/v1/worker/gpu/spec_decode/extract_hidden_states.py](https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e)
| Both support extract_hidden_states on MRV2; remove old unsupported
dispatch/skip. |
| #53508 |
`479eeb32d2b432dbb4e442fbd3f94ca2eca35d67`<br>[vllm/v1/worker/gpu_model_runner.py](https://github.com/vllm-project/vllm/pull/53508/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29)
| Both remove post_kv_cache_wake_up; retire release-only call. |
| #52494 |
`3ff4f02dfe69abc1a0375d1ea8d8d5cb25609fcc`<br>[vllm/models/kimi_k3/amd/mla.py](https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1)
| Only main provides KimiK3MultiHeadLatentAttentionWrapper; retain
release exclusion. |
| #52861 |
`b09bd69b5bf14911abef9a0e8e493b83c8a38fa6`<br>[vllm/v1/spec_decode/llm_base_proposer.py](https://github.com/vllm-project/vllm/pull/52861/files#diff-56fcad87dae9192cb0ab1643473765096cc44e1e744dccf3b99c4f2d85360d28)
| Both include DeepseekV32MTPModel in the two-hidden-state architecture
set. |
| #54713 |
`b28c3e1568bfae930f61d4b24940e47528c85d4a`<br>[vllm/v1/core/single_type_kv_cache_manager.py](https://github.com/vllm-project/vllm/pull/54713/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3)
| Only main takes replay_boundaries in compressed-prefix hit lookup;
preserve release calls without that keyword. |
| #42785 |
`442d36031ce710aa2a777c353becb660c57ab2bd`<br>[vllm/v1/worker/encoder_cudagraph.py](https://github.com/vllm-project/vllm/pull/42785/files#diff-287435cc78f753ae42acda542bca01c094557cd17db6bc2991d968fd012c1f18)
| Only main capture callers pass axis_keys; preserve the existing Ascend
rejection of nonempty axes. |
| #52358 |
`8f816a3f665489d7f0d222115d4f72ebab01076b`<br>[vllm/v1/worker/gpu/model_runner.py](https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0)
| Both ExecuteModelState have dp_sync; only main has cudagraph_stats. |
| #52789 |
`9eb9d9d3953959695108600c8ed33d36bc6a1e5f`<br>[vllm/v1/core/sched/scheduler.py](https://github.com/vllm-project/vllm/pull/52789/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4)
| Release already has mamba_has_prefill_checkpoint_blocks; later main
also has fine-grained prefix-cache state. |
| #51251 |
`7bbbf7c8e5040f7ebd374e8cbc657e01af1136dd`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/51251/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| v0.29.0 and main expose ec_manager_config; retire the old release-only
ScoreEncoder configuration skip. |
| #53240 |
`b2db227a7c4c5e55f85524a094b607e5d27408b4`<br>[vllm/model_executor/layers/fused_moe/routed_experts_capturer.py](https://github.com/vllm-project/vllm/pull/53240/files#diff-0bbfc19dd02619bbcc948fa0281d4ab9a5b77b7d5062cde9911319c93089535c)
| v0.29.0 supports MRV2 routed-expert capture and unwraps uniform cache
groups; retire the old release-only replay skip. |
| #53853 |
`e376d45e82cb7e220da430e3179e81eb0922cf56`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53853/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both delegate PCP compatibility validation to the PCP manager. |
| #53183 |
`4aab2b0ebed20343efe543c633f71b3c1336d5b8`<br>[vllm/config/vllm.py](https://github.com/vllm-project/vllm/pull/53183/files#diff-bee6813076031d3ca1edc903c1b02b81e4676519afc562ce3fefe37f20c7b650)
| Both expose the V1 unsupported-feature helper used by the existing
Ascend MRV1 feature filter. |






#### Additional inherited contracts











| vllm-ascend change | Why it is required | Upstream cause and direct
link | Lane | Verification |
|---|---|---|---|---|





| `_310p/worker/v2/model_state.py`: remove tuple-return and RecoverSSM
v0.28 fallbacks | v0.29.0 and fixed main both return `dict[MambaSpec,
list[int]]` from `get_mamba_groups` and both initialize `recoverssm`;
keeping the old fallback would preserve an unsupported third contract |
[v0.29.0 mamba
groups](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/mamba_utils.py#L691-L704),
[fixed-main mamba
groups](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/mamba_utils.py#L691-L704),
[v0.29.0
RecoverSSM](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103),
[fixed-main
RecoverSSM](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_states/mamba_hybrid.py#L80-L103)
| both, common implementation | Existing constructor UT now asserts the
parent-created RecoverSSM value is retained; Ruff and compileall pass |
| `worker/v2/model_runner.py`: always forward
`kv_cache_allocation_context` | v0.29.0 and fixed main both accept this
keyword; the new-base gate was specifically for v0.28.0 | [v0.29.0
signature](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/worker/gpu/model_runner.py#L533-L538),
[fixed-main
signature](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/v1/worker/gpu/model_runner.py#L561-L566),
[vllm-ascend
vllm-project#16791](https://github.com/vllm-project/vllm-ascend/pull/16791/files) |
both, common implementation | Existing UT continues to assert the exact
context object reaches the parent; Ruff and compileall pass |
| `_310p/worker/v2/model_runner.py`: select the release KV-zeroing
contract with `vllm_version_is("0.29.0")` | v0.29.0 exposes
`KVCacheGroupSpec.is_eagle_group` but lacks
`SpeculativeConfig.use_eagle_block_drop`; fixed main added the method |
vLLM [#53388
diff](https://github.com/vllm-project/vllm/pull/53388/files), commit
[`481839ad9e5ebf87aecb54fa5c9d986bd5ea4b81`](vllm-project/vllm@481839a);
[v0.29.0 group
field](https://github.com/vllm-project/vllm/blob/98dff2a81d747d1dba01a47f939f48c3526d4206/vllm/v1/kv_cache_interface.py#L1200),
[fixed-main
helper](https://github.com/vllm-project/vllm/blob/84030bbe3d74d99bad477a3d2e37a973ccd8865c/vllm/config/speculative.py#L1874-L1876)
| release differs from main | Existing two-path KV-zeroing UT retained
and renamed for v0.29.0; Ruff and compileall pass |
| existing DFlash kernel UT: remove v0.28-only kwarg omission | the
current Ascend kernel accepts the CP arguments and the only excluded
lane was v0.28.0, which this PR replaces | [vllm-ascend
vllm-project#15098](https://github.com/vllm-project/vllm-ascend/pull/15098/files) |
both, common invocation | Existing NPU test remains enabled with all
assertions; Current-head CI pending |
| Ascend change | Why / upstream cause | Lane | Verification |





|---|---|---|---|





| `patch/platform/patch_parallel_config.py`, registration, and global
patch documentation | Allow Ascend PCP+DP by removing the generic GPU
restriction, following vLLM
[#54523](https://github.com/vllm-project/vllm/pull/54523/files), commit
`7c2f1ff4958eaf0818405e9192c71608fe4a16b1`. Preserve dynamic
`parallel_config.current_platform` lookup and rebuild ParallelConfig →
SpeculativeConfig → VllmConfig. | v0.29.0 only | Existing
configuration/EPLB UTs and historical PCP+DP NPU execution passed;
current-head CI pending. |
| `ops/triton/v2/sample/categorical_sample.py`, existing Gumbel UT |
Both supported contracts require `is_drafting`, from
[#54282](https://github.com/vllm-project/vllm/pull/54282/files),
`fe755c88995ad468882517b6c4bdd60138d46a3a`; preserve the inherited
optimized kernel and use a common wrapper. | Both | Existing positive
drafting assertion retained; current-head CI pending. |
| `patch/platform/patch_kv_cache_coordinator.py`, existing prefix-cache
UT | Both support `cache_hit_alignment_tokens`, introduced by
[#53598](https://github.com/vllm-project/vllm/pull/53598/files),
`2ba984a5d06db414f3b2474fe9338faf6cd80a1c`; retire the v0.28-only
write-mask branch. | Both | Existing assertions retained; current-head
CI pending. |






#### Rebase and retired-fallback evidence











| Change | Exact source evidence | Decision |





|---|---|---|





| `tests/e2e/conftest.py::PROMPT_CONFIGS[HunyuanOCR]` | vLLM
[#53272](https://github.com/vllm-project/vllm/pull/53272/files#diff-0852f1e9753819abe4f85380abdc0660fd4a998f44bd7bb73d68f462bd776d95),
`d53b1c2efc0ca7161c3844f51ca9acbcdb7129d5` removes native Hunyuan V1/VL
before v0.29.0; both exact supported sources lack it. The old
conditional came from vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-824bc3e5c28226a3c0f4581750fa93a7da7001aff9e0ff4aee4131e3edd7b144),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Remove the v0.28.0
selector and use the existing exclusion for both supported lanes. This
does not migrate the protected SFA PCP skip or add a v0.29.0 accuracy
skip. |
| `vllm_ascend/worker/model_runner_v1.py` NaN helper imports and
`tests/ut/worker/test_model_runner_v1_nan_detection.py` | vLLM
[#50323](https://github.com/vllm-project/vllm/pull/50323/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29),
`12292d94b25869be2af6b6d4f8eea6c2445e935f`; both exact sources export
`nans_to_dict`, `gpu_sync_allowed`, and `raise_if_nan_logits`. The
v0.28.0 fallback originated in vllm-ascend
[vllm-project#14898](https://github.com/vllm-project/vllm-ascend/pull/14898/files#diff-c49594855b615477bbc34f06d2d423a7dd84c021a7925cd1f61fdb79cb814c08),
`e5118d151314ae18e56c0b63aa8dd00d294adc22`. | Delete
import-failure/`None` fallbacks and the existing UT's obsolete
availability skip; assertions remain unchanged and execute on both
lanes. |
| `vllm_ascend/worker/v2/aclgraph_utils.py` | vllm-ascend
[vllm-project#16409](https://github.com/vllm-project/vllm-ascend/pull/16409/files#diff-dc1c8a80dadf23627467bf04d6bc59fe13c6d075b7d2564d0acc07dc97d921f3),
`799801feef347469d5e9b39374b9210e0d4f7431`. | Preserve the upstream
revert while resolving the real rebase conflict; do not reintroduce the
reverted graph-update behavior. |
| `vllm_ascend/_310p/worker/v2/model_runner.py` | vllm-ascend
[vllm-project#16043](https://github.com/vllm-project/vllm-ascend/pull/16043/files#diff-2f27611a7cfbbdb780b9e88157f05d14d6dfe65406c8d66a5ba1f6949124f5d2),
`d4d2957e5208c2f464d4625c05920bd29ea233cb`. | Preserve the newly merged
310P MTP copy tracking and remove only the v0.28.0 descriptor/allocation
branch. |






</details>













</details>



### Does this PR introduce _any_ user-facing change?



Yes. The supported vLLM release and default container builds move to
v0.29.0, while the fixed main remains supported. The vllm-ascend package
version does not change. Existing 0.29-only PCP+DP compatibility is
restored.


### How was this patch tested?

- Final E2E result for head `a3b76b073a201851454e873aa89fc2254992fc06`:
[run
35504646654](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654)
succeeded (39 successful jobs, 6 skipped). Raw logs from all 32 selected
NPU jobs confirm the requested Ascend head, integration base
`2de71b594319bde52c8bded69eba154f50be8e75`, and the actual vLLM
pins/installations: fixed main
`84030bbe3d74d99bad477a3d2e37a973ccd8865c` /
`0.1.dev1+g84030bbe3.empty`, release
`98dff2a81d747d1dba01a47f939f48c3526d4206` / `0.29.0+empty`. Each lane
totals **563 passed, 35 skipped, 1 xfailed** across its selected pytest
invocations. Skips/xfails are not passes. The resulting rebased
integration commit is not printed and is not inferred. Actual release
CPU and failed/cancelled image variants remain gaps.


- Current-head CI update (2026-09-20): [main CPU
UT](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646654/job/106062951955)
passed **5157 tests**, with **67 skipped**; actual installed vLLM was
`0.1.dev1+g84030bbe3.empty`. Ascend checkout was the current PR head;
the log does not print the full resulting integration head/base.
Pre-commit and mypy passed. NPU/E2E results are recorded above; actual
release CPU remains unverified.
- Image build is partially blocked by infrastructure: A5 amd64
[Ubuntu](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362005)
and
[openEuler](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062362068)
failed before reading the Dockerfile because BuildKit could not create a
snapshot temporary directory (`no space left on device`). No release
compatibility code change is justified by this failure; cancelled
variants remain unverified. The successful [310P openEuler arm64
build](https://github.com/vllm-project/vllm-ascend/actions/runs/35504646536/job/106062361987)
explicitly checked out release
`98dff2a81d747d1dba01a47f939f48c3526d4206` and installed `0.29.0+empty`;
this is build evidence, not runtime UT coverage.


- Current head: `a3b76b073a201851454e873aa89fc2254992fc06`, rebased onto
the baseline above. Previous-head CPU [job
106060931458](https://github.com/vllm-project/vllm-ascend/actions/runs/35503843936/job/106060931458)
stopped during Ascend integration rebase after vllm-project#16803 merged, before any
UT ran. Its vLLM checkout was the fixed main and installation reported
`0.1.dev1+g84030bbe3.empty`. The two import conflicts are resolved;
fresh-head CI was retriggered by the push.


- Remote official v0.29.0 tag resolved to the exact SHA above; both
source pins inspected.
- All 95 changed Python files pass Ruff lint, Ruff format and AST
parsing; `git diff --check` passes. Dockerfile tag/marker consistency
checked across all eight variants.
- Full pre-commit invocation: Ruff, codespell, typos, clang-format,
markdownlint, actionlint, package-init, forbidden-import and
boolean-context checks passed. Bash-dependent hooks cannot run on this
Windows host; Python launcher hooks exit 9009. Full CI lint remains
pending.
- Actual main/release CPU/NPU and image builds are not claimed locally:
the required Linux/NPU/container environment is unavailable. New PR E2E
and image-build CI are requested. Existing CPU workflow runs fixed main
only, so actual release CPU remains a validation gap.
- vllm-project#16393's historical green CI is not a substitute for this new head. No
new test functions, workflow changes, golden/threshold changes or
additional skips are introduced beyond restoring vllm-project#16393. Inherited
0.28-only tests/skips are not counted as passes.

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…llm-project#16910)

### What this PR does / why we need it?

Brings GLM-5.3-Flash prefetch/decode disaggregation to the default model
runner (MRV1) with the MooncakeConnectorV2 pull engine, reusing the
kernel-block layout prefract landed for MRV2 (vllm-project#16755). Without this, PD
disaggregation only works with VLLM_USE_V2_MODEL_RUNNER=1; the default
runner rejects the layout at KV init and the cross-TP handshake fails.

Three adaptation points:
- V1 reshape: two small-slot view branches mirrored from
_reshape_kv_cache_v2 – the per-request kpool tail ring packs a
contiguous suffix of the shared small slot, and the compressed indexer
packs natural kernel blocks as a contiguous prefix, each bounded to half
the slot.
- GLM-Next cache layout: size the shared small slot by the unified page.
_align_glm5_next_cache_specs derived the small page from the small
candidates only, so with MRV1's unpadded worker specs the slot collapsed
to the indexer's own page and the half-slot bound rejected the layout at
KV init. MRV2 workers pre-pad their specs, so their behavior is
unchanged.
- V1 reshape: expose the NOPE main as natural 128-row kernel blocks
instead of one page-sized block per scheduler block, so dim0 blocks stay
byte-identical across engines with different TP. The page-sized form
made the MooncakeConnectorV2 handshake reject cross-TP transfers with
different local and remote kernel block size 1152 | 640.

Depends on vllm-project#16755.

### Does this PR introduce any user-facing change?

No new flags, config options, or API changes. For users running
GLM-5.3-Flash with PD disaggregation on the default model runner,
cross-TP deployments (e.g. P=TP8 + D=DP2xTP4) now work out of the box.
All other models and colocated workloads are unaffected – branch
selection is identical to before (only GLM-Next publishes the stride
marker).

How was this patch tested?

On GLM-5.3-Flash w8a8 real hardware (TP8 single instance, P=TP8 + D=TP8,
and P=TP8 → D=DP2xTP4):
- Indexer and tail views bit-identical with the MRV2 reshape.
- Equal-TP full-verify passes (Serial + concurrent, all match; per-rank
ms-level pulls).
- Unequal-TP chain: handshake clean, three factual questions all
answered correctly, zero pull/transfer errors.
- TextVQA authoritative rescore (per-row gold alignment, post-
extraction): PD 92.0% identical to colocated 90.0% in the same batch.


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: sunbaosong <13793883820@163.com>
lijiahang226 added a commit to lijiahang226/vllm-ascend that referenced this pull request Sep 25, 2026
Generate circular slot mappings for compressed cache groups, retain capture-safe metadata for graph and MTP execution, and reuse the native AttentionGroup builder hooks. Keep GLM cache views and copy layout in the existing model module. Refs vllm-project#15665 and vllm-project#16755.

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants