Skip to content

[Feature][MRV2] Adapt extract_hidden_states for Model Runner V2 - #14699

Merged
weiguihua2 merged 10 commits into
vllm-project:mainfrom
yjyang62:mrv2-extract-hidden-states-5b6a
Sep 9, 2026
Merged

weiguihua2 merged 10 commits into
vllm-project:mainfrom
yjyang62:mrv2-extract-hidden-states-5b6a

Conversation

@yjyang62

@yjyang62 yjyang62 commented Aug 21, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Upstream vLLM PR #49811 added Model Runner V2 support for extract_hidden_states. Ascend already supports this method on Model Runner V1, but Ascend MRV2 previously raised NotImplementedError in init_speculator, and the Ascend V2 KV allocate/reshape path treated HiddenStateCacheSpec like MLA K/V (split tensors).

This PR adapts Ascend MRV2:

  • thin-wraps upstream ExtractHiddenStatesSpeculator as AscendExtractHiddenStatesSpeculator and dispatches it from Ascend init_speculator (depends on upstream PR #49811 / a vLLM build that ships that module)
  • forces use_aux_hidden_state_outputs=True for extract_hidden_states in NPUModelRunner when the pinned vLLM omits the method from the GPU allow-list
  • keeps HiddenStateCacheSpec / cache_only_layers on a single-tensor allocate + reshape path in MRV2 attn_utils
  • adds MRV2 e2e cases and documents VLLM_USE_V2_MODEL_RUNNER=1

Does this PR introduce any user-facing change?

Yes. With VLLM_USE_V2_MODEL_RUNNER=1, users can run extract_hidden_states on Model Runner V2 (once the Ascend image/vLLM includes PR #49811).

How was this patch tested?

  • Unit tests (CPU), 6 passed:
    • tests/ut/worker/test_extract_hidden_states_speculator_v2.py
    • tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache
  • E2E cases added in tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py (dense_eager_mrv2, hybrid_dummy_eager_mrv2). NPU e2e was not run in the cloud agent environment (no Ascend device); please verify on Atlas hardware.
pytest -sv tests/ut/worker/test_extract_hidden_states_speculator_v2.py \
  tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache

export VLLM_USE_V2_MODEL_RUNNER=1
pytest -sv tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py -k mrv2

- vLLM main: https://github.com/vllm-project/vllm/commit/b2f685834a6456197e7033966fdef52a23f1abcd

@yjyang62
yjyang62 requested review from LCAIZJ and Yikun as code owners August 21, 2026 02:22
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enables extract_hidden_states functionality for Model Runner V2 on Ascend hardware. It integrates upstream vLLM changes by creating an Ascend-specific wrapper for the speculator, ensuring correct NPU memory handling for hidden state caches, and enforcing necessary configuration flags to support the feature when running in V2 mode.

Highlights

  • Model Runner V2 Support: Implemented support for extract_hidden_states in Model Runner V2 by wrapping the upstream ExtractHiddenStatesSpeculator as AscendExtractHiddenStatesSpeculator.
  • NPU Memory Management: Forced use_aux_hidden_state_outputs=True in NPUModelRunner and ensured HiddenStateCacheSpec maintains a single-tensor allocation and reshape path in MRV2.
  • Testing and Documentation: Added unit tests for the new speculator and NPU memory allocation, along with E2E cases and updated user documentation for enabling MRV2.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@cursor
cursor Bot force-pushed the mrv2-extract-hidden-states-5b6a branch 3 times, most recently from 66db379 to f9b2fc6 Compare August 21, 2026 02:23

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:\n\nmarkdown\n[Attention][Feature] Support Model Runner V2 for extract_hidden_states speculative decoding on Ascend\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis pull request adds support for Model Runner V2 (MRV2) for the `extract_hidden_states` speculative decoding mode on Ascend. It thin-wraps upstream vLLM's `ExtractHiddenStatesSpeculator`, adds dispatching in `init_speculator`, ensures auxiliary hidden state outputs are enabled, and handles allocation and reshaping for `HiddenStateCacheSpec` on a single-tensor path.\n\nFeedback on the changes suggests reusing the existing helper function `_allocate_int8_cache_tensor` in `vllm_ascend/worker/v2/attn_utils.py` to avoid duplicating aligned/unaligned tensor allocation logic.\n\n### Does this PR introduce _any_ user-facing change?\nYes, users can now enable Model Runner V2 for `extract_hidden_states` on Ascend by setting the environment variable `VLLM_USE_V2_MODEL_RUNNER=1`.\n\n### How was this patch tested?\nThe changes were tested with new end-to-end tests in `test_extract_hidden_states.py` covering dense and hybrid models with MRV2, as well as unit tests in `test_attn_utils_v2.py` and `test_extract_hidden_states_speculator_v2.py`.\n

Comment thread vllm_ascend/worker/v2/attn_utils.py Outdated
Comment on lines +649 to +678
if vllm_config.kv_transfer_config is None:
tensor = torch.zeros(kv_cache_tensor.size, dtype=torch.int8, device=device)
else:
tensor = torch.zeros(
kv_cache_tensor.size + alignment,
dtype=torch.int8,
device=device,
)
tensor = _align_memory(tensor, alignment)[: kv_cache_tensor.size]

if has_mamba and has_hidden:
# Keep Mamba and hidden-state dumps on separate physical buffers
# so float32 SSM writes cannot corrupt bfloat16 hidden states.
for layer_name in kv_cache_tensor.shared_by:
if is_hidden_state_cache_spec(layer_kv_cache_spec[layer_name]):
if vllm_config.kv_transfer_config is None:
hidden_tensor = torch.zeros(kv_cache_tensor.size, dtype=torch.int8, device=device)
else:
hidden_tensor = torch.zeros(
kv_cache_tensor.size + alignment,
dtype=torch.int8,
device=device,
)
hidden_tensor = _align_memory(hidden_tensor, alignment)[: kv_cache_tensor.size]
kv_cache_raw_tensors[layer_name] = hidden_tensor
else:
kv_cache_raw_tensors[layer_name] = tensor
else:
for layer_name in kv_cache_tensor.shared_by:
kv_cache_raw_tensors[layer_name] = tensor

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Instead of duplicating the aligned/unaligned tensor allocation logic for tensor and hidden_tensor, you can reuse the existing helper function _allocate_int8_cache_tensor defined in the same file. This improves code maintainability and readability.

            tensor = _allocate_int8_cache_tensor(kv_cache_tensor.size, alignment, device)

            if has_mamba and has_hidden:
                # Keep Mamba and hidden-state dumps on separate physical buffers
                # so float32 SSM writes cannot corrupt bfloat16 hidden states.
                for layer_name in kv_cache_tensor.shared_by:
                    if is_hidden_state_cache_spec(layer_kv_cache_spec[layer_name]):
                        hidden_tensor = _allocate_int8_cache_tensor(kv_cache_tensor.size, alignment, device)
                        kv_cache_raw_tensors[layer_name] = hidden_tensor
                    else:
                        kv_cache_raw_tensors[layer_name] = tensor
            else:
                for layer_name in kv_cache_tensor.shared_by:
                    kv_cache_raw_tensors[layer_name] = tensor

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests labels Aug 21, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@cursor
cursor Bot force-pushed the mrv2-extract-hidden-states-5b6a branch from f9b2fc6 to 10bcc54 Compare August 21, 2026 02:25
@yjyang62
yjyang62 force-pushed the mrv2-extract-hidden-states-5b6a branch from 10bcc54 to 4bddca1 Compare August 21, 2026 02:35
@yjyang62 yjyang62 changed the title [Feat][MRV2] Adapt extract_hidden_states for Model Runner V2 [Feature][MRV2] Adapt extract_hidden_states for Model Runner V2 Aug 21, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@cursor
cursor Bot force-pushed the mrv2-extract-hidden-states-5b6a branch from a7321c2 to 95e967e Compare August 25, 2026 02:05
@cursor
cursor Bot force-pushed the mrv2-extract-hidden-states-5b6a branch from 95e967e to fbe004a Compare August 25, 2026 02:05
@yjyang62
yjyang62 force-pushed the mrv2-extract-hidden-states-5b6a branch from fbe004a to e1209a8 Compare August 25, 2026 03:05
@cursor
cursor Bot force-pushed the mrv2-extract-hidden-states-5b6a branch from e1209a8 to e043262 Compare August 26, 2026 03:57
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@cursor
cursor Bot force-pushed the mrv2-extract-hidden-states-5b6a branch from cc88a4d to 749ae87 Compare September 9, 2026 01:54
@yjyang62 yjyang62 added the mrv2 label Sep 9, 2026

@Ronald1995 Ronald1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

yjyang62 and others added 9 commits September 9, 2026 07:59
Port extract_hidden_states to Model Runner V2 on the 0828 pin
(e6bfe03ad / vLLM #51718). Dispatch upstream ExtractHiddenStatesSpeculator,
keep HiddenStateCacheSpec on private [B, H, N, C] buffers so dumps cannot
overlay hybrid Attention/Mamba backing, and cover allocate/reshape plus
speculator dispatch in unit tests.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
The speculator crashed with "Expected 3 auxiliary hidden states, got 2"
on TP workers. NPU torch.compile graph-breaks at TP collectives can drop
early Python-list aux appends, and example ids such as [2, 18, 34] are
silently skipped on models with fewer than 35 layers.

Disable compile on the target backbone after load, reject out-of-range
layer ids with the EAGLE3 default, enable MiniMax aux collection for
this method, and tolerate residual=None in the DeepSeek V2 aux path.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Drop the Ascend speculator subclass and call upstream
ExtractHiddenStatesSpeculator from init_speculator. Keep only NPU glue:
disable target compile so aux list collection survives TP graph-breaks,
and reject out-of-range eagle_aux_hidden_state_layer_ids after load.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Yesterday's 0828 port already reused upstream ExtractHiddenStatesSpeculator
and only needed private HiddenStateCacheSpec allocate/reshape after
vLLM #51718. The later helper, load_model hook, MiniMax method list, and
DeepSeek residual=None patch were added for a compile/OOB path that the
proven Qwen3.5-35B-A3B TP=8 enforce-eager dump does not need.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
CacheOnlyAttentionBackend dropped get_kv_cache_shape in #51718 and has
not restored it. Reshape the dump cache from HiddenStateCacheSpec
[B, H, N, C] properties at the call site instead of a helper that
could pick up a pre-51718 backend layout.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
The cache_only_layers name check was a V1 fallback for shadowed spec
classes. V2 get_kv_cache_spec does not rewrite CacheOnly layers, and
allocate already keys off HiddenStateCacheSpec, so the layer-name OR
is redundant.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Cover skip_tokenizer_init + TokensPrompt on the generic generate path
and on extract_hidden_states dumps, including Model Runner V2.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep token-in/token-out coverage inside extract_hidden_states e2e, but
remove the extra one-card case that was not part of the MRV2 dump work.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
@Ronald1995 Ronald1995 added ready-precise run selected e2e test for pr and removed ready-all run all e2e test for pr labels Sep 9, 2026
@cursor
cursor Bot force-pushed the mrv2-extract-hidden-states-5b6a branch from c3b6bfc to 7def8c6 Compare September 9, 2026 08:00
Qwen3.5-0.8B is multimodal, so skip_tokenizer_init leaves tokenizer=None
and Qwen3VLProcessor crashes during LLM init. Keep the token-in/token-out
path, but cover it with dummy Qwen3-8B instead of the hybrid VL model.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
@weiguihua2
weiguihua2 merged commit 58076f2 into vllm-project:main Sep 9, 2026
23 checks passed
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
…-project#14699)

### What this PR does / why we need it?

Upstream [vLLM PR
#49811](vllm-project/vllm#49811) added Model
Runner V2 support for `extract_hidden_states`. Ascend already supports
this method on Model Runner V1, but Ascend MRV2 previously raised
`NotImplementedError` in `init_speculator`, and the Ascend V2 KV
allocate/reshape path treated `HiddenStateCacheSpec` like MLA K/V (split
tensors).

This PR adapts Ascend MRV2:

- thin-wraps upstream `ExtractHiddenStatesSpeculator` as
`AscendExtractHiddenStatesSpeculator` and dispatches it from Ascend
`init_speculator` (**depends on upstream PR #49811 / a vLLM build that
ships that module**)
- forces `use_aux_hidden_state_outputs=True` for `extract_hidden_states`
in `NPUModelRunner` when the pinned vLLM omits the method from the GPU
allow-list
- keeps `HiddenStateCacheSpec` / `cache_only_layers` on a single-tensor
allocate + reshape path in MRV2 `attn_utils`
- adds MRV2 e2e cases and documents `VLLM_USE_V2_MODEL_RUNNER=1`

### Does this PR introduce _any_ user-facing change?

Yes. With `VLLM_USE_V2_MODEL_RUNNER=1`, users can run
`extract_hidden_states` on Model Runner V2 (once the Ascend image/vLLM
includes PR #49811).

### How was this patch tested?

- Unit tests (CPU), 6 passed:
  - `tests/ut/worker/test_extract_hidden_states_speculator_v2.py`
-
`tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache`
- E2E cases added in
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
(`dense_eager_mrv2`, `hybrid_dummy_eager_mrv2`). NPU e2e was not run in
the cloud agent environment (no Ascend device); please verify on Atlas
hardware.

```bash
pytest -sv tests/ut/worker/test_extract_hidden_states_speculator_v2.py \
  tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache

export VLLM_USE_V2_MODEL_RUNNER=1
pytest -sv tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py -k mrv2

- vLLM main: vllm-project/vllm@b2f6858

---------

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
weiguihua2 pushed a commit that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [#14068][ascmoon],
[#14699][ascextract] and [#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[#16306](#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend #14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend #14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend #14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend #14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend #14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend #15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend #15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend #14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…-project#14699)

### What this PR does / why we need it?

Upstream [vLLM PR
#49811](vllm-project/vllm#49811) added Model
Runner V2 support for `extract_hidden_states`. Ascend already supports
this method on Model Runner V1, but Ascend MRV2 previously raised
`NotImplementedError` in `init_speculator`, and the Ascend V2 KV
allocate/reshape path treated `HiddenStateCacheSpec` like MLA K/V (split
tensors).

This PR adapts Ascend MRV2:

- thin-wraps upstream `ExtractHiddenStatesSpeculator` as
`AscendExtractHiddenStatesSpeculator` and dispatches it from Ascend
`init_speculator` (**depends on upstream PR #49811 / a vLLM build that
ships that module**)
- forces `use_aux_hidden_state_outputs=True` for `extract_hidden_states`
in `NPUModelRunner` when the pinned vLLM omits the method from the GPU
allow-list
- keeps `HiddenStateCacheSpec` / `cache_only_layers` on a single-tensor
allocate + reshape path in MRV2 `attn_utils`
- adds MRV2 e2e cases and documents `VLLM_USE_V2_MODEL_RUNNER=1`

### Does this PR introduce _any_ user-facing change?

Yes. With `VLLM_USE_V2_MODEL_RUNNER=1`, users can run
`extract_hidden_states` on Model Runner V2 (once the Ascend image/vLLM
includes PR #49811).

### How was this patch tested?

- Unit tests (CPU), 6 passed:
  - `tests/ut/worker/test_extract_hidden_states_speculator_v2.py`
-
`tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache`
- E2E cases added in
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
(`dense_eager_mrv2`, `hybrid_dummy_eager_mrv2`). NPU e2e was not run in
the cloud agent environment (no Ascend device); please verify on Atlas
hardware.

```bash
pytest -sv tests/ut/worker/test_extract_hidden_states_speculator_v2.py \
  tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache

export VLLM_USE_V2_MODEL_RUNNER=1
pytest -sv tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py -k mrv2

- vLLM main: vllm-project/vllm@b2f6858

---------

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…-project#14699)

### What this PR does / why we need it?

Upstream [vLLM PR
#49811](vllm-project/vllm#49811) added Model
Runner V2 support for `extract_hidden_states`. Ascend already supports
this method on Model Runner V1, but Ascend MRV2 previously raised
`NotImplementedError` in `init_speculator`, and the Ascend V2 KV
allocate/reshape path treated `HiddenStateCacheSpec` like MLA K/V (split
tensors).

This PR adapts Ascend MRV2:

- thin-wraps upstream `ExtractHiddenStatesSpeculator` as
`AscendExtractHiddenStatesSpeculator` and dispatches it from Ascend
`init_speculator` (**depends on upstream PR #49811 / a vLLM build that
ships that module**)
- forces `use_aux_hidden_state_outputs=True` for `extract_hidden_states`
in `NPUModelRunner` when the pinned vLLM omits the method from the GPU
allow-list
- keeps `HiddenStateCacheSpec` / `cache_only_layers` on a single-tensor
allocate + reshape path in MRV2 `attn_utils`
- adds MRV2 e2e cases and documents `VLLM_USE_V2_MODEL_RUNNER=1`

### Does this PR introduce _any_ user-facing change?

Yes. With `VLLM_USE_V2_MODEL_RUNNER=1`, users can run
`extract_hidden_states` on Model Runner V2 (once the Ascend image/vLLM
includes PR #49811).

### How was this patch tested?

- Unit tests (CPU), 6 passed:
  - `tests/ut/worker/test_extract_hidden_states_speculator_v2.py`
-
`tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache`
- E2E cases added in
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
(`dense_eager_mrv2`, `hybrid_dummy_eager_mrv2`). NPU e2e was not run in
the cloud agent environment (no Ascend device); please verify on Atlas
hardware.

```bash
pytest -sv tests/ut/worker/test_extract_hidden_states_speculator_v2.py \
  tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache

export VLLM_USE_V2_MODEL_RUNNER=1
pytest -sv tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py -k mrv2

- vLLM main: vllm-project/vllm@b2f6858

---------

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Signed-off-by: tianming2009 <13246728590@163.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…-project#14699)

### What this PR does / why we need it?

Upstream [vLLM PR
#49811](vllm-project/vllm#49811) added Model
Runner V2 support for `extract_hidden_states`. Ascend already supports
this method on Model Runner V1, but Ascend MRV2 previously raised
`NotImplementedError` in `init_speculator`, and the Ascend V2 KV
allocate/reshape path treated `HiddenStateCacheSpec` like MLA K/V (split
tensors).

This PR adapts Ascend MRV2:

- thin-wraps upstream `ExtractHiddenStatesSpeculator` as
`AscendExtractHiddenStatesSpeculator` and dispatches it from Ascend
`init_speculator` (**depends on upstream PR #49811 / a vLLM build that
ships that module**)
- forces `use_aux_hidden_state_outputs=True` for `extract_hidden_states`
in `NPUModelRunner` when the pinned vLLM omits the method from the GPU
allow-list
- keeps `HiddenStateCacheSpec` / `cache_only_layers` on a single-tensor
allocate + reshape path in MRV2 `attn_utils`
- adds MRV2 e2e cases and documents `VLLM_USE_V2_MODEL_RUNNER=1`

### Does this PR introduce _any_ user-facing change?

Yes. With `VLLM_USE_V2_MODEL_RUNNER=1`, users can run
`extract_hidden_states` on Model Runner V2 (once the Ascend image/vLLM
includes PR #49811).

### How was this patch tested?

- Unit tests (CPU), 6 passed:
  - `tests/ut/worker/test_extract_hidden_states_speculator_v2.py`
-
`tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache`
- E2E cases added in
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
(`dense_eager_mrv2`, `hybrid_dummy_eager_mrv2`). NPU e2e was not run in
the cloud agent environment (no Ascend device); please verify on Atlas
hardware.

```bash
pytest -sv tests/ut/worker/test_extract_hidden_states_speculator_v2.py \
  tests/ut/worker/test_attn_utils_v2.py::test_mrv2_allocates_and_reshapes_hidden_state_cache

export VLLM_USE_V2_MODEL_RUNNER=1
pytest -sv tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py -k mrv2

- vLLM main: vllm-project/vllm@b2f6858

---------

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Signed-off-by: like-0517 <ithwlike@126.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:tests mrv2 ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants