Skip to content

[MRV2][Performance][Kernel] Use windowed block-table gather for V2 slot mappings - #16120

Merged
zzzzwwjj merged 1 commit into
vllm-project:mainfrom
Liamup777:perf/ascend-v2-slot-mappings-clean
Sep 11, 2026
Merged

zzzzwwjj merged 1 commit into
vllm-project:mainfrom
Liamup777:perf/ascend-v2-slot-mappings-clean

Conversation

@Liamup777

@Liamup777 Liamup777 commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Replace full block-table row staging in the Ascend V2 slot-mapping Triton kernel with a bounded contiguous window load followed by tl.gather.

This bounds UB use independently of block-table row length. The window is sized from the smallest kernel block size, because the upstream V0.28.0 block_sizes_tensor used by slot mapping stores kernel_block_sizes.

The kernel preserves the upstream CP mapping behavior, narrows positions to INT32, and computes block offsets with multiply/subtract instead of remainder. The V2 operator documentation is moved under ops/triton/docs/v2.

Does this PR introduce any user-facing change?

No.

How was this patch tested?

  • bash format.sh — passed.

  • bash format.sh ci — passed.

  • SHELLCHECK_OPTS="--exclude=SC2046,SC2006,SC2086" pre-commit run --all-files --hook-stage manual --show-diff-on-failure — passed.

  • python -m compileall -q vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py vllm_ascend/worker/v2/block_table.py tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py — passed.

  • NPU UT on lab-worker-a3-01 passed on the preceding revision:

    python -m pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py

    Result: 4 passed, 14 warnings in 17.74s. This covers CP=1/2/4, empty requests, cross-block/cross-tile ranges, mixed kernel block sizes, padding, and the out interface. Re-run is pending after the reviewer-requested tile-size source deduplication; the change preserves the configured value and kernel algorithm.

  • Per-case NPU msprof op comparison against the V0.28.0 upstream kernel. Task Duration(us):

    Case Grid Upstream Ascend Speedup
    tile-1 1x2 114.598 3.680 31.14x
    tile-64 1x65 229.255 6.680 34.32x
    tile-1024 1x1025 2934.781 39.399 74.49x
    long-1x8192 1x2 786.244 10.920 72.00x
    long-64x8192 1x65 1573.109 20.960 75.05x
    long-64x32768 1x65 6181.396 70.719 87.41x
    cp2-tile-64-rank0 1x65 498.110 43.579 11.43x
    cp4-tile-64-rank2 1x65 493.110 44.019 11.20x
  • vLLM main: vllm-project/vllm@b2f6858

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests module:ops labels Sep 9, 2026
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request optimizes the Ascend V2 slot-mapping Triton kernel by transitioning from full block-table row staging to a bounded contiguous window load approach. This change significantly improves memory efficiency by decoupling Unified Buffer (UB) usage from the length of block-table rows. The implementation now dynamically determines window sizes based on kernel block sizes, ensuring consistent performance. Additionally, documentation has been reorganized to align with the new structure, and the test suite has been updated to verify the correctness of the new windowing logic.

Highlights

  • Kernel Optimization: Implemented a windowed block-table gather in the Ascend V2 slot-mapping Triton kernel to replace full-row staging.
  • UB Usage Management: Bounded Unified Buffer (UB) usage independently of block-table row length by utilizing a contiguous window load.
  • Dynamic Window Sizing: Refactored AscendBlockTables to calculate window sizes dynamically based on kernel block sizes rather than static staging limits.
  • Documentation: Migrated and updated documentation for the slot-mapping kernel to the new ops/triton/docs/v2 directory.
  • Test Suite Updates: Updated test cases to validate the new windowing logic and removed obsolete staging parameters.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Implement bounded window gather for Ascend block table slot mapping

Suggested PR Summary:

### What this PR does / why we need it?
This PR replaces the full-row block-table staging mechanism with a bounded window gather approach in `_compute_slot_mappings_kernel` for Ascend. Instead of staging the entire block-table row (which could exceed UB for large rows), it calculates a power-of-two `BLOCK_TABLE_WINDOW_SIZE` based on the smallest kernel block size and stages only the contiguous portion of the request row used by the current token tile. This removes the need for `_MAX_STAGED_BLOCK_TABLE_PAD_SIZE` and `USE_BLOCK_TABLE_STAGING`.

Feedback:
The reviewer suggests avoiding hardcoding the token tile size `1024` in both `__init__` and the kernel launch by defining an instance variable `self._triton_block_size = 1024` to ensure consistency and prevent potential out-of-bounds indexing.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
The patch was tested using the existing single-operator test suite: `pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`.

Comment thread vllm_ascend/worker/v2/block_table.py Outdated
Comment on lines +57 to +63
# block_sizes_tensor stores kernel_block_sizes, which determine the
# number of block-table entries touched by one token tile. Use the
# smallest kernel block size to form one safe constexpr window for all
# groups, without staging a whole row.
min_kernel_block_size = min(kernel_block_sizes)
window_size = (1024 + min_kernel_block_size - 1) // min_kernel_block_size + 1
self._block_table_window_size = triton.next_power_of_2(window_size)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The token tile size 1024 is hardcoded both in the window size calculation here and in the kernel launch parameters (TRITON_BLOCK_SIZE=1024 on line 101). If the block size is changed in one place but not the other, it can lead to silent out-of-bounds indexing or correctness bugs.

To prevent this, define an instance variable self._triton_block_size = 1024 in __init__ and use it in both places.

Suggested change
# block_sizes_tensor stores kernel_block_sizes, which determine the
# number of block-table entries touched by one token tile. Use the
# smallest kernel block size to form one safe constexpr window for all
# groups, without staging a whole row.
min_kernel_block_size = min(kernel_block_sizes)
window_size = (1024 + min_kernel_block_size - 1) // min_kernel_block_size + 1
self._block_table_window_size = triton.next_power_of_2(window_size)
self._triton_block_size = 1024
# block_sizes_tensor stores kernel_block_sizes, which determine the
# number of block-table entries touched by one token tile. Use the
# smallest kernel block size to form one safe constexpr window for all
# groups, without staging a whole row.
min_kernel_block_size = min(kernel_block_sizes)
window_size = (self._triton_block_size + min_kernel_block_size - 1) // min_kernel_block_size + 1
self._block_table_window_size = triton.next_power_of_2(window_size)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, agreed. I introduced self._triton_block_size = 1024 in init and use it for both window-size calculation and the kernel TRITON_BLOCK_SIZE launch argument. The manually constructed out-interface test fixture now initializes the same attribute.

Comment on lines 99 to 103
CP_INTERLEAVE=self.cp_interleave,
PAD_ID=PAD_SLOT_ID,
TRITON_BLOCK_SIZE=1024,
BLOCK_TABLE_PAD_SIZE=self._block_table_pad_size,
USE_BLOCK_TABLE_STAGING=(self._block_table_pad_size <= _MAX_STAGED_BLOCK_TABLE_PAD_SIZE),
BLOCK_TABLE_WINDOW_SIZE=self._block_table_window_size,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Use the newly defined self._triton_block_size instead of the hardcoded 1024 constant to ensure consistency with the window size calculation.

Suggested change
CP_INTERLEAVE=self.cp_interleave,
PAD_ID=PAD_SLOT_ID,
TRITON_BLOCK_SIZE=1024,
BLOCK_TABLE_PAD_SIZE=self._block_table_pad_size,
USE_BLOCK_TABLE_STAGING=(self._block_table_pad_size <= _MAX_STAGED_BLOCK_TABLE_PAD_SIZE),
BLOCK_TABLE_WINDOW_SIZE=self._block_table_window_size,
)
CP_INTERLEAVE=self.cp_interleave,
PAD_ID=PAD_SLOT_ID,
TRITON_BLOCK_SIZE=self._triton_block_size,
BLOCK_TABLE_WINDOW_SIZE=self._block_table_window_size,
)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated in the same change: the launch now passes TRITON_BLOCK_SIZE=self._triton_block_size, so it cannot diverge from the window-size calculation.

# groups, without staging a whole row.
min_kernel_block_size = min(kernel_block_sizes)
window_size = (1024 + min_kernel_block_size - 1) // min_kernel_block_size + 1
self._block_table_window_size = triton.next_power_of_2(window_size)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Check if Gemini say it right : )

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed. Gemini’s consistency concern was valid, and the requested tile-size deduplication has now been applied.

@AuroraEmiya

Copy link
Copy Markdown
Contributor

Check if there is any performance fallback compared to which you already submitted

@Liamup777
Liamup777 force-pushed the perf/ascend-v2-slot-mappings-clean branch 2 times, most recently from 592f8a4 to eb9e3f9 Compare September 9, 2026 06:48
@Ronald1995 Ronald1995 added the ready-precise run selected e2e test for pr label Sep 9, 2026

@AuroraEmiya AuroraEmiya left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

…ot mappings

Signed-off-by: Liam <ml646@duke.edu>
@Liamup777
Liamup777 force-pushed the perf/ascend-v2-slot-mappings-clean branch from eb9e3f9 to 2cf6685 Compare September 10, 2026 09:14
@zzzzwwjj
zzzzwwjj merged commit be42704 into vllm-project:main Sep 11, 2026
24 checks passed
weiguihua2 pushed a commit that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [#14068][ascmoon],
[#14699][ascextract] and [#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[#16306](#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend #14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend #14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend #14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend #14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend #14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend #15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend #15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend #14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…ot mappings (vllm-project#16120)

### What this PR does / why we need it?

Replace full block-table row staging in the Ascend V2 slot-mapping
Triton kernel with a bounded contiguous window load followed by
`tl.gather`.

This bounds UB use independently of block-table row length. The window
is sized from the smallest kernel block size, because the upstream
V0.28.0 `block_sizes_tensor` used by slot mapping stores
`kernel_block_sizes`.

The kernel preserves the upstream CP mapping behavior, narrows positions
to INT32, and computes block offsets with multiply/subtract instead of
remainder. The V2 operator documentation is moved under
`ops/triton/docs/v2`.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- `bash format.sh` — passed.
- `bash format.sh ci` — passed.
- `SHELLCHECK_OPTS="--exclude=SC2046,SC2006,SC2086" pre-commit run
--all-files --hook-stage manual --show-diff-on-failure` — passed.
- `python -m compileall -q
vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py
vllm_ascend/worker/v2/block_table.py
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
— passed.
- NPU UT on `lab-worker-a3-01` passed on the preceding revision:
  ```bash
python -m pytest -sv
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py
  ```
Result: `4 passed, 14 warnings in 17.74s`. This covers CP=1/2/4, empty
requests, cross-block/cross-tile ranges, mixed kernel block sizes,
padding, and the `out` interface. Re-run is pending after the
reviewer-requested tile-size source deduplication; the change preserves
the configured value and kernel algorithm.
- Per-case NPU `msprof op` comparison against the V0.28.0 upstream
kernel. `Task Duration(us)`:

  | Case | Grid | Upstream | Ascend | Speedup |
  | --- | ---: | ---: | ---: | ---: |
  | tile-1 | 1x2 | 114.598 | 3.680 | 31.14x |
  | tile-64 | 1x65 | 229.255 | 6.680 | 34.32x |
  | tile-1024 | 1x1025 | 2934.781 | 39.399 | 74.49x |
  | long-1x8192 | 1x2 | 786.244 | 10.920 | 72.00x |
  | long-64x8192 | 1x65 | 1573.109 | 20.960 | 75.05x |
  | long-64x32768 | 1x65 | 6181.396 | 70.719 | 87.41x |
  | cp2-tile-64-rank0 | 1x65 | 498.110 | 43.579 | 11.43x |
  | cp4-tile-64-rank2 | 1x65 | 493.110 | 44.019 | 11.20x |

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Liam <ml646@duke.edu>
Co-authored-by: Liam <ml646@duke.edu>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…ot mappings (vllm-project#16120)

### What this PR does / why we need it?

Replace full block-table row staging in the Ascend V2 slot-mapping
Triton kernel with a bounded contiguous window load followed by
`tl.gather`.

This bounds UB use independently of block-table row length. The window
is sized from the smallest kernel block size, because the upstream
V0.28.0 `block_sizes_tensor` used by slot mapping stores
`kernel_block_sizes`.

The kernel preserves the upstream CP mapping behavior, narrows positions
to INT32, and computes block offsets with multiply/subtract instead of
remainder. The V2 operator documentation is moved under
`ops/triton/docs/v2`.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- `bash format.sh` — passed.
- `bash format.sh ci` — passed.
- `SHELLCHECK_OPTS="--exclude=SC2046,SC2006,SC2086" pre-commit run
--all-files --hook-stage manual --show-diff-on-failure` — passed.
- `python -m compileall -q
vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py
vllm_ascend/worker/v2/block_table.py
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
— passed.
- NPU UT on `lab-worker-a3-01` passed on the preceding revision:
  ```bash
python -m pytest -sv
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py
  ```
Result: `4 passed, 14 warnings in 17.74s`. This covers CP=1/2/4, empty
requests, cross-block/cross-tile ranges, mixed kernel block sizes,
padding, and the `out` interface. Re-run is pending after the
reviewer-requested tile-size source deduplication; the change preserves
the configured value and kernel algorithm.
- Per-case NPU `msprof op` comparison against the V0.28.0 upstream
kernel. `Task Duration(us)`:

  | Case | Grid | Upstream | Ascend | Speedup |
  | --- | ---: | ---: | ---: | ---: |
  | tile-1 | 1x2 | 114.598 | 3.680 | 31.14x |
  | tile-64 | 1x65 | 229.255 | 6.680 | 34.32x |
  | tile-1024 | 1x1025 | 2934.781 | 39.399 | 74.49x |
  | long-1x8192 | 1x2 | 786.244 | 10.920 | 72.00x |
  | long-64x8192 | 1x65 | 1573.109 | 20.960 | 75.05x |
  | long-64x32768 | 1x65 | 6181.396 | 70.719 | 87.41x |
  | cp2-tile-64-rank0 | 1x65 | 498.110 | 43.579 | 11.43x |
  | cp4-tile-64-rank2 | 1x65 | 493.110 | 44.019 | 11.20x |

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Liam <ml646@duke.edu>
Co-authored-by: Liam <ml646@duke.edu>
Signed-off-by: tianming2009 <13246728590@163.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…ot mappings (vllm-project#16120)

### What this PR does / why we need it?

Replace full block-table row staging in the Ascend V2 slot-mapping
Triton kernel with a bounded contiguous window load followed by
`tl.gather`.

This bounds UB use independently of block-table row length. The window
is sized from the smallest kernel block size, because the upstream
V0.28.0 `block_sizes_tensor` used by slot mapping stores
`kernel_block_sizes`.

The kernel preserves the upstream CP mapping behavior, narrows positions
to INT32, and computes block offsets with multiply/subtract instead of
remainder. The V2 operator documentation is moved under
`ops/triton/docs/v2`.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- `bash format.sh` — passed.
- `bash format.sh ci` — passed.
- `SHELLCHECK_OPTS="--exclude=SC2046,SC2006,SC2086" pre-commit run
--all-files --hook-stage manual --show-diff-on-failure` — passed.
- `python -m compileall -q
vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py
vllm_ascend/worker/v2/block_table.py
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
— passed.
- NPU UT on `lab-worker-a3-01` passed on the preceding revision:
  ```bash
python -m pytest -sv
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py
  ```
Result: `4 passed, 14 warnings in 17.74s`. This covers CP=1/2/4, empty
requests, cross-block/cross-tile ranges, mixed kernel block sizes,
padding, and the `out` interface. Re-run is pending after the
reviewer-requested tile-size source deduplication; the change preserves
the configured value and kernel algorithm.
- Per-case NPU `msprof op` comparison against the V0.28.0 upstream
kernel. `Task Duration(us)`:

  | Case | Grid | Upstream | Ascend | Speedup |
  | --- | ---: | ---: | ---: | ---: |
  | tile-1 | 1x2 | 114.598 | 3.680 | 31.14x |
  | tile-64 | 1x65 | 229.255 | 6.680 | 34.32x |
  | tile-1024 | 1x1025 | 2934.781 | 39.399 | 74.49x |
  | long-1x8192 | 1x2 | 786.244 | 10.920 | 72.00x |
  | long-64x8192 | 1x65 | 1573.109 | 20.960 | 75.05x |
  | long-64x32768 | 1x65 | 6181.396 | 70.719 | 87.41x |
  | cp2-tile-64-rank0 | 1x65 | 498.110 | 43.579 | 11.43x |
  | cp4-tile-64-rank2 | 1x65 | 493.110 | 44.019 | 11.20x |

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Liam <ml646@duke.edu>
Co-authored-by: Liam <ml646@duke.edu>
Signed-off-by: like-0517 <ithwlike@126.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

#### Summary

Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.

| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |

Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.

#### Scope and version handling

- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.

#### File-by-file changes and upstream evidence

Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.

##### Dependency pin (1 file)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Device test adaptations (4 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

##### Existing CPU test adaptations (24 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |

##### Runtime adaptations (32 files)

| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |

#### Interface review and limitations

The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.

That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.

### Does this PR introduce _any_ user-facing change?

Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.

### How was this patch tested?

- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.

[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290

[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:ops module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants