Repository navigation
[Feature][Refactor][P/D] Add MooncakeConnectorV2 - #14068
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request refactors the Mooncake KV transfer connector to improve modularity and support a pull-based transfer mechanism. By introducing dedicated base classes for schedulers and workers, the implementation cleanly separates control plane logic from data plane transfer operations. This change enhances maintainability and provides a foundation for future connector types. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Implement Mooncake KV transfer connectors for AscendSuggested PR Summary:
### What this PR does / why we need it?
This PR implements the Mooncake KV transfer connectors for the Ascend backend, enabling direct device-to-device (D2D) KV cache transfers. It introduces base and pull-based implementations for both schedulers and workers, along with metadata, statistics, and Prometheus metrics tracking.
Several critical issues were identified during the review:
- In `pull_worker.py`, the source and destination address lists (`src_list` and `dst_list`) are swapped in multiple places during pull transfers, which would cause severe runtime crashes or memory corruption.
- The retry logic in `utils.py` for `zmq.REQ` sockets is broken because ZeroMQ REQ sockets cannot simply retry `send`/`recv` on failure without being recreated or using a `DEALER` socket.
- Type annotations and comparisons for `remote_num_kv_heads` in `pull_worker.py` can lead to a `TypeError` if it is `None`.
- Prompt length calculations in `pull_scheduler.py` rely on `prompt_token_ids`, which breaks embedding-based requests where token IDs are `None`.
- A potential `IndexError` exists in `base_scheduler.py` when popping from `_all_token_ids` if it is empty.
### Does this PR introduce _any_ user-facing change?
Yes, it introduces the Mooncake KV transfer connector for Ascend, allowing users to configure and use Mooncake for KV cache transfers.
### How was this patch tested?
No tests were added in this PR. It is recommended to add unit and integration tests for the Mooncake scheduler and worker transfer paths, particularly covering the edge cases identified in the review (e.g., embedding-based requests, empty token IDs, and connection retries).| src_list.append(local_block_addr + local_inner_offset) | ||
| dst_list.append(remote_block_addr + remote_inner_offset) |
There was a problem hiding this comment.
In MooncakePullRecvingThread, which runs on the D-side (consumer), the local memory is the destination and the remote memory is the source. Therefore, src_list (source) should contain the remote addresses, and dst_list (destination) should contain the local addresses. Currently, they are swapped here, which will cause the transfer engine to read from local memory and write to remote memory, leading to severe runtime crashes or memory corruption.
src_list.append(remote_block_addr + remote_inner_offset)\n dst_list.append(local_block_addr + local_inner_offset)| src_list.append(local_base_addr + local_block_id * local_block_stride) | ||
| dst_list.append(remote_base_addr + remote_block_id * remote_block_stride) |
There was a problem hiding this comment.
In MooncakePullRecvingThread, which runs on the D-side (consumer), the local memory is the destination and the remote memory is the source. Therefore, src_list (source) should contain the remote addresses, and dst_list (destination) should contain the local addresses. Currently, they are swapped here, which will cause the transfer engine to read from local memory and write to remote memory, leading to severe runtime crashes or memory corruption.
src_list.append(remote_base_addr + remote_block_id * remote_block_stride)\n dst_list.append(local_base_addr + local_block_id * local_block_stride)| src_list.append(local_base_addr + local_block_id * local_block_stride + local_inner_offset) | ||
| dst_list.append(remote_base_addr + remote_block_id * remote_block_stride + remote_inner_offset) |
There was a problem hiding this comment.
In MooncakePullRecvingThread, which runs on the D-side (consumer), the local memory is the destination and the remote memory is the source. Therefore, src_list (source) should contain the remote addresses, and dst_list (destination) should contain the local addresses. Currently, they are swapped here, which will cause the transfer engine to read from local memory and write to remote memory, leading to severe runtime crashes or memory corruption.
src_list.append(remote_base_addr + remote_block_id * remote_block_stride + remote_inner_offset)\n dst_list.append(local_base_addr + local_block_id * local_block_stride + local_inner_offset)| def ensure_zmq_send( | ||
| socket: zmq.Socket, # type: ignore | ||
| data: bytes, | ||
| path: str, | ||
| max_retries: int = 3, | ||
| ): | ||
| retries_left = max_retries | ||
| while True: | ||
| try: | ||
| socket.send(data) | ||
| return | ||
| except zmq.ZMQError as e: # type: ignore | ||
| retries_left -= 1 | ||
| if retries_left > 0: | ||
| logger.warning("Send failed. error=%s, attempts_left=%d. ", e, retries_left) | ||
| time.sleep(0.1) | ||
| else: | ||
| logger.error("Send failed after all retries. error=%s. ", e) | ||
| raise RuntimeError(f"Failed to send data to {path} after {max_retries} retries: {e}") | ||
|
|
||
|
|
||
| def ensure_zmq_recv( | ||
| socket: zmq.Socket, # type: ignore | ||
| path: str, | ||
| max_retries: int = 3, | ||
| ) -> bytes: | ||
| retries_left = max_retries | ||
| while True: | ||
| try: | ||
| return socket.recv() | ||
| except zmq.ZMQError as e: # type: ignore | ||
| retries_left -= 1 | ||
| if retries_left > 0: | ||
| logger.warning("Receive failed. error=%s, attempts_left=%d. ", e, retries_left) | ||
| time.sleep(0.1) | ||
| else: | ||
| logger.error("Receive failed after all retries. source=%s, error=%s. ", path, e) | ||
| raise RuntimeError(f"Failed to receive data after {max_retries} retries: {e}") | ||
|
|
There was a problem hiding this comment.
The retry logic in ensure_zmq_send and ensure_zmq_recv is broken for zmq.REQ sockets. In ZeroMQ, a REQ socket strictly enforces a send-receive state machine. If a recv() or send() times out or fails, the socket enters an exceptional state and any subsequent calls on the same socket will immediately fail with EFSM (state flow error). Therefore, simply retrying recv() or send() in a loop on the same socket will not work and will just repeatedly fail. The socket must be closed and recreated to retry, or a zmq.DEALER socket should be used instead.
| def _infer_total_num_kv_heads( | ||
| self, | ||
| local_num_kv_heads: int, | ||
| remote_num_kv_heads: int | None, |
There was a problem hiding this comment.
In _infer_total_num_kv_heads, remote_num_kv_heads is typed as int | None. However, the code performs comparisons like max(local_num_kv_heads, remote_num_kv_heads) and remote_num_kv_heads > 1 without checking if it is None. If remote_num_kv_heads is None, this will raise a TypeError and crash the worker. Since it is always an int in practice, the type annotation should be updated to int.
| remote_num_kv_heads: int | None, | |
| remote_num_kv_heads: int, |
| token_ids = request.prompt_token_ids or [] | ||
| actual = self._state_prefill_token_count(len(token_ids)) |
There was a problem hiding this comment.
The code uses len(request.prompt_token_ids or []) to determine the prompt length. However, if the request uses prompt embeddings instead of token IDs, request.prompt_token_ids will be None, resulting in a prompt length of 0 even though request.num_prompt_tokens is greater than 0. This will break KV transfer for embedding-based requests. request.num_prompt_tokens should be used instead.
| token_ids = request.prompt_token_ids or [] | |
| actual = self._state_prefill_token_count(len(token_ids)) | |
| actual = self._state_prefill_token_count(request.num_prompt_tokens) |
| prompt_token_ids = request.prompt_token_ids or [] | ||
| prompt_len = len(prompt_token_ids) |
There was a problem hiding this comment.
The code uses len(request.prompt_token_ids or []) to determine the prompt length. However, if the request uses prompt embeddings instead of token IDs, request.prompt_token_ids will be None, resulting in a prompt length of 0 even though request.num_prompt_tokens is greater than 0. This will break KV transfer for embedding-based requests. request.num_prompt_tokens should be used instead.
| prompt_token_ids = request.prompt_token_ids or [] | |
| prompt_len = len(prompt_token_ids) | |
| prompt_len = request.num_prompt_tokens |
| else: | ||
| return | ||
|
|
||
| request._all_token_ids.pop() |
There was a problem hiding this comment.
In _truncate_request_for_prefill, request._all_token_ids.pop() is called. If prompt_token_ids is None (e.g., when using prompt embeddings), _all_token_ids might be empty, and calling pop() on it will raise an IndexError and crash the engine. This should be guarded with an if request._all_token_ids: check.
if request._all_token_ids:\n request._all_token_ids.pop()7d3d5db to
e63d65a
Compare
70e23b5 to
a1484c9
Compare
…t#14068) Squashed from the 40 commits in upstream PR vllm-project#14068 and rebased onto 26190b4.
|
|
||
| logger.info("Initializing Mooncake Scheduler %s", engine_id) | ||
|
|
||
| def _needs_prefill_token_truncation(self) -> bool: |
There was a problem hiding this comment.
Is this method necessary?
There was a problem hiding this comment.
Yes. For state model like Qwen3.5 or DSV4, it's necessary to truncate the last token.
bab558d to
3c05c01
Compare
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com>
Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com>
fdcccf5 to
5f12fba
Compare
Patch loaded Mooncake modules directly instead of resolving dotted paths through parent packages that other CPU tests can replace. Cover a detached kv_p2p parent attribute in the socket context test. Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com>
### What this PR does / why we need it? Add `MooncakeConnectorV2`. More information and RFC will be added later. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? By CI. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com> Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Co-authored-by: wangxiaoteng <wangxiaoteng@huawei.com>
### What this PR does / why we need it?
#### Summary
Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.
| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |
Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.
#### Scope and version handling
- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [#14068][ascmoon],
[#14699][ascextract] and [#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[#16306](#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.
#### File-by-file changes and upstream evidence
Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.
##### Dependency pin (1 file)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Device test adaptations (4 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend #14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
##### Existing CPU test adaptations (24 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend #14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend #14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Runtime adaptations (32 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend #14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend #14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend #15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend #15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend #15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend #14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
#### Interface review and limitations
The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.
That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.
### Does this PR introduce _any_ user-facing change?
Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.
### How was this patch tested?
- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.
[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290
[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files
- vLLM main:
vllm-project/vllm@b2f6858
---------
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
### What this PR does / why we need it? Add `MooncakeConnectorV2`. More information and RFC will be added later. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? By CI. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com> Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Co-authored-by: wangxiaoteng <wangxiaoteng@huawei.com>
### What this PR does / why we need it?
#### Summary
Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.
| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |
Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.
#### Scope and version handling
- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.
#### File-by-file changes and upstream evidence
Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.
##### Dependency pin (1 file)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Device test adaptations (4 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
##### Existing CPU test adaptations (24 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Runtime adaptations (32 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
#### Interface review and limitations
The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.
That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.
### Does this PR introduce _any_ user-facing change?
Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.
### How was this patch tested?
- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.
[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290
[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files
- vLLM main:
vllm-project/vllm@b2f6858
---------
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
### What this PR does / why we need it? Add `MooncakeConnectorV2`. More information and RFC will be added later. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? By CI. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com> Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Co-authored-by: wangxiaoteng <wangxiaoteng@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
### What this PR does / why we need it?
#### Summary
Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.
| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |
Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.
#### Scope and version handling
- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.
#### File-by-file changes and upstream evidence
Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.
##### Dependency pin (1 file)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Device test adaptations (4 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
##### Existing CPU test adaptations (24 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Runtime adaptations (32 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
#### Interface review and limitations
The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.
That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.
### Does this PR introduce _any_ user-facing change?
Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.
### How was this patch tested?
- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.
[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290
[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files
- vLLM main:
vllm-project/vllm@b2f6858
---------
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
### What this PR does / why we need it? Add `MooncakeConnectorV2`. More information and RFC will be added later. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? By CI. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com> Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Co-authored-by: wangxiaoteng <wangxiaoteng@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
### What this PR does / why we need it?
#### Summary
Upgrade the verified vLLM main revision while retaining v0.28.0
compatibility.
| Input | Exact revision |
|---|---|
| Previous vLLM main | `b2f685834a6456197e7033966fdef52a23f1abcd` |
| Target vLLM main | `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d`
(September 9 target; unchanged by this revision) |
| Retained release | `v0.28.0` /
`2cf0a6915ce544dc493a0990f2ea38d81601128a` |
| Current Ascend PR base | `1b1709b9227ad3e0db2d82b9231c4962919b86f5` |
| Reviewed PR head | `225f7cfdbf1c479798643a7d70a7e2031e5466f2` |
Current PR-wide diff: **61 files, +819/-418**. The numbered inventory
below matches the current GitHub changed-file list one-to-one. Removed
changes and historical commit-by-commit progress are not part of this
inventory.
#### Scope and version handling
- Adapt the exact upstream contracts linked below. Where main and
release differ, select the lane with `vllm_version_is("0.28.0")`; share
logic where both contracts match.
- The owner also included baseline dual-version compatibility gaps. In
particular, #51031's KV/kernel block-size split predates this main
upgrade. Newly merged Ascend consumers from [vllm-project#14068][ascmoon],
[vllm-project#14699][ascextract] and [vllm-project#15969][ascpcp] require compatibility
handling; they are not mislabeled as newly introduced vLLM breaks.
- **No workflow changes remain in this PR.** CPU UT uses the inherited
single-main configuration. The existing `main2main` label continues to
select main/tag NPU validation.
- Keep existing tests and necessary test adaptations; do not add
guard-test files or cases. On release only, skip the three unsupported
MRV2 extract-hidden-states cases at the owner's request, rather than
asserting an initialization error. V1 on both versions and MRV2 on main
remain enabled.
- The existing owner-authorized release SFA PCP precision exclusion
remains a temporary exception, not a precision fix or passing result.
The rebased Ascend main now inherits
[vllm-project#16306](vllm-project#16306), which
temporarily skips DSpark PCP. This PR does not modify that test file or
add the skip. A green CI with this inherited skip is not an actual
DSpark pass; the independent precision investigation/fix remains outside
this PR.
- No new patch file is added. This does not add DBO, CUDA Triton
attention, NPU ReplaySSM or release MRV2 extract-hidden-states support.
#### File-by-file changes and upstream evidence
Lane key: **M** = main, **R** = release, **B** = both (possibly
version-selected); **T** = test adaptation; **E** = explicitly
authorized exception. Upstream links for baseline compatibility or test
exclusions document the relevant contract, not a claim that this upgrade
introduced the issue.
##### Dependency pin (1 file)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 1. `.github/vllm-main-verified.commit` | M | Advance the verified main
pin from b2f6858 to a97dacb; the release tag is unchanged. | [exact
upgrade range][range] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Device test adaptations (4 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 2.
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_compute_slot_mapping.py`
| B/T | Keep the inherited windowed-kernel test inputs. Main passes
separate KV/kernel sizes and enablement to its reference kernel. Release
lacks that reference ABI: first derive CP-local positions, then invoke
the release CP=1 physical-slot reference and mask non-local tokens.
Adapt the existing out-buffer fixture; add no cases. | [#53896
ABI][uniform]; [main source][blockM]; [release source][blockR]; [Ascend
vllm-project#16120][ascwindow] | 10 actual-source CPU simulations passed, including
both lanes; selected NPU CI passed; coverage limits in the testing
section |
| 3.
`tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py`
| R/E | Temporarily skip only
test_dsv3_2_sfa_pcp_model_runner_v2_graph_accuracy on v0.28.0 at the
owner's request. Keep main and other cases unchanged; this is NOT a
precision fix or an upstream-break claim. | [baseline test being
isolated][sfatest] | Owner-authorized precision exception; unresolved |
| 4.
`tests/e2e/pull_request/one_card/spec_decode/test_extract_hidden_states.py`
| R/T | Skip only the three existing cases with use_v2_model_runner=True
on v0.28.0, whose upstream configuration rejects extract_hidden_states
with MRV2. Keep all V1 cases and main MRV2 execution unchanged; no
rejection-contract guard is added. | [#49811 implementation][extract];
[release rejection][configR]; [Ascend vllm-project#14699 cases][ascextract] | Four
source-isolated version/runner combinations and Ruff passed; selected
device CI passed; coverage limits in the testing section |
| 5. `tests/e2e/pull_request/one_card/test_gumbel_sampling.py` | B/T |
Adapt the existing Gumbel tests to each lane's signature and installed
implementation identity. | [#54282 Gumbel signature][gumbel] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
##### Existing CPU test adaptations (24 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 6. `tests/ut/core/test_dyntra_lb_scheduler.py` | B/T | Build the
actual SchedulerOutput fields for each lane; verify boundary offers only
on main while retaining connector/EC metadata checks. This is fixture
compatibility, not new Dyntra behavior. | [#51358 scheduler
output][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 7. `tests/ut/core/test_recompute_scheduler.py` | B/T | Adapt the
existing scheduler test's drain fixture and assertion to main boundary
offers versus release partial-tail offers. | [#51358][boundary] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 8. `tests/ut/core/test_scheduler_connector_block_state.py` | B/T |
Verify Recompute/Dyntra handoff using real lane-specific output fields
and cleared scheduler-local state, not only dispatch mocks. |
[#51358][boundary] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 9. `tests/ut/distributed/ascend_store/test_pool_worker.py` | B/T | Use
torch.float32 in the Mamba spec fixture; uniformity/page-size evaluation
reaches get_dtype_size and cannot consume the old NumPy dtype. |
[#53896][uniform]; [spec definition][cacheM] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 10. `tests/ut/kv_offload/mooncake_v2/helpers.py` | B/T | Add a real
KVCacheTensor fixture builder: release shared_by/block-stride/offset
fields versus main layers/layer-stride fields. | [#51718
descriptor][layout]; [Ascend vllm-project#14068 consumer][ascmoon] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 11. `tests/ut/kv_offload/mooncake_v2/test_base_worker.py` | B/T | Use
the real descriptor helper while retaining ordering, packed-view, SFA
virtual-block and invalid-input assertions. | [#51718][layout]; [Ascend
vllm-project#14068][ascmoon] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 12. `tests/ut/kv_offload/mooncake_v2/test_utils.py` | B/T | Replace
main-shaped descriptor stand-ins with real lane descriptors; preserve
storage deduplication, independence and alignment checks. |
[#51718][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 13. `tests/ut/kv_offload/test_native_cpu_offload.py` | B/T | Supply
data_parallel_size and data_parallel_rank_local in both fixtures: both
pinned versions read them. Correct a false release assumption without
changing offload behavior. | [main config][offloadM]; [release
config][offloadR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 14. `tests/ut/models/test_deepseek_v4_vision_preprocess.py` | B/T |
Give _StubInfo and its context a working shared tokenizer so the new
processing-info path is tested without an incomplete mock. | [#54886
processor contract][vision] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 15. `tests/ut/models/test_glm5next_kv_cache.py` | B/T | Assert 16
physical storage rows through get_storage_block_size rather than main's
nullable raw field; retain the new GLM architecture and its tests. |
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 16. `tests/ut/ops/test_routed_experts.py` | B/T | Initialize
quant_method.moe_kernel in both fixtures because expert_map already
reads it on release too; no production MoE change. | [main
expert_map][expertsM]; [release expert_map][expertsR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 17. `tests/ut/patch/platform/test_deepseek_v4_thinking.py` | B/T |
Assert the actual reasoning-mode mapping on both pinned tokenizers;
remove false release-only expectations while retaining all eight cases.
| [main tokenizer][thinkingM]; [release tokenizer][thinkingR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 18. `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py`
| B/T | Supply the common use_eagle_block_drop fixture field and
main-only mamba_fine_grained_prefix_cache=False, preserving current
indexer configuration. | [#53388 scheduler][schedfixture];
[#53945][replay] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 19. `tests/ut/patch/platform/test_patch_structured_output.py` | B/T |
Provide actual request/stop-token fields and validation exceptions on
both lanes; retain the {2} stop-token payload and backend-locking
assertions. | [main grammar creation][grammarM]; [release grammar
creation][grammarR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 20. `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` | B/T
| Adapt the existing V1 helper test to the installed lane. | [main
configuration][configM]; [release configuration][configR] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 21. `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | B/T |
Use the version-aware max-layers helper in the existing shared-tuple
planner assertion. | [#53614][eagle]; [#53906][storage];
[#53896][packed] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 22. `tests/ut/patch/worker/test_patch_mamba_utils_uniform_groups.py` |
B/T | Verify release's Ascend uniform-group override and main's upstream
spec-keyed dictionary/MambaCopyBuffers contract. | [#53896][mamba];
[both-lane source][mambasrcR] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 23. `tests/ut/spec_decode/test_extract_hidden_states_proposer.py` |
B/T | Create CPU buffers without pinned memory on both lanes instead of
patching a PIN_MEMORY symbol absent from release; retain proposer
behavior assertions. | [main proposer][proposerM]; [release
proposer][proposerR] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 24. `tests/ut/test_compressed_prefix_cache.py` | B/T | Use physical
storage rows and pass replay_boundary to the three real manager calls
only on main; retain all cache/hash assertions. | [#53906][storage];
[#53945 cache_blocks][replay] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 25. `tests/ut/worker/test_attn_utils_v2.py` | B/T | Adapt existing
storage/binding fixtures and the expected hidden-state cache layout to
each lane. | [#53906][storage]; [#52506][bind]; [#51718 hidden-state
path][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 26. `tests/ut/worker/test_extract_hidden_states_speculator_v2.py` |
B/T | Run main's extract-hidden-states speculator tests only where that
module exists; explicitly assert release dispatch raises
NotImplementedError before import. | [#49811][extract]; [Ascend
vllm-project#14699][ascextract] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 27. `tests/ut/worker/test_model_runner_v1.py` | B/T | Mock the
version-aware Ascend Mamba copy-function producer in the real V1 runner
fixture. | [#53896][mamba] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 28. `tests/ut/worker/test_model_runner_v2.py` | B/T | Adapt existing
profiling/dummy-call kwargs and PCP buffer access to each lane. Preserve
original test inputs. | [#54436][input]; [#52506][dummy]; [#53515][pcp]
| Source reviewed; selected current-head CI passed; coverage limits in
the testing section |
| 29. `tests/ut/worker/test_pcp_manager_v2.py` | B/T | Construct main
ExecuteModelState with cudagraph_stats and other required current
fields; omit fields absent from release. | [#52358
ExecuteModelState][stats] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
##### Runtime adaptations (32 files)
| # / File | Lane | Change and reason | Upstream code / cause |
Verification |
|---|---|---|---|---|
| 30. `vllm_ascend/_310p/worker/v2/block_table.py` | B | Accept and
validate optional per-group slot enablement; write PAD for disabled
groups in the 310P-owned buffers on both lanes. | [#53896
BlockTables][uniform] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 31. `vllm_ascend/_310p/worker/v2/model_runner.py` | B | Retain release
InputBatch max_seq_len_np; forward the shared dummy-state flag; gate
main-only circular specs and derive physical storage rows. |
[#54436][input]; [#52506][dummy]; [#53896][uniform]; [#53906][storage] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 32. `vllm_ascend/_310p/worker/v2/model_state.py` | B | Forward
ubatch_id through the Ascend parent MRO; preserve the existing no-DBO
limitation. | [#50945][ubatch] | Source reviewed; selected current-head
CI passed; coverage limits in the testing section |
| 33. `vllm_ascend/attention/context_parallel/common_cp.py` | B |
Explicitly declare supports_dcp=True only on the mixin whose
implementations already perform DCP, because upstream now defaults to
False. | [#55780 capability default][dcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 34. `vllm_ascend/attention/context_parallel/dsa_cp.py` | B | Use
get_storage_block_size for DSA CP physical cache geometry instead of
directly reading the version-dependent spec field. | [#53906 storage
contract][storage] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 35. `vllm_ascend/attention/dsa_v1.py` | B | Use the same physical-row
helper in DSA attention cache geometry; preserve existing attention
behavior. | [#53906 storage contract][storage] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 36. `vllm_ascend/core/kv_cache_interface.py` | B | Keep release's
derived storage property without overriding main's optional dataclass
field; centralize physical-row derivation, including
UniformTypeKVCacheSpecs. | [#53906][storage]; [#51718 spec
fields][layout] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 37. `vllm_ascend/core/recompute_scheduler.py` | B | Explicitly
separate release producer partial-tail handoff from main boundary
snapshots and scheduler-local connector state. Pass CoW copy tasks and
EC manager metadata on both lanes: both pins define these fields. Keep
CoW lifetime and scheduling policy unchanged. | [#51358
handoff][boundary]; [main producer][schedM]; [release producer][schedR];
[main fields][schedoutM]; [release fields][schedoutR] | Source reviewed;
8 isolated output cases and local formatting passed; selected
current-head CI passed; coverage limits in the testing section |
| 38. `vllm_ascend/distributed/device_communicators/npu_communicator.py`
| B | Initialize fi_pcie_ipc_ar_comm=None because new upstream
cleanup/communication code reads it; NPU does not create the CUDA
communicator. | [#53576 reader][ipc] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 39.
`vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/base_worker.py` | B
| Use the existing version-aware get_kv_cache_tensor_layers helper
instead of main-only .layers when registering Mooncake cache views. |
[#51718 descriptor][layout]; [Ascend vllm-project#14068 new consumer][ascmoon] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 40. `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake/utils.py` | B
| Use that same layer helper when collecting/deduplicating physical
cache storage; preserve packed-layout semantics. | [#51718
descriptor][layout]; [Ascend vllm-project#14068][ascmoon] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 41. `vllm_ascend/models/deepseek_v4/model.py` | B | Advertise the
already implemented auxiliary-state-over-PP capability and expose
Ascend's existing pp_transport_aux_hidden_states_ prefix to the upstream
relay; release safely retains the declarations. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 42. `vllm_ascend/models/minimax_m3/minimax_m3.py` | B | Align existing
auxiliary hidden-state PP capability and transport prefix with the new
upstream interface, without adding a new model feature. | [#50514 model
protocol][aux] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 43. `vllm_ascend/ops/kimi_mla.py` | M | Relocate the existing concrete
Kimi MLA replacement into a lightweight ops module; preserve
Ascend-first double inheritance for OOT initialization and forward. No
model-module registration side effect. | [#52494 concrete caller][kimi];
[upstream wrapper][kimiwrap] | Source-isolated dispatch/MRO checks and
Ruff passed; selected current-head CI passed; coverage limits in the
testing section |
| 44. `vllm_ascend/ops/rel_pos_attention.py` | B | Delete the
constructor that only forwarded super incorrectly; inherit each lane's
exact upstream signature/initialization, including use_triton_attention.
NPU forward is unchanged. | [#55629 constructor/caller][relpos] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 45. `vllm_ascend/ops/triton/v2/block_table/compute_slot_mappings.py` |
M | Add only the version-selected disabled-group PAD path. The inherited
base now already implements separate KV/kernel addressing with a bounded
window; preserve that algorithm, rather than restoring the old full-row
staging/direct-load paths. | [#53896 enablement][uniform]; [Ascend
vllm-project#16120 inherited mapping][ascwindow] | Source diff verified: only
enablement arguments and disabled-group path differ from base; CPU
simulations passed; selected NPU CI passed; coverage limits in the
testing section |
| 46. `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | M |
Forward drop_eagle_checkpoint_block through the CP coordinator only on
main; retain release's older call contract. | [#53614
coordinator][eagle] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 47. `vllm_ascend/patch/platform/patch_kv_cache_utils.py` | B | Bind
main's live _get_packed_kv_cache_groups and matching max-layers helper;
retain release's old grouping hooks and base GLM/Kimi grouping fixes. |
[#53896 planner hooks][packed]; [Ascend vllm-project#15913 architecture][ascglm] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 48. `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` | B | On
release, override only the GPU-specific non-MLA PCP restriction already
supported by Ascend; keep all other config rejections. Select the
V1-helper patch by version. This is baseline compatibility, not a newly
introduced upstream break. | [main config][configM]; [release
config][configR]; [Ascend PCP consumer][ascpcp] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 49. `vllm_ascend/patch/worker/patch_bind_kv_cache.py` | B | Accept
kv_cache_groups, retain stable layer binding and call
share_replayssm_ring_trackers only on main after binding. Do not claim
new NPU ReplaySSM support. | [#52506 binding/ring metadata][bind] |
Source reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 50. `vllm_ascend/patch/worker/patch_mamba_utils.py` | B | Keep release
tuple grouping/copy functions and uniform-group override; on main
preserve upstream spec-keyed groups and select copy functions by
mamba_type in each consumer. | [#53896 grouping][mamba]; [main
source][mambasrcM]; [release source][mambasrcR] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 51. `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | B | Gate old
module-level get_pp_group/_should_share bindings versus main's local
imports. Remove obsolete temporary EP-config mutation now that upstream
preserves target settings; retain quantization inheritance, PP guard
restoration and the original target weight-attribute guard. | [#50514
bindings][dspp]; [#55472 config preservation][dspark] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 52. `vllm_ascend/utils.py` | M | Add the concrete Kimi wrapper mapping
to the existing central custom-op registration under an explicit
main-only version gate; retain the shared one-time guard and generic
release mapping. | [#52494 caller][kimi]; [upstream wrapper][kimiwrap] |
Source-isolated main/release import, mapping and repeat-call checks
passed; selected current-head CI passed; coverage limits in the testing
section |
| 53. `vllm_ascend/worker/model_runner_v1.py` | B | Produce the
lane-correct Mamba copy-function container and use physical storage rows
in cache views, preserving the rebased per-layer attention-backend
dispatch. | [#53896 V1 producer][producer]; [copy consumer][mamba];
[#53906][storage]; [Ascend vllm-project#15913][ascglm] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 54. `vllm_ascend/worker/v2/attn_utils.py` | B | Allocate hidden-state
views in release [B,N,H,C] versus main [B,H,N,C] order after removal of
the old shape hook; honor padded page size. | [#51718 hidden-state
layout][hidden] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 55. `vllm_ascend/worker/v2/block_table.py` | B | Version-gate the
parent enablement argument and normalize release KV/kernel size tensors
at startup and wake-up. Preserve the inherited window sizing and launch
configuration, forwarding enablement only on main. | [#53896][uniform];
[#51031][slot]; [main fields][blockM]; [release fields][blockR]; [Ascend
vllm-project#16120][ascwindow] | Both-lane actual-source output/padding simulations
and Ruff passed; selected NPU CI passed; coverage limits in the testing
section |
| 56. `vllm_ascend/worker/v2/model_runner.py` | B | Install the existing
combined-broadcast patch only on release; main inherits upstream
separate PP draft receive/update/broadcast. Retain release InputBatch
arguments and dummy-state/PCP buffer adaptations required by the new
Ascend override. | [#50514][pp]; [#54436][input]; [#52506][dummy];
[Ascend vllm-project#15969 new override][ascpcp] | Source reviewed; selected
current-head CI passed; coverage limits in the testing section |
| 57. `vllm_ascend/worker/v2/model_states/default.py` | B | Accept
upstream's optional ubatch_id in prepare_attn, defaulting to zero; keep
explicit rejection of unsupported DBO/nonzero microbatches. | [#50945
default model state][ubatch] | Source reviewed; selected current-head CI
passed; coverage limits in the testing section |
| 58. `vllm_ascend/worker/v2/model_states/mamba_hybrid.py` | B | Apply
the same prepare_attn signature contract to Mamba hybrid state,
preserving the no-DBO guard. | [#50945 Mamba state][ubatchm] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
| 59. `vllm_ascend/worker/v2/pcp_manager.py` | B | Use the already-owned
_input_buffers with an explicit non-None assertion instead of a
main-only convenience property; this adapts the PCP consumer newly added
to the Ascend base. | [#53515 buffer ownership][pcp]; [Ascend
vllm-project#15969][ascpcp] | Source reviewed; selected current-head CI passed;
coverage limits in the testing section |
| 60. `vllm_ascend/worker/v2/sample/gumbel.py` | B | Select exact public
signatures: release's positional logits_cache versus main's required
is_drafting before logits_cache. Forward into a shared internal
implementation while preserving existing salt, stride, temperature and
cache behavior. | [#54282 positional ABI][gumbel] | Source reviewed;
selected current-head CI passed; coverage limits in the testing section
|
| 61. `vllm_ascend/worker/v2/spec_decode/__init__.py` | B | Reject
release extract-hidden-states MRV2 dispatch explicitly before importing
the main-only module; keep main's supported implementation. Do not
backport the feature or merely hide collection errors. | [#49811
speculator][extract]; [Ascend vllm-project#14699 new dispatch][ascextract] | Source
reviewed; selected current-head CI passed; coverage limits in the
testing section |
#### Interface review and limitations
The completed b2f6858-to-a97dacb interface report used Ascend baseline
`9dfdd529e9bec72d452ecac22d4c512332323a5f`. Its findings were
source-reviewed. It is reused at the owner's request; no expensive
analyzer rerun is performed for this revision.
That report does not automatically cover later rebased Ascend consumers.
Those adaptations are linked individually above. The selected
current-head integration suite passed; this does not establish
exhaustive regression freedom. The release SFA exclusion and inherited
upstream DSpark skip remain disclosed limitations; successful required
checks will not establish DSpark precision recovery.
### Does this PR introduce _any_ user-facing change?
Advance the verified main dependency while preserving supported v0.28.0
behavior through explicit version handling. The release-only
unsupported-feature skip changes test reporting to SKIPPED; it does not
backport MRV2 extract-hidden-states support or change model execution.
### How was this patch tested?
- **Versions:** vLLM main `a97dacb7106ee49f39f3d1fc6ae1800ff724e01d` and
release `v0.28.0`. NPU CI covers both lanes; CPU UT uses main only.
- **CI:** [E2E run
34592721264](https://github.com/vllm-project/vllm-ascend/actions/runs/34592721264)
passed on attempt 2 for PR head `225f7cfdb`. This is a rerun success,
not a precision-fix claim.
- **Coverage limits:** the inherited DSpark PCP skip (vllm-project#16306),
owner-authorized release SFA PCP exclusion, and three unsupported
release MRV2 extract-hidden-states skips remain. Passing CI does not
validate these excluded cases or all PP topologies.
- **Local checks:** Ruff lint/format and `git diff --check` passed.
Earlier source-isolated CPU checks are noted in the file table; no local
NPU or full installed-vLLM validation is claimed. Full `bash format.sh
ci` was unavailable on Windows.
[slot]:
https://github.com/vllm-project/vllm/pull/51031/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[uniform]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-21052649468f36e592c9ed378a9cd7b4615c558fb4d10f034f7b4268d3b6f9e1
[mamba]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-8ab86225fa08e9a6700851191abcb3f4f8b0bba40e97a82aad700d7a23a4d7ec
[packed]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-fb2b41380ca86adbba904aff40b18a0815beb73c1198818912e71723383ab604
[storage]:
https://github.com/vllm-project/vllm/pull/53906/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[layout]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-f76cdfbf02dacd9dffbfbd0d9ad68a7a6ac0d8aed70834f95a2ae8ccd2e333cb
[hidden]:
https://github.com/vllm-project/vllm/pull/51718/files#diff-57af80f3ae4b519ad1ac9d1338716420aed243a8f13d652a122ea4661e93708b
[boundary]:
https://github.com/vllm-project/vllm/pull/51358/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[ipc]:
https://github.com/vllm-project/vllm/pull/53576/files#diff-cf486d32b11454e89103000ac0f9a93ee1ccbb8983c02e794b734a65301cd98b
[vision]:
https://github.com/vllm-project/vllm/pull/54886/files#diff-740d0a4b454d681455df7c9f56361deaafa7a5baaa9dce45e7fa0aefec3e8cbd
[kimi]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-49368b83ccd7bf5f68204323e211a87440e53b608f5ecf91ff25d4832535b1b9
[kimiwrap]:
https://github.com/vllm-project/vllm/pull/52494/files#diff-aad295e6de3607a46cea22f9f1cca768d4937dc5e1d3ca58105513358ed4e4c1
[dcp]:
https://github.com/vllm-project/vllm/pull/55780/files#diff-bdb2df4662d59d54517931b86406644ee1e0c33cb3e00afac078d2a9aa550200
[relpos]:
https://github.com/vllm-project/vllm/pull/55629/files#diff-53dffba908ca88d6f312c7e27606f8f6fa11fb476e279efa11d712576809e1dd
[eagle]:
https://github.com/vllm-project/vllm/pull/53614/files#diff-43875c71daa893ef7567e21633d9988c2baf95bef61e3a334a6d584d6444d725
[replay]:
https://github.com/vllm-project/vllm/pull/53945/files#diff-97c184a680b7a4bd7d58b11aa0073706533cc887d990eddb98469e7374025ab3
[bind]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-645d58630d5acf3a0b07226bfef1e890a584c32502ab97c3d4642070f39a783c
[dummy]:
https://github.com/vllm-project/vllm/pull/52506/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[pp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-8330021c29feb32718335218ea994b6a770dd182a94b1baeacb7f7f42214afb7
[aux]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2ae789d702f81147ec584f3c0077aa3dedb6ebd424320dc607a44b86e3780874
[dspark]:
https://github.com/vllm-project/vllm/pull/55472/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[dspp]:
https://github.com/vllm-project/vllm/pull/50514/files#diff-2223c5e6b44b302db9bff4023777cc88b4427f4fb9d4a04476f6d6e4bd687393
[ubatch]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-aa6dbff87b81bae23ded2803a01a8c4913fceb6fe39c6d9423a20ad4b89859be
[ubatchm]:
https://github.com/vllm-project/vllm/pull/50945/files#diff-d6ed06f4278c0f7cd58e0eae8a019b484fba760424e9604284b3529843ba2384
[input]:
https://github.com/vllm-project/vllm/pull/54436/files#diff-106f39c08266f186830bb8fcd7fb1df35c0aaa5fd0ac5c17aac64aeddee48902
[pcp]:
https://github.com/vllm-project/vllm/pull/53515/files#diff-45dad5fe975c2ec8d02124fa65579cf11bfe9169a94217ceae1de0ea29518127
[gumbel]:
https://github.com/vllm-project/vllm/pull/54282/files#diff-02f7c5a06a2d23d544a07d57af16c7ffb74886469c8844385f51a719e329e6b8
[extract]:
https://github.com/vllm-project/vllm/pull/49811/files#diff-b13061cd1be0bc80aaa9fce208fe9128aaaad25b3832c5f1693da1d4e8263f5e
[schedfixture]:
https://github.com/vllm-project/vllm/pull/53388/files#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4
[stats]:
https://github.com/vllm-project/vllm/pull/52358/files#diff-5823f988fc0264681a80db24ccaba4d364f14394815d83ec1d90944c09f571f0
[offloadM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_offload/config.py
[offloadR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/kv_offload/config.py
[expertsM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/model_executor/layers/fused_moe/routed_experts.py
[expertsR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/model_executor/layers/fused_moe/routed_experts.py
[thinkingM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/tokenizers/deepseek_v4.py
[thinkingR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/tokenizers/deepseek_v4.py
[grammarM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/structured_output/__init__.py
[grammarR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/structured_output/__init__.py
[proposerM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/spec_decode/extract_hidden_states.py
[proposerR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/spec_decode/extract_hidden_states.py
[configM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/config/vllm.py
[configR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/config/vllm.py
[cacheM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/kv_cache_interface.py
[blockM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/gpu/block_table.py
[blockR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/gpu/block_table.py
[mambasrcM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/worker/mamba_utils.py
[mambasrcR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/worker/mamba_utils.py
[sfatest]:
https://github.com/vllm-project/vllm-ascend/blob/b8b230112d9d1edb13b8df2cd4d422074c17d640/tests/e2e/pull_request/four_card/context_parallel/test_accuracy_v2.py
[ascglm]: https://github.com/vllm-project/vllm-ascend/pull/15913/files
[ascmoon]: https://github.com/vllm-project/vllm-ascend/pull/14068/files
[ascpcp]: https://github.com/vllm-project/vllm-ascend/pull/15969/files
[ascextract]:
https://github.com/vllm-project/vllm-ascend/pull/14699/files
[range]:
vllm-project/vllm@b2f6858...a97dacb
[producer]:
https://github.com/vllm-project/vllm/pull/53896/files#diff-80ee7e2a62f9dcfbb8a312dc4e3948557e97ef187290daebbcae1e28596bda29
[schedoutM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/output.py#L282-L292
[schedoutR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/output.py#L252-L265
[schedM]:
https://github.com/vllm-project/vllm/blob/a97dacb7106ee49f39f3d1fc6ae1800ff724e01d/vllm/v1/core/sched/scheduler.py#L1354-L1450
[schedR]:
https://github.com/vllm-project/vllm/blob/2cf0a6915ce544dc493a0990f2ea38d81601128a/vllm/v1/core/sched/scheduler.py#L1233-L1290
[ascwindow]:
https://github.com/vllm-project/vllm-ascend/pull/16120/files
- vLLM main:
vllm-project/vllm@b2f6858
---------
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
Add
MooncakeConnectorV2. More information and RFC will be added later.Does this PR introduce any user-facing change?
No.
How was this patch tested?
By CI.