Skip to content

[BugFix][v0.23.0][KV Pool] Include MTP KV in layerwise AscendStore transfer - #13454

Merged
yiz-liu merged 1 commit into
vllm-project:releases/v0.23.0from
tyy0829:rebase-pr-13384-v0.23.0
Aug 4, 2026
Merged

yiz-liu merged 1 commit into
vllm-project:releases/v0.23.0from
tyy0829:rebase-pr-13384-v0.23.0

Conversation

@tyy0829

@tyy0829 tyy0829 commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Rebase of #13384 onto releases/v0.23.0 to resolve merge conflicts.

Build layerwise execution and group offsets from the KV caches registered on each worker. This includes MTP caches while keeping PP-local task indices and dense per-group storage indices for hybrid models.

The conflict was in tests/ut/distributed/ascend_store/test_pool_worker.py where both the target branch and the PR added different test methods at the same location. Both tests are kept since they test independent functionality.

Does this PR introduce any user-facing change?

No user-facing API changes. Internal fix for KV pool layer registration including MTP layers.

How was this patch tested?

  • Unit tests added/updated: tests/ut/distributed/ascend_store/test_pool_worker.py
    • test_registered_layer_layout_includes_mtp_multi_group_and_pp
    • test_lookup_reuses_grouped_hashes_for_hit_resolution
  • CI verification
  • vLLM version: v0.23.0
  • vLLM main: vllm-project/vllm@ee0da84

Build layerwise execution and group offsets from the KV caches registered on each worker. This includes MTP caches while keeping PP-local task indices and dense per-group storage indices for hybrid models.

Signed-off-by: Pz1116 <zpbzpb123123@gmail.com>
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request updates the KV pool worker's layer registration logic to support MTP layers. By moving away from implicit physical layer indexing toward an explicit registration-based configuration, the system can now correctly manage layerwise execution and group offsets for hybrid models, including those with MTP draft layers. This change ensures that layerwise tasks are correctly mapped and processed even when the model architecture includes additional non-standard layers.

Highlights

  • KV Pool Layer Registration: Refactored the KV pool worker to properly register and handle MTP (Multi-Token Prediction) layers alongside standard attention layers.
  • Configuration Logic: Replaced the previous physical layer mapping with a more robust registration-based approach to ensure accurate layer layout for hybrid models.
  • Testing: Added a new unit test, test_registered_layer_layout_includes_mtp_multi_group_and_pp, to verify correct layer configuration in complex multi-group and PP scenarios.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the KVPoolWorker to dynamically configure layerwise execution and group offsets from the registered KV caches instead of statically initializing them from the configuration. This ensures that registered MTP draft layers are correctly included alongside pipeline-parallel local target layers. Additionally, a new unit test has been added to verify this layout. The reviewer's feedback correctly identifies that the PR title and summary do not adhere to the repository's style guide and provides compliant suggestions.

Comment on lines +278 to +281
# Registration provides the authoritative layer layout. In
# particular, model_config.get_num_layers() only reports the local PP
# target layers and does not include registered MTP draft layers.
self.local_layer_to_group_layers: dict[int, list[tuple[int, int]]] = {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request title and summary do not adhere to the repository's style guide. Please update the PR title and summary to match the following suggested formats:\n\nSuggested PR Title:\n\nmarkdown\n[releases/v0.23.0][Attention][BugFix] transfer registered MTP layers\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\n\nThis PR builds layerwise execution and group offsets from the KV caches registered on each worker. This includes MTP caches while keeping PP-local task indices and dense per-group storage indices for hybrid models.\n\n### Does this PR introduce _any_ user-facing change?\n\nNo.\n\n### How was this patch tested?\n\nTested with new unit tests in `tests/ut/distributed/ascend_store/test_pool_worker.py`:\n- `test_registered_layer_layout_includes_mtp_multi_group_and_pp`\n- `test_lookup_reuses_grouped_hashes_for_hit_resolution`\n

References
  1. The PR title and summary must follow the format specified in the Repository Style Guide. (link)

@Pz1116 Pz1116 added the ready label Aug 4, 2026
@tyy0829 tyy0829 changed the title [BugFix][KV Pool] transfer registered MTP layers [BugFix][v0.23.0][KV Pool] transfer registered MTP layers Aug 4, 2026
@tyy0829 tyy0829 changed the title [BugFix][v0.23.0][KV Pool] transfer registered MTP layers [BugFix][v0.23.0][KV Pool] Include MTP KV in layerwise AscendStore transfer Aug 4, 2026
@tyy0829

tyy0829 commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor Author

/cherry-pick main
[Bot]: cherry-pick command failed. Please check the workflow run for details.

physical_layers = set()
for layer_name in layer_names:
phys = self._extract_physical_layer_index(layer_name)
if phys >= getattr(self.hf_config, "num_hidden_layers", self.num_layers):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the draft layers will not be discarded now.
LGTM.

@yiz-liu
yiz-liu merged commit 37e09e3 into vllm-project:releases/v0.23.0 Aug 4, 2026
21 checks passed
Wyz-134 pushed a commit to Wyz-134/vllm-ascend that referenced this pull request Aug 5, 2026
* releases/v0.23.0: (104 commits)
  [Doc][BugFix] Update proxy script name in DeepSeek-V3.2 tutorial (vllm-project#13537)
  [Doc] Fix link errors and update documentation structure (vllm-project#13483)
  [BugFix][releases/v0.23.0] fix fiaV2 contiguous err in GQA (vllm-project#13458)
  [v0.23.0][BugFix] Isolate layerwise GVA keys by parallel rank (vllm-project#13513)
  [Doc][Feature] Add model support of Ascend 950 (vllm-project#13525)
  [Doc] fix DeepSeek V4 Flash&Pro model tutorial docs link error (vllm-project#13497)
  [Cherry-pick][releases/v0.23.0][Doc][Misc] Add limitation for reduce sample (from vllm-project#13468) (vllm-project#13469)
  [BugFix][v0.23.0][KV Pool] Include MTP KV in layerwise AscendStore transfer (vllm-project#13454)
  [Doc][Misc] Standardize TorchNPU capitalization and update Ascend 950 product terminology (vllm-project#13089)
  [v0.23.0][Doc] Translated Doc files 2026-08-04 (vllm-project#13437)
  [Misc][v0.23.0] Fix translation extraction for tables nested in tabs (vllm-project#13413)
  [Doc] Fix translation and formatting in documentation (vllm-project#13390)
  [releases/v0.23.0][Doc][Misc] Backport Kimi-K2-Thinking tuning docs to v0.23.0 (vllm-project#13361)
  [Doc] Deployment key parameter supplement- vllm-project#13297 (vllm-project#13299)
  [v0.23.0][Doc] Translated Doc files 2026-07-31 (vllm-project#13283)
  [Doc][Misc] Update max-num-seqs configurations in GLM5 tutorial (vllm-project#13203)
  [Cherry-pick][releases/v0.23.0][Doc][Misc] Add deployment reference notice for GLM-5 (from vllm-project#12958) (vllm-project#12960)
  [BugFix][v0.23.0][KV Pool] Guard batch_get_key_info before memcache backend init (vllm-project#13307)
  [DOC]Modify the scope of scenarios supported by CP (vllm-project#13303)
  Revert "[cherry-pick][v0.23.0][Performance] remove D2H sync in QLIMetadata builder for DSA_CP" (vllm-project#13289)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants