Skip to content

[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay. - #15908

Merged
yiz-liu merged 5 commits into
vllm-project:mainfrom
zhiyu-wa:new-refactor
Sep 11, 2026
Merged

yiz-liu merged 5 commits into
vllm-project:mainfrom
zhiyu-wa:new-refactor

Conversation

@zhiyu-wa

@zhiyu-wa zhiyu-wa commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?
Implements the UpdatableGraph design discussed in #13058.
Refactors the FIA and Speculative Decoding to work with the new UpdatableGraph design.

Manually validated on A3 with the following scenarios:

  • FIA: Qwen3-0.6B with MRV1 and MRV2
  • MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:

  • FIA: Qwen3-0.6B with MRV1 and MRV2
  • MTP: Qwen3.5-9B with MRV1 and MRV2
  • PA: gemma4 with MRV1 and MRV2

Does this PR introduce any user-facing change?
No.

How was this patch tested?
Temporarily validated on some model.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces an UpdatableGraph architecture to replace static ACL graph parameter updates on Ascend NPUs. By decoupling parameter resolution from graph execution, this change enables more flexible and efficient handling of dynamic model parameters, specifically for Fused Infer Attention (FIA) and speculative decoding scenarios. The implementation includes a new task registration system and ensures compatibility with existing graph replay mechanisms across the Ascend backend.

Highlights

  • UpdatableGraph Architecture: Introduced a new UpdatableGraph class to manage dynamic parameter updates for ACL graph replay, decoupling parameter resolution from graph execution.
  • FIA and Speculative Decoding Refactor: Refactored AscendAttentionBackendImpl and speculative decoding workflows to utilize the UpdatableGraph design for improved flexibility.
  • Dynamic Parameter Resolution: Added FIAParamProvider to handle dynamic parameter resolution, ensuring correct metadata handling during graph replay.
  • Utility Enhancements: Updated weak_ref_tensors to support recursive weak referencing of tensors within nested containers.
  • Integration: Integrated UpdatableGraph into existing worker and graph manager workflows, including patching torch.cuda.CUDAGraph.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Introduce UpdatableGraph for Fused Infer Attention on Ascend

Suggested PR Summary:

### What this PR does / why we need it?
This PR introduces `UpdatableGraph` to support updatable graphs for Fused Infer Attention (FIA) on Ascend NPUs, replacing the manual graph parameter updates with a task-registration and resolution system. This simplifies the graph replay path and improves maintainability.

Feedback has been provided to address several critical issues, including a signature mismatch in `use_updatable_graph`, potential `AttributeError`s when accessing uncompiled `_runnable` methods or `None` layers, unsafe `issubclass` checks, and the use of `assert` for runtime validation.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested via existing CI and dummy runs for speculative decoding.

Comment thread vllm_ascend/compilation/acl_graph.py Outdated
Comment thread vllm_ascend/compilation/acl_graph.py Outdated
Comment thread vllm_ascend/spec_decode/llm_base_proposer.py Outdated
Comment thread vllm_ascend/spec_decode/llm_base_proposer.py Outdated
Comment thread vllm_ascend/spec_decode/llm_base_proposer.py Outdated
Comment thread vllm_ascend/compilation/updatable_graph.py
Comment thread vllm_ascend/attention/attention_v1.py
@zhiyu-wa
zhiyu-wa force-pushed the new-refactor branch 2 times, most recently from 8e3d3ef to 9d9c169 Compare September 7, 2026 11:27
@yiz-liu yiz-liu added the ready-precise run selected e2e test for pr label Sep 8, 2026
@zhiyu-wa
zhiyu-wa force-pushed the new-refactor branch 2 times, most recently from c7f3e4a to 6d59749 Compare September 8, 2026 03:32
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@zhiyu-wa
zhiyu-wa force-pushed the new-refactor branch 2 times, most recently from 8cf4361 to 50d2574 Compare September 9, 2026 01:19
@yiz-liu yiz-liu added ready-a5 ready-all run all e2e test for pr labels Sep 9, 2026
Comment thread tests/ut/compilation/test_acl_graph.py
Comment thread vllm_ascend/attention/attention_v1.py Outdated
Comment thread vllm_ascend/attention/attention_v1.py
Comment thread vllm_ascend/attention/attention_v1.py Outdated
Comment thread vllm_ascend/attention/attention_v1.py
Comment thread vllm_ascend/worker/v2/aclgraph_utils.py Outdated
Comment thread vllm_ascend/worker/model_runner_v1.py Outdated
Comment thread vllm_ascend/utils.py Outdated
Comment thread vllm_ascend/utils.py Outdated
Comment thread vllm_ascend/utils.py Outdated
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@zhiyu-wa
zhiyu-wa force-pushed the new-refactor branch 2 times, most recently from 2c82adc to a897bdb Compare September 9, 2026 07:14
@yiz-liu

yiz-liu commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

/weekly GLM-5_1-W8A8_A3_weekly
weekly command triggered.

@yiz-liu

yiz-liu commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

/weekly Minimax_m2.7_in198_prefix99
weekly command triggered.

@yiz-liu

yiz-liu commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

/weekly Qwen3_5_397B_A17B_w8a8_mxfp8_A5
weekly command triggered.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: zhiyu-wa <1959864813@qq.com>
Signed-off-by: zhiyu-wa <1959864813@qq.com>
Signed-off-by: zhiyu-wa <1959864813@qq.com>
Signed-off-by: zhiyu-wa <1959864813@qq.com>
Signed-off-by: zhiyu-wa <1959864813@qq.com>
@yiz-liu
yiz-liu enabled auto-merge (squash) September 11, 2026 09:51
@yiz-liu
yiz-liu merged commit daa6441 into vllm-project:main Sep 11, 2026
31 checks passed
czydyy added a commit to czydyy/vllm-ascend that referenced this pull request Sep 12, 2026
…r ACL Graph Replay. (vllm-project#15908)"

This reverts commit daa6441.

Signed-off-by: chenzeyu <2978509328@qq.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…aph Replay. (vllm-project#15908)

**What this PR does / why we need it?**
Implements the `UpdatableGraph` design discussed in
vllm-project#13058.
Refactors the `FIA` and `Speculative Decoding` to work with the new
`UpdatableGraph` design.

Manually validated on A3 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-9B with MRV1 and MRV2
- PA: gemma4 with MRV1 and MRV2
 
**Does this PR introduce any user-facing change?**
No.

**How was this patch tested?**
Temporarily validated on some model.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: zhiyu-wa <1959864813@qq.com>
@yiz-liu

yiz-liu commented Sep 12, 2026 •

Copy link
Copy Markdown
Collaborator

/revert
[Bot]: revert completed successfully. New PR created: #16409

@czydyy

czydyy commented Sep 12, 2026 •

Copy link
Copy Markdown
Collaborator

/revert
[Bot]: revert branch has been refreshed. Existing PR: #16409

vllm-ascend-ci pushed a commit to vllm-ascend-ci/vllm-ascend that referenced this pull request Sep 12, 2026
yiz-liu pushed a commit that referenced this pull request Sep 12, 2026
…pdates for ACL Graph Replay." (#15908) (#16409)

Revert of PR #15908 (merged onto `main`).

Original PR: #15908
Original author: @zhiyu-wa
Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920`

---
**What this PR does / why we need it?**
Implements the `UpdatableGraph` design discussed in
#13058.
Refactors the `FIA` and `Speculative Decoding` to work with the new
`UpdatableGraph` design.

Manually validated on A3 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-9B with MRV1 and MRV2
- PA: gemma4 with MRV1 and MRV2
 
**Does this PR introduce any user-facing change?**
No.

**How was this patch tested?**
Temporarily validated on some model.


- vLLM main:
vllm-project/vllm@a97dacb

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
windshado added a commit to windshado/vllm-ascend that referenced this pull request Sep 14, 2026
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba

* 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits)
  Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py
  [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344)
  [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300)
  [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253)
  [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415)
  [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251)
  [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343)
  [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422)
  [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232)
  [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319)
  [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189)
  [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043)
  [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)
  [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243)
  [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252)
  [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321)
  [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918)
  [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067)
  [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214)
  [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259)
  ...
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…aph Replay. (vllm-project#15908)

**What this PR does / why we need it?**
Implements the `UpdatableGraph` design discussed in
vllm-project#13058.
Refactors the `FIA` and `Speculative Decoding` to work with the new
`UpdatableGraph` design.

Manually validated on A3 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-9B with MRV1 and MRV2
- PA: gemma4 with MRV1 and MRV2
 
**Does this PR introduce any user-facing change?**
No.

**How was this patch tested?**
Temporarily validated on some model.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: zhiyu-wa <1959864813@qq.com>
Signed-off-by: tianming2009 <13246728590@163.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…pdates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)

Revert of PR vllm-project#15908 (merged onto `main`).

Original PR: vllm-project#15908
Original author: @zhiyu-wa
Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920`

---
**What this PR does / why we need it?**
Implements the `UpdatableGraph` design discussed in
vllm-project#13058.
Refactors the `FIA` and `Speculative Decoding` to work with the new
`UpdatableGraph` design.

Manually validated on A3 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-9B with MRV1 and MRV2
- PA: gemma4 with MRV1 and MRV2
 
**Does this PR introduce any user-facing change?**
No.

**How was this patch tested?**
Temporarily validated on some model.


- vLLM main:
vllm-project/vllm@a97dacb

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…aph Replay. (vllm-project#15908)

**What this PR does / why we need it?**
Implements the `UpdatableGraph` design discussed in
vllm-project#13058.
Refactors the `FIA` and `Speculative Decoding` to work with the new
`UpdatableGraph` design.

Manually validated on A3 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-9B with MRV1 and MRV2
- PA: gemma4 with MRV1 and MRV2

**Does this PR introduce any user-facing change?**
No.

**How was this patch tested?**
Temporarily validated on some model.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: zhiyu-wa <1959864813@qq.com>
Signed-off-by: like-0517 <ithwlike@126.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…pdates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)

Revert of PR vllm-project#15908 (merged onto `main`).

Original PR: vllm-project#15908
Original author: @zhiyu-wa
Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920`

---
**What this PR does / why we need it?**
Implements the `UpdatableGraph` design discussed in
vllm-project#13058.
Refactors the `FIA` and `Speculative Decoding` to work with the new
`UpdatableGraph` design.

Manually validated on A3 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-27B with MRV1 and MRV2

Manually validated on A5 with the following scenarios:
- FIA: Qwen3-0.6B with MRV1 and MRV2
- MTP: Qwen3.5-9B with MRV1 and MRV2
- PA: gemma4 with MRV1 and MRV2

**Does this PR introduce any user-facing change?**
No.

**How was this patch tested?**
Temporarily validated on some model.

- vLLM main:
vllm-project/vllm@a97dacb

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Signed-off-by: like-0517 <ithwlike@126.com>
ZT-AIA pushed a commit that referenced this pull request Sep 19, 2026
### What this PR does / why we need it?

**Depends on #15908 for Full Decode Only graph replay. Please merge
after that dependency.**

Two independent MRV2 Qwen3.5 PP fixes:

- **Graph backend selection:** skip metadata-only attention backends
whose `get_impl_cls()` raises `NotImplementedError` while collecting
graph-update backends, so the PP stage can select an executable
attention backend (GDN-style backends have no attention impl).
- **Draft inputs:** disable external multimodal embedding collection on
the speculator when PP is enabled, set right after draft loading and
before profiling. The last PP stage has no encoder cache; ordinary token
embedding remains in the draft model. Non-PP behavior is unchanged.

Plus one four-card E2E (`test_qwen35_pp_mtp.py`) covering serial,
concurrent, and chunked-prefill requests.

### Does this PR introduce _any_ user-facing change?

Addresses PP + MTP backend-selection and draft-input failures for
Qwen3.5 under MRV2. Full Decode Only support still relies on #15908;
this PR is not a standalone replacement for that dependency.

### How was this patch tested?

- **E2E (CI four-card):**
`tests/e2e/pull_request/four_card/model_runner_v2/test_qwen35_pp_mtp.py`
— serial answers, concurrency, and a prompt exceeding the prefill
budget.

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
### What this PR does / why we need it?

**Depends on vllm-project#15908 for Full Decode Only graph replay. Please merge
after that dependency.**

Two independent MRV2 Qwen3.5 PP fixes:

- **Graph backend selection:** skip metadata-only attention backends
whose `get_impl_cls()` raises `NotImplementedError` while collecting
graph-update backends, so the PP stage can select an executable
attention backend (GDN-style backends have no attention impl).
- **Draft inputs:** disable external multimodal embedding collection on
the speculator when PP is enabled, set right after draft loading and
before profiling. The last PP stage has no encoder cache; ordinary token
embedding remains in the draft model. Non-PP behavior is unchanged.

Plus one four-card E2E (`test_qwen35_pp_mtp.py`) covering serial,
concurrent, and chunked-prefill requests.

### Does this PR introduce _any_ user-facing change?

Addresses PP + MTP backend-selection and draft-input failures for
Qwen3.5 under MRV2. Full Decode Only support still relies on vllm-project#15908;
this PR is not a standalone replacement for that dependency.

### How was this patch tested?

- **E2E (CI four-card):**
`tests/e2e/pull_request/four_card/model_runner_v2/test_qwen35_pp_mtp.py`
— serial answers, concurrency, and a prompt exceeding the prefill
budget.

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: LostFox11 <wangziyue17@huawei.com>
Co-authored-by: LostFox11 <wangziyue17@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants