[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay. - #15908
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces an UpdatableGraph architecture to replace static ACL graph parameter updates on Ascend NPUs. By decoupling parameter resolution from graph execution, this change enables more flexible and efficient handling of dynamic model parameters, specifically for Fused Infer Attention (FIA) and speculative decoding scenarios. The implementation includes a new task registration system and ensures compatibility with existing graph replay mechanisms across the Ascend backend. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Introduce UpdatableGraph for Fused Infer Attention on AscendSuggested PR Summary:
### What this PR does / why we need it?
This PR introduces `UpdatableGraph` to support updatable graphs for Fused Infer Attention (FIA) on Ascend NPUs, replacing the manual graph parameter updates with a task-registration and resolution system. This simplifies the graph replay path and improves maintainability.
Feedback has been provided to address several critical issues, including a signature mismatch in `use_updatable_graph`, potential `AttributeError`s when accessing uncompiled `_runnable` methods or `None` layers, unsafe `issubclass` checks, and the use of `assert` for runtime validation.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested via existing CI and dummy runs for speculative decoding.8e3d3ef to
9d9c169
Compare
c7f3e4a to
6d59749
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
1 similar comment
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
8cf4361 to
50d2574
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
8684f76 to
9c73783
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
2c82adc to
a897bdb
Compare
|
/weekly GLM-5_1-W8A8_A3_weekly
|
|
/weekly Minimax_m2.7_in198_prefix99
|
|
/weekly Qwen3_5_397B_A17B_w8a8_mxfp8_A5 |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Signed-off-by: zhiyu-wa <1959864813@qq.com>
Signed-off-by: zhiyu-wa <1959864813@qq.com>
95d2be8 to
a82894f
Compare
…r ACL Graph Replay. (vllm-project#15908)" This reverts commit daa6441. Signed-off-by: chenzeyu <2978509328@qq.com>
…aph Replay. (vllm-project#15908) **What this PR does / why we need it?** Implements the `UpdatableGraph` design discussed in vllm-project#13058. Refactors the `FIA` and `Speculative Decoding` to work with the new `UpdatableGraph` design. Manually validated on A3 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-27B with MRV1 and MRV2 Manually validated on A5 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-9B with MRV1 and MRV2 - PA: gemma4 with MRV1 and MRV2 **Does this PR introduce any user-facing change?** No. **How was this patch tested?** Temporarily validated on some model. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: zhiyu-wa <1959864813@qq.com>
|
/revert |
|
/revert |
…r ACL Graph Replay. (vllm-project#15908)" This reverts commit daa6441.
…pdates for ACL Graph Replay." (#15908) (#16409) Revert of PR #15908 (merged onto `main`). Original PR: #15908 Original author: @zhiyu-wa Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920` --- **What this PR does / why we need it?** Implements the `UpdatableGraph` design discussed in #13058. Refactors the `FIA` and `Speculative Decoding` to work with the new `UpdatableGraph` design. Manually validated on A3 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-27B with MRV1 and MRV2 Manually validated on A5 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-9B with MRV1 and MRV2 - PA: gemma4 with MRV1 and MRV2 **Does this PR introduce any user-facing change?** No. **How was this patch tested?** Temporarily validated on some model. - vLLM main: vllm-project/vllm@a97dacb Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba * 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits) Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344) [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300) [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253) [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415) [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251) [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343) [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422) [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232) [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319) [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189) [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043) [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243) [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252) [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321) [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918) [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067) [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214) [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259) ...
…aph Replay. (vllm-project#15908) **What this PR does / why we need it?** Implements the `UpdatableGraph` design discussed in vllm-project#13058. Refactors the `FIA` and `Speculative Decoding` to work with the new `UpdatableGraph` design. Manually validated on A3 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-27B with MRV1 and MRV2 Manually validated on A5 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-9B with MRV1 and MRV2 - PA: gemma4 with MRV1 and MRV2 **Does this PR introduce any user-facing change?** No. **How was this patch tested?** Temporarily validated on some model. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: zhiyu-wa <1959864813@qq.com> Signed-off-by: tianming2009 <13246728590@163.com>
…pdates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) Revert of PR vllm-project#15908 (merged onto `main`). Original PR: vllm-project#15908 Original author: @zhiyu-wa Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920` --- **What this PR does / why we need it?** Implements the `UpdatableGraph` design discussed in vllm-project#13058. Refactors the `FIA` and `Speculative Decoding` to work with the new `UpdatableGraph` design. Manually validated on A3 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-27B with MRV1 and MRV2 Manually validated on A5 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-9B with MRV1 and MRV2 - PA: gemma4 with MRV1 and MRV2 **Does this PR introduce any user-facing change?** No. **How was this patch tested?** Temporarily validated on some model. - vLLM main: vllm-project/vllm@a97dacb Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Signed-off-by: tianming2009 <13246728590@163.com>
…aph Replay. (vllm-project#15908) **What this PR does / why we need it?** Implements the `UpdatableGraph` design discussed in vllm-project#13058. Refactors the `FIA` and `Speculative Decoding` to work with the new `UpdatableGraph` design. Manually validated on A3 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-27B with MRV1 and MRV2 Manually validated on A5 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-9B with MRV1 and MRV2 - PA: gemma4 with MRV1 and MRV2 **Does this PR introduce any user-facing change?** No. **How was this patch tested?** Temporarily validated on some model. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: zhiyu-wa <1959864813@qq.com> Signed-off-by: like-0517 <ithwlike@126.com>
…pdates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) Revert of PR vllm-project#15908 (merged onto `main`). Original PR: vllm-project#15908 Original author: @zhiyu-wa Merge commit: `daa644121e1142821777e3ba60a49ab6dd354920` --- **What this PR does / why we need it?** Implements the `UpdatableGraph` design discussed in vllm-project#13058. Refactors the `FIA` and `Speculative Decoding` to work with the new `UpdatableGraph` design. Manually validated on A3 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-27B with MRV1 and MRV2 Manually validated on A5 with the following scenarios: - FIA: Qwen3-0.6B with MRV1 and MRV2 - MTP: Qwen3.5-9B with MRV1 and MRV2 - PA: gemma4 with MRV1 and MRV2 **Does this PR introduce any user-facing change?** No. **How was this patch tested?** Temporarily validated on some model. - vLLM main: vllm-project/vllm@a97dacb Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Signed-off-by: like-0517 <ithwlike@126.com>
### What this PR does / why we need it? **Depends on #15908 for Full Decode Only graph replay. Please merge after that dependency.** Two independent MRV2 Qwen3.5 PP fixes: - **Graph backend selection:** skip metadata-only attention backends whose `get_impl_cls()` raises `NotImplementedError` while collecting graph-update backends, so the PP stage can select an executable attention backend (GDN-style backends have no attention impl). - **Draft inputs:** disable external multimodal embedding collection on the speculator when PP is enabled, set right after draft loading and before profiling. The last PP stage has no encoder cache; ordinary token embedding remains in the draft model. Non-PP behavior is unchanged. Plus one four-card E2E (`test_qwen35_pp_mtp.py`) covering serial, concurrent, and chunked-prefill requests. ### Does this PR introduce _any_ user-facing change? Addresses PP + MTP backend-selection and draft-input failures for Qwen3.5 under MRV2. Full Decode Only support still relies on #15908; this PR is not a standalone replacement for that dependency. ### How was this patch tested? - **E2E (CI four-card):** `tests/e2e/pull_request/four_card/model_runner_v2/test_qwen35_pp_mtp.py` — serial answers, concurrency, and a prompt exceeding the prefill budget. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: LostFox11 <wangziyue17@huawei.com> Co-authored-by: LostFox11 <wangziyue17@huawei.com>
### What this PR does / why we need it? **Depends on vllm-project#15908 for Full Decode Only graph replay. Please merge after that dependency.** Two independent MRV2 Qwen3.5 PP fixes: - **Graph backend selection:** skip metadata-only attention backends whose `get_impl_cls()` raises `NotImplementedError` while collecting graph-update backends, so the PP stage can select an executable attention backend (GDN-style backends have no attention impl). - **Draft inputs:** disable external multimodal embedding collection on the speculator when PP is enabled, set right after draft loading and before profiling. The last PP stage has no encoder cache; ordinary token embedding remains in the draft model. Non-PP behavior is unchanged. Plus one four-card E2E (`test_qwen35_pp_mtp.py`) covering serial, concurrent, and chunked-prefill requests. ### Does this PR introduce _any_ user-facing change? Addresses PP + MTP backend-selection and draft-input failures for Qwen3.5 under MRV2. Full Decode Only support still relies on vllm-project#15908; this PR is not a standalone replacement for that dependency. ### How was this patch tested? - **E2E (CI four-card):** `tests/e2e/pull_request/four_card/model_runner_v2/test_qwen35_pp_mtp.py` — serial answers, concurrency, and a prompt exceeding the prefill budget. - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: LostFox11 <wangziyue17@huawei.com> Co-authored-by: LostFox11 <wangziyue17@huawei.com>
What this PR does / why we need it?
Implements the
UpdatableGraphdesign discussed in #13058.Refactors the
FIAandSpeculative Decodingto work with the newUpdatableGraphdesign.Manually validated on A3 with the following scenarios:
Manually validated on A5 with the following scenarios:
Does this PR introduce any user-facing change?
No.
How was this patch tested?
Temporarily validated on some model.