[Feature][MRV2] expand default MRv2 architecture whitelist and add dspark - #16626
ningjingbengxiaohai merged 37 commits into
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request expands the default Model Runner V2 (MRv2) support on Ascend platforms by updating the architecture and speculative method whitelists. It introduces a centralized utility module to manage MRv2 enablement, which decouples the logic from upstream validation and provides a more robust mechanism for applying configuration patches. Additionally, the PR improves the detection of ACL graph capturing to ensure compatibility with the V2 runner, preventing issues where GPU-specific flags were incorrectly interpreted in the NPU environment. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Misc][Feature] Implement Ascend-owned whitelist heuristics for Model Runner V2 enablementSuggested PR Summary:
### What this PR does / why we need it?
This PR implements Ascend-owned whitelist heuristics for Model Runner V2 (MRv2) enablement instead of relying on the upstream GPU-specific defaults, which can cause crashes on unsupported configurations. It introduces `vllm_ascend/mrv2_utils.py` to manage the whitelist of architectures (e.g., Qwen3, DeepseekV3) and speculative decoding methods (e.g., eagle3, mtp, dflash, dspark), ensuring MRv2 is only enabled when the environment is ready (non-310P and Triton available).
Additionally, it resolves an issue where GPU V2's `capturing` flag was incorrectly treated as ACL stream capture, which caused hangs. It introduces `is_acl_full_graph_capturing()` to verify actual NPU stream capture status.
Feedback on the code changes:
- Safely retrieve `model_config` and `speculative_config` using `getattr` to prevent potential `AttributeError`s.
- Add a warning log when Model Runner V2 is disabled due to an unsupported speculative decoding method to improve debuggability.
### Does this PR introduce _any_ user-facing change?
Yes, Model Runner V2 will now be enabled by default for whitelisted architectures (such as Qwen3) and supported speculative decoding methods on compatible Ascend platforms.
### How was this patch tested?
The changes are covered by new unit tests in `tests/ut/test_mrv2_utils.py` and updates to existing tests in `tests/ut/test_ascend_forward_context.py` and `tests/e2e/`.|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
1 similar comment
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
4fdff6c to
ae37a4e
Compare
eec9224 to
1355a74
Compare
|
/nightly qwen3-30b-a3b-w8a8 kimi25_w4a8_step_3_7 kimi-k2.6_max_model_len acceptace_rate_deepseekv4-flash-w8a8-mtp Qwen3-235B-A22B-W8A8_piecewise_fullgraph_A3 Qwen3_32b_w8a8_ml_A3 qwen3-235b-a22b-w8a8 glm-5.2-w4a8-mtp glm-5.2-w4a8-dspark glm-5.2-w4a8c8-sfa-dcp compressor-metadata-cross-stream Qwen3-32B-QuaRot Qwen3_32b_W8A8_wl_A3 qwen3-32b-int8 qwen3-32b-int8-prefix-cache Qwen3-30B-A3B-W4A8-llm-compressor Qwen3-30B-QuaRot qwen3-30b-acc Minimax_m2.7_w8a8_A3 deepseek-v3-2-w8a8 Deepseek_R1_W8A8_Reasoning_output_A3 deepseek-r1-0528-w8a8-prefix-cache mtpx-deepseek-r1-0528-w8a8 --a3-560t
|
|
/nightly kimi-k2-thinking Qwq_32b_A3 MiniMax-M3-BF16-A3 Qwen3.8-27B-w8a8-A3 qwen_3_6_27b_w8a8_A3 Qwen3.5-397B-A17B-w8a8-mtp Qwen3.5-27B-w8a8-A3 Qwen3.5-122B-A10B-W8A8-A3 qwen3-vl-235b-a22b-instruct-w8a8 Qwen3-32B-W8A8C8-A3 MiniMax-M3-W8A8-A3 Kimi-K2.6-w4a8-A3 kimi-k2.5 glm-5.1-w8a8-prefill-mc2 glm-4.7-w8a8 DeepSeek-V4-Flash-W8A8-A3 Deepseek_R1_W8A8_A3 rejection-sample gemma4-31b-dense gemma4 qwen3-30b-a3b-w8a8-a2-performance qwen3-32b-int8 Qwen3-Next-80B-A3B-Instruct Qwen3-ASR-1.7B Qwen3-8B qwen3-30b-a3b-bf16-a2-performance multi-node-qwen3-235b-dp Qwen3.5-397B-A17B-w4a8-mtp Qwen3.5-27B-w8a8-A2 qwen3-vl-32b-instruct-w8a8 MiniMax-M2.5-w8a8-QuaRot-A2 multi-node-Kimi-K2.5-W4A8-A2 multi-node-GLM-5.1-w8a8-A2 --a3-560t
|
|
/nightly qwen3-30b-a3b-w8a8 kimi25_w4a8_step_3_7 kimi-k2.6_max_model_len acceptace_rate_deepseekv4-flash-w8a8-mtp Qwen3-235B-A22B-W8A8_piecewise_fullgraph_A3 Qwen3_32b_w8a8_ml_A3 qwen3-235b-a22b-w8a8 glm-5.2-w4a8-mtp glm-5.2-w4a8-dspark glm-5.2-w4a8c8-sfa-dcp compressor-metadata-cross-stream Qwen3-32B-QuaRot Qwen3_32b_W8A8_wl_A3 qwen3-32b-int8 qwen3-32b-int8-prefix-cache Qwen3-30B-A3B-W4A8-llm-compressor Qwen3-30B-QuaRot qwen3-30b-acc Minimax_m2.7_w8a8_A3 deepseek-v3-2-w8a8 Deepseek_R1_W8A8_Reasoning_output_A3 deepseek-r1-0528-w8a8-prefix-cache mtpx-deepseek-r1-0528-w8a8 --a3-560t
|
|
/nightly kimi-k2-thinking Qwq_32b_A3 MiniMax-M3-BF16-A3 Qwen3.8-27B-w8a8-A3 qwen_3_6_27b_w8a8_A3 Qwen3.5-397B-A17B-w8a8-mtp Qwen3.5-27B-w8a8-A3 Qwen3.5-122B-A10B-W8A8-A3 qwen3-vl-235b-a22b-instruct-w8a8 Qwen3-32B-W8A8C8-A3 MiniMax-M3-W8A8-A3 Kimi-K2.6-w4a8-A3 kimi-k2.5 glm-5.1-w8a8-prefill-mc2 glm-4.7-w8a8 DeepSeek-V4-Flash-W8A8-A3 Deepseek_R1_W8A8_A3 rejection-sample gemma4-31b-dense gemma4 qwen3-30b-a3b-w8a8-a2-performance qwen3-32b-int8 Qwen3-Next-80B-A3B-Instruct Qwen3-ASR-1.7B Qwen3-8B qwen3-30b-a3b-bf16-a2-performance multi-node-qwen3-235b-dp Qwen3.5-397B-A17B-w4a8-mtp Qwen3.5-27B-w8a8-A2 qwen3-vl-32b-instruct-w8a8 MiniMax-M2.5-w8a8-QuaRot-A2 multi-node-Kimi-K2.5-W4A8-A2 multi-node-GLM-5.1-w8a8-A2 --a3-560t
|
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
ba671cc to
ae0186a
Compare
|
/nightly QWEN3_235B_PD DeepSeek-V4-Pro-w4a8-1M-PD multi-node-deepseek-v3.2-W8A8-EP QWEN3_235B_PD_3_5K_1_5k multi-node-dpsk3.2-2node multi-node-GLM-5.1-w8a8-A3 multi-node-GLM-5.2-w8a8-A3 multi-node-qwenw8a8-2node-eplb multi-node-GLM-5.1-W8A8C8-A3_128k_90_50 multi-node-GLM-5.1-W8A8C8-MTP-A3_198k_function multi-node-deepseek-v3.1 Minimax_m2.7_in128k_1k_prefix90_tpot50 DeepSeek-V4-flash-w8a8-PD-prefix |
Signed-off-by: ZhangwenTaoHW <zhangwentao101@huawei.com>
Pass optional reasoning_effort through AISBench chat request generation. Set reasoning_effort: low for the DeepSeek-V4 Flash W8A8 A3 nightly GPQA case. Keep thinking: true and the performance case unchanged. Cherry-picked from vllm-project#16721 Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
29c01ae to
3c33450
Compare
Qwen3_5ForConditionalGeneration hybrid VL hits MRv2 encoder graph capture (CUDA stream assert) and hybrid KV copy. Keep it on V1 by default. Pin dflash2 PIECEWISE acceptance to V1 after vllm-project#16726 broke dummy propose without set_forward_context; V2 eager remains covered. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
…telist and add dspark (vllm-project#16626)" This reverts commit c7ca0b6. Signed-off-by: chenzeyu <2978509328@qq.com>
Keep DeepSeek V4.1 forward-context moe_comm_methods from main and drop the reverted _USE_V2_EXTRA_KWARGS assertion from the unit test. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
…telist and add dspark (vllm-project#16626)" This reverts commit c7ca0b6. Signed-off-by: chenzeyu <2978509328@qq.com>
…d add dspark" (#16832) Reverts #16626 - vLLM main: vllm-project/vllm@84030bb Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
vllm-project#16544) Reverts commit 200309d on main (cherry-picked from main_verify 82145c7). Conflict resolution: ascend_forward_context.py, patch/__init__.py, worker.py and test_ascend_forward_context.py were entangled with vllm-project#16626 (merged before vllm-project#16544, already reverted on this branch); they are restored to the pre-vllm-project#16626 state (c7ca0b6~1), the correct composition of both reverts. dsa_cp.py and test_model_runner_v2.py keep later upstream changes (vllm-project#16168 metadata buffer sizing fix, vllm-project#15747 spec-pp protocol rename). All other surviving files match the pre-PR state. Signed-off-by: chenzeyu <2978509328@qq.com>
Reapply the code removed by the upstream revert of vllm-project#16626 while retaining the newer blacklist-based runner selection from vllm-project#16848. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and vllm-project#15514 --kv-cache-dtype/--indexer_kv_dtype fp8 without the retired enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Restore the optional-mask access used in upstream vllm-project#16626 to resolve the shared mypy failure tracked in vllm-project#16658. Preserve configured cache routing and the all-layer fallback. Five isolated cache-routing cases pass. Assisted-by: OpenAI Codex Signed-off-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com>
Restore the optional-mask access used in upstream vllm-project#16626 to resolve the shared mypy failure tracked in vllm-project#16658. Preserve configured cache routing and the all-layer fallback. Five isolated cache-routing cases pass. Assisted-by: OpenAI Codex Signed-off-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and enable_sparse_*_c8 additional-config keys. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Sage Martin <lethamannodaiu@outlook.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
What this PR does / why we need it?
This PR expands the default Model Runner V2 model and feature whitelists so more architectures, plus dspark, default to V2.
A request uses Model Runner V2 by default only when all of the following hold:
The model architecture is on the default-V2 whitelist:
Enabled features are on the default-V2 feature whitelist. Static eagle3 / mtp / dflash / dspark are included. LoRA and dynamic speculative decoding (num_speculative_tokens_per_batch_size) stay on V1.
Triton is available.
Otherwise the request stays on Model Runner V1. Explicit VLLM_USE_V2_MODEL_RUNNER=0/1 still overrides the whitelist.
Does this PR introduce any user-facing change?
Yes. The architectures above now default to Model Runner V2 when the conditions above are satisfied. Other models still default to V1. Explicit VLLM_USE_V2_MODEL_RUNNER=0/1 is unchanged.
How was this patch tested?
Unit tests updated: tests/ut/test_mrv2_utils.py (architecture parametrize, dspark on the feature whitelist)
CI: cpu-ut / e2e on this PR.