Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request refactors how the V2 model runner is enabled for Ascend NPU devices. Instead of relying on a simple environment variable override, the system now dynamically evaluates model architecture compatibility, Triton availability, and feature support (such as speculative decoding or parallel execution modes) to determine if the V2 runner can be safely used. This approach improves stability by preventing crashes on unsupported configurations while allowing the V2 runner to be utilized where it is fully compatible. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
This pull request refactors how the V2 model runner is enabled on Ascend NPU by removing legacy monkey-patches and implementing robust, Ascend-specific heuristics directly in vllm_ascend/platform.py. The review feedback suggests formatting the PR title and summary according to the repository style guide, using existing helper functions (is_moe_model and enable_sp) to improve configuration checks, and addressing a Python 3.9 compatibility issue with importlib.metadata.entry_points.
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
2b14fb7 to
d509585
Compare
|
cc @wangxiyuan This pr enable model runner v2 for specific models and features. These models and features have e2e tests. PTAL. |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
d62ddb3 to
1e4f81e
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
7a0edb4 to
25193fc
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
f3d8c55 to
0fc0799
Compare
7868953 to
bb2bce6
Compare
|
/rerun Rerun (failed jobs only):
|
bb2bce6 to
f839075
Compare
f839075 to
6037ed0
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Signed-off-by: zouzy <zouzongyu@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
What this PR does / why we need it?
Replace the
use_v2_model_runnermonkey-patch with an Ascend-native property onVllmConfig,applied inNPUPlatform.check_and_update_config.Does this PR introduce any user-facing change?
N/A
How was this patch tested?
CI passed with new added/existing test.
Co-authored-by: zouzy