Skip to content

[Misc][Platform]enable model runner v2 by default via whitelists - #11692

Closed
wxsIcey wants to merge 19 commits into
vllm-project:mainfrom
wxsIcey:wxs/mrv2/enable-v2
Closed

wxsIcey wants to merge 19 commits into
vllm-project:mainfrom
wxsIcey:wxs/mrv2/enable-v2

Conversation

@wxsIcey

@wxsIcey wxsIcey commented Jul 9, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Replace the use_v2_model_runner monkey-patch with an Ascend-native property on VllmConfig, applied in NPUPlatform.check_and_update_config.

  • Model whitelist: default-V2 architectures (Qwen3ForCausalLM) .
  • Feature whitelist: whitelisted models may stack eagle/mtp/dflash spec decoding.
  • V2 requires Triton; fall back to V1 otherwise.

Does this PR introduce any user-facing change?

N/A

How was this patch tested?

CI passed with new added/existing test.

Co-authored-by: zouzy

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request refactors how the V2 model runner is enabled for Ascend NPU devices. Instead of relying on a simple environment variable override, the system now dynamically evaluates model architecture compatibility, Triton availability, and feature support (such as speculative decoding or parallel execution modes) to determine if the V2 runner can be safely used. This approach improves stability by preventing crashes on unsupported configurations while allowing the V2 runner to be utilized where it is fully compatible.

Highlights

  • Model Runner V2 Enablement: Replaced the static monkey-patch for V2 model runner enablement with a dynamic, feature-aware logic implementation in the NPU platform configuration.
  • Feature Compatibility Checks: Introduced comprehensive compatibility checks to verify if the V2 model runner supports specific model architectures and features, falling back to V1 when necessary.
  • Cleanup of Legacy Patches: Removed redundant monkey-patch files and initialization imports that were previously used to force-enable or gate the V2 model runner.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors how the V2 model runner is enabled on Ascend NPU by removing legacy monkey-patches and implementing robust, Ascend-specific heuristics directly in vllm_ascend/platform.py. The review feedback suggests formatting the PR title and summary according to the repository style guide, using existing helper functions (is_moe_model and enable_sp) to improve configuration checks, and addressing a Python 3.9 compatibility issue with importlib.metadata.entry_points.

Comment thread vllm_ascend/platform.py Outdated
Comment thread vllm_ascend/platform.py Outdated
Comment thread vllm_ascend/platform.py Outdated
Comment thread vllm_ascend/platform.py Outdated
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@wxsIcey
wxsIcey force-pushed the wxs/mrv2/enable-v2 branch from 2b14fb7 to d509585 Compare August 25, 2026 07:55
@wxsIcey wxsIcey changed the title [Misc] enable model runner v2 [platform] enable model runner v2 by default via whitelists Aug 26, 2026
@wxsIcey
wxsIcey marked this pull request as ready for review August 26, 2026 02:51
@wxsIcey wxsIcey changed the title [platform] enable model runner v2 by default via whitelists [Misc] [Platform] enable model runner v2 by default via whitelists Aug 26, 2026
@wxsIcey wxsIcey changed the title [Misc] [Platform] enable model runner v2 by default via whitelists [Misc][Platform]enable model runner v2 by default via whitelists Aug 26, 2026
@wxsIcey

wxsIcey commented Aug 26, 2026 •

Copy link
Copy Markdown
Collaborator Author

cc @wangxiyuan This pr enable model runner v2 for specific models and features. These models and features have e2e tests. PTAL.

Comment thread vllm_ascend/platform.py Outdated
Comment thread vllm_ascend/platform.py Outdated
Comment thread vllm_ascend/platform.py Outdated
@wxsIcey wxsIcey added the ready-precise run selected e2e test for pr label Aug 26, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@wxsIcey wxsIcey added ready-all run all e2e test for pr and removed ready-precise run selected e2e test for pr labels Sep 1, 2026
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

root and others added 5 commits September 4, 2026 00:01
Signed-off-by: root <root@vllm-ascend.local>
Signed-off-by: root <root@vllm-ascend.local>
Signed-off-by: root <root@vllm-ascend.local>
Signed-off-by: root <root@vllm-ascend.local>
Signed-off-by: zouzy <zouzongyu@huawei.com>
@zouzy5137
zouzy5137 force-pushed the wxs/mrv2/enable-v2 branch 2 times, most recently from 7868953 to bb2bce6 Compare September 4, 2026 01:15
@wxsIcey

wxsIcey commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

Signed-off-by: zouzy <zouzongyu@huawei.com>
@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: zouzy <zouzongyu@huawei.com>
@wxsIcey wxsIcey closed this Sep 10, 2026
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 10, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 10, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 10, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 14, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 14, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 15, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 15, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 15, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 16, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 16, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 17, 2026
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants