Skip to content

[Misc][MRv2]enable model runner v2 by default via whitelists - #16203

Closed
yjyang62 wants to merge 17 commits into
vllm-project:mainfrom
yjyang62:enable-mrv2-whitelist-deae
Closed

yjyang62 wants to merge 17 commits into
vllm-project:mainfrom
yjyang62:enable-mrv2-whitelist-deae

Conversation

@yjyang62

@yjyang62 yjyang62 commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Rebase #11692 onto current main.

main currently overrides VllmConfig.use_v2_model_runner with an env-only gate: unless VLLM_USE_V2_MODEL_RUNNER is set, Ascend always uses Model Runner V1. Upstream vLLM instead defaults V2 from its own architecture / Triton / feature checks. Following those GPU defaults on NPU can crash, because the Ascend V2 runner does not yet cover the same model and feature set.

This PR replaces that env-only property with an Ascend-owned default in vllm_ascend/mrv2_utils.py, installed wherever VllmConfig is created or unpickled (frontend NPUPlatform.check_and_update_config, EngineCore via vllm_ascend/__init__.py, and worker init).

Default V2 when all of the following hold; otherwise V1:

  • Model architecture whitelist: Qwen3ForCausalLM (generate, not hybrid, not attention-free).
  • Feature whitelist: no speculative decoding, or eagle3 / mtp / dflash.
  • Platform is not 310P and Triton is available.

VLLM_USE_V2_MODEL_RUNNER=0/1 still wins over the whitelist.

Related changes in the same patch file:

  • Neutralize upstream _validate_v2_model_runner. GPU V2 feature/Triton checks must not reject an Ascend enablement decision (including an explicit env override).
  • Keep the existing V2 unsupported-feature exceptions: drop prefill context parallelism on v0.28.0, and honor resolve_spec_pp_support for spec+PP.
  • Keep the V1 unsupported-feature exceptions for dspark speculative decoding and dflash2 drafts, but install that patch with hasattr instead of vllm_version_is("0.28.0") so SHA / __version__="dev" CI checkouts do not patch a missing API.

Does this PR introduce any user-facing change?

Yes. vllm serve of Qwen3ForCausalLM (for example Qwen3 dense) now defaults to Model Runner V2 on A2/A3/A5 when Triton is present and speculative decoding is absent or eagle3/mtp/dflash. Models not on the architecture whitelist (DeepSeek, GLM, Qwen3-MoE, hybrid models, 310P) still default to V1. Explicit VLLM_USE_V2_MODEL_RUNNER is unchanged.

How was this patch tested?

  • Unit tests added/updated:

    • tests/ut/test_mrv2_utils.py (architecture / feature / env / 310P / Triton gates, validation no-op)
    • tests/ut/patch/platform/test_patch_use_v2_model_runner.py (property wiring, V1 feature exceptions, v0.28.0 PCP exception)
  • tests/e2e/pull_request/one_card/test_attention_v1_precision.py: mock TP/DP groups because Qwen3 now defaults to V2 and NPUPlatform.set_additional_forward_context reads those groups.

  • CI: cpu-ut / e2e on this PR.

  • vLLM main: vllm-project/vllm@84030bb

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request transitions the vLLM-Ascend platform from a simple environment-variable-based toggle for the V2 model runner to a more robust, whitelist-driven mechanism. By centralizing the enablement logic, the platform now intelligently selects the V2 runner for supported architectures and features while maintaining safe fallbacks to V1. This ensures consistent configuration behavior across distributed processes and improves maintainability by decoupling the enablement decision from upstream vLLM defaults.

Highlights

  • Centralized V2 Logic: Introduced a new utility module, mrv2_utils.py, to encapsulate logic for V2 model runner enablement, replacing scattered monkey-patches.
  • Whitelist-Driven Enablement: Implemented a whitelist-based approach for model architectures (e.g., Qwen3) and features (e.g., eagle/mtp/dflash) to determine V2 runner eligibility, with a fallback to V1 when requirements like Triton are not met.
  • Consistent Patch Application: Ensured the V2 model runner configuration patch is consistently applied across frontend, engine, and worker processes by integrating it into platform hooks and worker initialization.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:\n\nmarkdown\n[Platform][Feature] Implement Ascend-owned whitelist heuristics for Model Runner V2 enablement\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis PR introduces Ascend-owned whitelist heuristics for enabling the Model Runner V2 by default. Instead of relying purely on the 'VLLM_USE_V2_MODEL_RUNNER' environment variable, it now enables V2 by default for whitelisted architectures (currently 'Qwen3ForCausalLM') and supported speculative decoding methods ('eagle3', 'mtp', 'dflash'), provided Triton is available and the platform is not 310P. It also decouples and neutralizes the upstream V2 validation checks which do not apply to Ascend.\n\nFeedback: Two potential 'TypeError' issues were identified in 'vllm_ascend/mrv2_utils.py' where 'architectures' could be 'None' if explicitly configured as such. It is recommended to use 'or []' to guarantee an iterable list.\n\n### Does this PR introduce _any_ user-facing change?\nYes, Model Runner V2 is now enabled by default for Qwen3 models on supported Ascend platforms (non-310P) with Triton, whereas previously it required explicitly setting 'VLLM_USE_V2_MODEL_RUNNER=1'.\n\n### How was this patch tested?\nThe changes are covered by new unit tests in 'tests/ut/test_mrv2_utils.py' and 'tests/ut/patch/platform/test_patch_use_v2_model_runner.py', as well as updates to existing end-to-end tests.\n

Comment thread vllm_ascend/mrv2_utils.py
if getattr(model_config, "is_attention_free", False):
return False

architectures = getattr(model_config, "architectures", [])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

If architectures is explicitly set to None in the model configuration, getattr(model_config, "architectures", []) will return None. This will cause a TypeError when iterating over it in the subsequent any(...) expression. Using or [] ensures that architectures is always an iterable list.

Suggested change
architectures = getattr(model_config, "architectures", [])
architectures = getattr(model_config, "architectures", []) or []

Comment thread vllm_ascend/mrv2_utils.py

if is_default_v2_model_runner_model(vllm_config):
if _v2_model_runner_environment_ready(vllm_config):
architectures = getattr(vllm_config.model_config, "architectures", [])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

If architectures is explicitly set to None in the model configuration, getattr(vllm_config.model_config, "architectures", []) will return None. This will cause a TypeError when attempting to join the elements with ", ".join(architectures). Using or [] ensures that architectures is always an iterable list.

Suggested change
architectures = getattr(vllm_config.model_config, "architectures", [])
architectures = getattr(vllm_config.model_config, "architectures", []) or []

@yjyang62 yjyang62 added the ready-all run all e2e test for pr label Sep 10, 2026
@yjyang62
yjyang62 force-pushed the enable-mrv2-whitelist-deae branch from c27a1b5 to 4261b44 Compare September 10, 2026 09:38
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch from 4b32ac4 to bc07dbe Compare September 14, 2026 03:11
@yjyang62
yjyang62 force-pushed the enable-mrv2-whitelist-deae branch from bc07dbe to 0480dbd Compare September 14, 2026 06:54
@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch from 7d8e38b to 5e04d59 Compare September 15, 2026 01:41
@cursor
cursor Bot requested a review from weijinqian0 as a code owner September 15, 2026 04:11
@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch from ba90e6a to fd9007a Compare September 15, 2026 06:24
@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch from a2e172f to fd9007a Compare September 15, 2026 09:13
@yjyang62
yjyang62 force-pushed the enable-mrv2-whitelist-deae branch 2 times, most recently from 1525bae to aa9b367 Compare September 15, 2026 11:43
@yjyang62 yjyang62 changed the title [Misc][Platform]enable model runner v2 by default via whitelists [Misc][MRv2]enable model runner v2 by default via whitelists Sep 15, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch from d4fd328 to da0c16d Compare September 16, 2026 04:03
yjyang62 and others added 13 commits September 16, 2026 04:41
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Qwen3ForCausalLM now defaults to MRv2. The RLHF sleep/wake suite still
depends on the V1 generate path after CuMem remap, so keep the subprocess
server on VLLM_USE_V2_MODEL_RUNNER=0.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 stores capturing on the forward-context object. Ascend FIA treats
_EXTRA_CTX.capturing as ACL graph capture. When Qwen3 defaults to V2 via
the whitelist (env unset), extras leaked onto ctx.capturing and
graph_task_group_begin ran on a non-capturing stream (error 107029).

Route extras through additional_kwargs whenever VllmConfig.use_v2_model_runner
is true, not only when VLLM_USE_V2_MODEL_RUNNER is set.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Whitelist-enabled V2 routed extras through additional_kwargs whenever
ctx.vllm_config.use_v2_model_runner was truthy. cpu-ut fixtures pass a
bare MagicMock forward context, which auto-creates a truthy flag and
hides capturing / max_tokens_across_dp. Require an actual bool True.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 sets forward_context.capturing during piecewise warmup. Ascend FIA
treated that as ACL capture and called graph_task_group_begin, which
failed with 107029 and hung Qwen3 whitelist-MRv2 e2e.

Use the live NPU stream capture state (and skip PIECEWISE) instead of
_EXTRA_CTX.capturing when wrapping FIA/PA kernels.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cpu-ut stubs torch.npu as MagicMock, so is_current_stream_capturing()
is truthy and FIA entered full_graph_fia. Require an actual True,
matching use_v2_model_runner.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
A2 CI hung on LoRA-only generate and failed dflash acceptance when
batch-size-based dynamic K (num_speculative_tokens_per_batch_size) was
enabled under the Qwen3 default-V2 whitelist. Drop both from the default
feature whitelist; static eagle3/mtp/dflash stay enabled, and
VLLM_USE_V2_MODEL_RUNNER still overrides.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 PrefetchOffloader calls torch.cuda.is_current_stream_capturing
during load_model. V1 already remaps that CUDA dummy to torch.npu;
V2 torch_cuda_wrapper did not, so Qwen3 default-V2 prefetch e2e
crashed on NPU.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Qwen3ForCausalLM now defaults to Model Runner V2. GPU V2 prefetch plus
NZ graph capture matches the eager baseline, so the strict xfail on
test_prefetch_offload_accuracy[NZ-graph] XPASS-fails a2-1 CI.
Keep the V1 AscendPrefetchOffloader fail-fast for the unsupported combo.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Route _EXTRA_CTX through additional_kwargs when
use_v2_model_runner(get_current_vllm_config()) is true. Whitelist-default
V2 leaves VLLM_USE_V2_MODEL_RUNNER unset and GPU ForwardContext has no
vllm_config, so env/ctx checks leaked GPU capturing onto FIA.

Restore FIA/PA graph_task_group gating to _EXTRA_CTX.capturing.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Replace `_extra_ctx_uses_additional_kwargs` with
`if use_v2_model_runner(get_current_vllm_config()) is True`.

Guard xlite `index_full_mask` for mypy after merging main.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep getattr/setattr the same as before, only replacing the helper
with `if use_v2_model_runner(get_current_vllm_config()) is True`.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep getattr/setattr the same as main. Only replace the V2 condition
with `use_v2_model_runner(get_current_vllm_config()) is True`.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch from da0c16d to 3538c10 Compare September 16, 2026 04:44
yjyang62 and others added 4 commits September 16, 2026 04:58
get_current_vllm_config() raises in cpu-ut attention fixtures.
Fall back to V1 extra-ctx attrs so FIA can read capturing.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Compiled FIA/MoE read _EXTRA_CTX, which called use_v2_model_runner()
and hit logger.warning_once / info_once. Dynamo cannot trace those
logs, so LoRA and non-whitelist models failed compile, and V2 MoE
baked additional_kwargs.get("moe_comm_method") as None.

Disable Dynamo on the extras helper, proxy getattr/setattr, and
use_v2_model_runner so isolation still uses get_current_vllm_config
at runtime.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Decorating __getattr__/__setattr__ with torch._dynamo.disable made
mypy treat _EXTRA_CTX as having no dynamic attributes. Keep the
dunders undecorated and disable the helpers they call instead.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cpu-ut installs torch 2.10, which marks @torch._dynamo.disable
with _torchdynamo_disable instead of _dynamo_disable.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
ningjingbengxiaohai pushed a commit that referenced this pull request Sep 17, 2026
…park (#16626)

### What this PR does / why we need it?

Stacked on #16203.

#16203 currently defaults Model Runner V2 only for `Qwen3ForCausalLM`
plus static `eagle3` / `mtp` / `dflash`. This PR only extends those two
whitelists so more architectures and `dspark` can default to V2. Gating
logic is otherwise unchanged: LoRA and dynamic speculative decoding
(`num_speculative_tokens_per_batch_size`) still stay on V1, and
`VLLM_USE_V2_MODEL_RUNNER` still overrides the whitelist.

Default-V2 model architectures:

- `Qwen3ForCausalLM`
- `Qwen3MoeForCausalLM`
- `MiniMaxM2ForCausalLM`
- `DeepseekV3ForCausalLM`
- `DeepseekV32ForCausalLM`
- `GlmMoeDsaForCausalLM`
- `DeepseekV4ForCausalLM`

Default-V2 speculative methods: `eagle3` / `mtp` / `dflash` / `dspark`.

### Does this PR introduce _any_ user-facing change?

Yes. The architectures above now default to Model Runner V2 when Triton
is present, the platform is not 310P, and the feature whitelist is
satisfied. Other models still default to V1. Explicit
`VLLM_USE_V2_MODEL_RUNNER=0/1` is unchanged.

### How was this patch tested?

- Unit tests updated: `tests/ut/test_mrv2_utils.py` (architecture
parametrize, `dspark` on the feature whitelist)
- CI: cpu-ut / e2e on this PR.

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Spicy-Stick <873805887@qq.com>
Signed-off-by: ZhangwenTaoHW <zhangwentao101@huawei.com>
Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Co-authored-by: Spicy-Stick <873805887@qq.com>
Co-authored-by: ZhangwenTaoHW <zhangwentao101@huawei.com>
Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.


idx = 0
index_mask = self.xlite_config.index_full_mask or [True] * self.xlite_config.n_layers
index_mask = getattr(self.xlite_config, "index_full_mask", None) or [True] * self.xlite_config.n_layers

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems unnecessary. Update xlite with pip install xlite==0.2.0rc1 and the lint issue should be gone.

@yjyang62 yjyang62 closed this Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants