Skip to content

[Feature][MRV2] expand default MRv2 architecture whitelist and add dspark - #148

Draft
yjyang62 wants to merge 37 commits into
mainfrom
cursor/expand-mrv2-whitelist-9095
Draft

yjyang62 wants to merge 37 commits into
mainfrom
cursor/expand-mrv2-whitelist-9095

Conversation

@yjyang62

@yjyang62 yjyang62 commented Sep 15, 2026 •

Copy link
Copy Markdown
Owner

What this PR does / why we need it?

Stacked on vllm-project#16203.

vllm-project#16203 introduces Ascend-owned default Model Runner V2 selection in vllm_ascend/mrv2_utils.py (model + feature whitelists, env override, upstream validation no-op). This PR expands that default path to more production models and dspark, and adjusts related platform / spec-decode / CI coverage so default-V2 stays safe on NPU.

Scope (范围)

  • Default-V2 model whitelist (DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES in vllm_ascend/mrv2_utils.py): add Qwen3MoeForCausalLM, MiniMaxM2ForCausalLM, DeepseekV3ForCausalLM, DeepseekV32ForCausalLM, GlmMoeDsaForCausalLM, DeepseekV4ForCausalLM, and Qwen3_5MoeForCausalLM on top of [Misc][MRv2]enable model runner v2 by default via whitelists vllm-project/vllm-ascend#16203’s Qwen3ForCausalLM.
  • Explicit default-V1 model exclusion: Qwen3_5ForConditionalGeneration (Qwen3.5/3.6 dense and VL hybrid) is not on the whitelist. It stays on V1 by default because MRv2 hits hybrid VL encoder-graph capture and hybrid KV copy-with-prefix-cache paths that are not ready for default enablement.
  • Default-V2 feature whitelist: add static dspark alongside eagle3 / mtp / dflash. Still default V1 when LoRA is enabled, when num_speculative_tokens_per_batch_size (dynamic spec decode) is set, or when DSpark KV sliding window (additional_config.draft_window_size) is set.
  • Hybrid on whitelist: for architectures in the list above, is_hybrid=True does not force V1 (e.g. Qwen3_5MoeForCausalLM). Attention-free models and non-generate runner types (except runner_type="draft", which inherits the target’s V2 decision) remain V1.
  • Spec-decode / worker alignment: keep MTP draft workers on the same default MRv2 path as the target model; blacklist DSpark sliding window from default V2; pin selected e2e cases that still require V1 (EPLB V1-only paths, p-eagle parallel drafting, SFA V1 precision, test_dflash2_acceptance PIECEWISE after [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (#16425) vllm-project/vllm-ascend#16726, and similar) via VLLM_USE_V2_MODEL_RUNNER=0 while moving validated nightly cases to V2 where appropriate.
  • CI / tooling: extend tests/ut/test_mrv2_utils.py; refresh or pin e2e goldens for default-V2 behavior; nightly DeepSeek-V4 Flash GPQA uses reasoning_effort: low in aisbench config ([CI] Pin DeepSeek-V4 Flash GPQA nightly to low reasoning effort vllm-project/vllm-ascend#16721).

Out of scope for the whitelist change itself: general MRv2 feature development on main that is not required to land the expanded defaults (reviewers should still treat this branch as stacked on vllm-project#16203 for the core gating logic).

MRv2 switching scheme (切换方案)

Decision order in use_v2_model_runner() (vllm_ascend/mrv2_utils.py), installed on VllmConfig by apply_v2_model_runner_config_patch() (vllm-project#16203):

  1. Env override (highest priority): if VLLM_USE_V2_MODEL_RUNNER is 0 or 1, use that value (explicit opt-out or opt-in for any model/feature).
  2. Model gate: default V2 only when is_default_v2_model_runner_model() is true:
    • runner_type == "generate" and architecture ∈ whitelist, or runner_type == "draft" (inherit target whitelist; do not re-check draft architecture).
    • Reject: is_attention_free, embedding/other non-generate runners, architectures not in the whitelist (including Qwen3_5ForConditionalGeneration).
  3. Feature gate: _v2_model_runner_environment_ready() requires is_supported_v2_model_runner_feature():
    • Allow: no speculative config, or method ∈ {eagle3, mtp, dflash, dspark} without the exclusions below.
    • Deny (stay V1): LoRA; num_speculative_tokens_per_batch_size; DSpark with draft_window_size; any other speculative method.
  4. Platform gate: not Ascend 310P; HAS_TRITON must be true.
  5. Otherwise: default Model Runner V1.

Default-V2 model architectures

  • Qwen3ForCausalLM
  • Qwen3MoeForCausalLM
  • MiniMaxM2ForCausalLM
  • DeepseekV3ForCausalLM
  • DeepseekV32ForCausalLM
  • GlmMoeDsaForCausalLM
  • DeepseekV4ForCausalLM
  • Qwen3_5MoeForCausalLM

Default-V2 speculative methods: eagle3, mtp, dflash, dspark (subject to the feature exclusions above).

Does this PR introduce any user-facing change?

Yes. The architectures above now default to Model Runner V2 on A2/A3/A5 when Triton is present, the platform is not 310P, and the feature whitelist passes. Whitelisted hybrid variants (for example Qwen3.5 MoE) also default to V2; Qwen3_5ForConditionalGeneration dense/VL stays on V1. LoRA, dynamic speculative decoding, and DSpark sliding window stay on V1 unless VLLM_USE_V2_MODEL_RUNNER=1. All other models remain V1 by default. Explicit VLLM_USE_V2_MODEL_RUNNER=0/1 behavior is unchanged.

How was this patch tested?

  • Unit tests added/updated:

    • tests/ut/test_mrv2_utils.py (expanded architecture list, hybrid whitelisted models, Qwen3_5ForConditionalGeneration excluded, dspark and DSpark sliding-window / LoRA / dynamic-spec gates, draft runner inheritance)
    • tests/ut/tools/test_aisbench.py (reasoning_effort passthrough) where included on this branch
  • E2e: selective VLLM_USE_V2_MODEL_RUNNER pins and updated goldens for default-V2 paths; test_dflash2_v2_acceptance still exercises V2 eager while test_dflash2_acceptance stays on V1 for PIECEWISE ACL graph.

  • CI: cpu-ut / NPU e2e on this PR.

  • vLLM main: vllm-project/vllm@84030bb

Open in Web Open in Cursor 

@cursor
cursor Bot changed the base branch from enable-mrv2-whitelist-deae to main September 16, 2026 02:53
@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch from 4fdff6c to ae37a4e Compare September 16, 2026 04:05
@cursor
cursor Bot changed the base branch from main to enable-mrv2-whitelist-deae September 16, 2026 04:05
@cursor
cursor Bot force-pushed the enable-mrv2-whitelist-deae branch from da0c16d to 3538c10 Compare September 16, 2026 04:44
@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch from ae37a4e to df01f7c Compare September 16, 2026 04:53
@cursor
cursor Bot changed the base branch from enable-mrv2-whitelist-deae to main September 16, 2026 04:53
@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch 3 times, most recently from ba671cc to ae0186a Compare September 16, 2026 12:58
@cursor cursor Bot changed the title [Feat][Platform] expand default MRv2 architecture whitelist and add dspark [Feature][MRV2] expand default MRv2 architecture whitelist and add dspark Sep 17, 2026
yjyang62 and others added 20 commits September 17, 2026 13:52
Rebase vllm-project#11692 onto
current vllm-project/vllm-ascend main.

Replace the env-only use_v2_model_runner override with Ascend-owned
whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton;
non-310P). Keep later unsupported-feature patches for spec-PP and
Ascend-supported V1 features (dspark/dflash2). Explicit
VLLM_USE_V2_MODEL_RUNNER still wins when set.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Qwen3ForCausalLM now defaults to MRv2. The RLHF sleep/wake suite still
depends on the V1 generate path after CuMem remap, so keep the subprocess
server on VLLM_USE_V2_MODEL_RUNNER=0.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 stores capturing on the forward-context object. Ascend FIA treats
_EXTRA_CTX.capturing as ACL graph capture. When Qwen3 defaults to V2 via
the whitelist (env unset), extras leaked onto ctx.capturing and
graph_task_group_begin ran on a non-capturing stream (error 107029).

Route extras through additional_kwargs whenever VllmConfig.use_v2_model_runner
is true, not only when VLLM_USE_V2_MODEL_RUNNER is set.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Whitelist-enabled V2 routed extras through additional_kwargs whenever
ctx.vllm_config.use_v2_model_runner was truthy. cpu-ut fixtures pass a
bare MagicMock forward context, which auto-creates a truthy flag and
hides capturing / max_tokens_across_dp. Require an actual bool True.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 sets forward_context.capturing during piecewise warmup. Ascend FIA
treated that as ACL capture and called graph_task_group_begin, which
failed with 107029 and hung Qwen3 whitelist-MRv2 e2e.

Use the live NPU stream capture state (and skip PIECEWISE) instead of
_EXTRA_CTX.capturing when wrapping FIA/PA kernels.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cpu-ut stubs torch.npu as MagicMock, so is_current_stream_capturing()
is truthy and FIA entered full_graph_fia. Require an actual True,
matching use_v2_model_runner.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
A2 CI hung on LoRA-only generate and failed dflash acceptance when
batch-size-based dynamic K (num_speculative_tokens_per_batch_size) was
enabled under the Qwen3 default-V2 whitelist. Drop both from the default
feature whitelist; static eagle3/mtp/dflash stay enabled, and
VLLM_USE_V2_MODEL_RUNNER still overrides.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 PrefetchOffloader calls torch.cuda.is_current_stream_capturing
during load_model. V1 already remaps that CUDA dummy to torch.npu;
V2 torch_cuda_wrapper did not, so Qwen3 default-V2 prefetch e2e
crashed on NPU.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Qwen3ForCausalLM now defaults to Model Runner V2. GPU V2 prefetch plus
NZ graph capture matches the eager baseline, so the strict xfail on
test_prefetch_offload_accuracy[NZ-graph] XPASS-fails a2-1 CI.
Keep the V1 AscendPrefetchOffloader fail-fast for the unsupported combo.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Route _EXTRA_CTX through additional_kwargs when
use_v2_model_runner(get_current_vllm_config()) is true. Whitelist-default
V2 leaves VLLM_USE_V2_MODEL_RUNNER unset and GPU ForwardContext has no
vllm_config, so env/ctx checks leaked GPU capturing onto FIA.

Restore FIA/PA graph_task_group gating to _EXTRA_CTX.capturing.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Replace `_extra_ctx_uses_additional_kwargs` with
`if use_v2_model_runner(get_current_vllm_config()) is True`.

Guard xlite `index_full_mask` for mypy after merging main.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep getattr/setattr the same as before, only replacing the helper
with `if use_v2_model_runner(get_current_vllm_config()) is True`.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep getattr/setattr the same as main. Only replace the V2 condition
with `use_v2_model_runner(get_current_vllm_config()) is True`.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
get_current_vllm_config() raises in cpu-ut attention fixtures.
Fall back to V1 extra-ctx attrs so FIA can read capturing.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Compiled FIA/MoE read _EXTRA_CTX, which called use_v2_model_runner()
and hit logger.warning_once / info_once. Dynamo cannot trace those
logs, so LoRA and non-whitelist models failed compile, and V2 MoE
baked additional_kwargs.get("moe_comm_method") as None.

Disable Dynamo on the extras helper, proxy getattr/setattr, and
use_v2_model_runner so isolation still uses get_current_vllm_config
at runtime.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Decorating __getattr__/__setattr__ with torch._dynamo.disable made
mypy treat _EXTRA_CTX as having no dynamic attributes. Keep the
dunders undecorated and disable the helpers they call instead.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cpu-ut installs torch 2.10, which marks @torch._dynamo.disable
with _torchdynamo_disable instead of _dynamo_disable.

Signed-off-by: yjyang62 <yangjinyang5@huawei.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Add MiniMaxM2ForCausalLM, DeepseekV3ForCausalLM, GlmMoeDsaForCausalLM,
DeepseekV4ForCausalLM, and Qwen3_5MoeForCausalLM to the default V2 model
whitelist, and allow dspark on the feature whitelist.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
yjyang62 and others added 16 commits September 17, 2026 13:52
Qwen3MoeForCausalLM now defaults to MRv2, which rejects V1
dynamic EPLB fields and DYNAMIC_EPLB. Keep the two-card EPLB
serve tests on V1; MRv2 EPLB remains covered by
test_qwen3_mrv2_eplb.py.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3 + dspark now defaults to Model Runner V2. Update the
single-block sliding-window acceptance curve to the V2 CI
measurement; the no-window dspark baseline is unchanged.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Calling use_v2_model_runner() from compiled _EXTRA_CTX getattr
traces logger.warning_once and graph-breaks with @torch._dynamo.disable.
Eager-cache the decision as _USE_V2_EXTRA_KWARGS so compiled FIA/MoE
constant-folds the extras path like VLLM_USE_V2_MODEL_RUNNER.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
DeepseekV32ForCausalLM now defaults to MRv2, so set_additional_forward_context
takes the V2 TP-group path. This kernel-only test never initializes
distributed groups. Pin V1 to keep the existing SFA V1 precision coverage.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
DeepseekV32 now defaults to V2, so SFA kernel precision called
set_additional_forward_context without TP and crashed. Qwen3.5 MTP
hang regression and eagle3 draft_window_size still belong on V1.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Drop the LoRA exclusion from is_supported_v2_model_runner_feature so
whitelisted models keep Model Runner V2 when LoRA is enabled. Dynamic
speculative decoding remains excluded.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3MoeForCausalLM now defaults to MRv2, which rejects V1 dynamic
EPLB fields. Keep the original 235B EPLB nightly case on V1; MRv2 EPLB
is covered by Qwen3-235B-A22B-W8A8-MRV2-EPLB.yaml.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
DSpark draft_window_size is not supported on Model Runner V2, so the
default V2 path falls back to V1 when it is set. Restore the
dspark_sliding_window golden to the V1 CI baseline.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
MRv2 FULL decode ACL graph capture hits a D2H .tolist() on device
seq_lens when parallel_drafting is enabled. Keep this acceptance
case on V1 until that capture path is safe.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3Moe now defaults to MRv2, which rejects V1 DYNAMIC_EPLB fields.
Switch the W8A8 EPLB case to --enable-eplb + load_collection_phase so
it starts on V2. Leave the PIECEWISE and FULL cases unchanged.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Draft VllmConfig uses runner_type="draft" and a draft architecture such
as DeepSeekV4MTPModel, so the generate-only whitelist dropped the draft
worker to V1 while the target stayed on V2. Treat draft configs as
inheriting the target's V2 decision, and normalize hf_overrides to {}
before replace so ModelSlim/pydantic accept the draft copy.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Spicy-Stick <873805887@qq.com>
Default MRv2 hangs on LoRA-only Qwen3 generate (copy_event.synchronize
107020). Exclude LoRA from the feature whitelist again so those e2e
paths stay on V1. Add Qwen3_5ForConditionalGeneration to the model
whitelist, including the GDN hybrid case.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Drop the is_hybrid gate so any architecture on the default-V2
whitelist uses Model Runner V2 even when the model is hybrid
(e.g. Qwen3.5 GDN). Attention-free models remain on V1. LoRA
and dynamic spec still stay on V1 via the feature whitelist.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: ZhangwenTaoHW <zhangwentao101@huawei.com>
Pass optional reasoning_effort through AISBench chat request generation.
Set reasoning_effort: low for the DeepSeek-V4 Flash W8A8 A3 nightly GPQA
case. Keep thinking: true and the performance case unchanged.

Cherry-picked from vllm-project#16721

Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch from 29c01ae to 3c33450 Compare September 17, 2026 13:53
Qwen3_5ForConditionalGeneration hybrid VL hits MRv2 encoder graph
capture (CUDA stream assert) and hybrid KV copy. Keep it on V1 by
default. Pin dflash2 PIECEWISE acceptance to V1 after vllm-project#16726 broke
dummy propose without set_forward_context; V2 eager remains covered.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants