Conversation
cursor
Bot
force-pushed
the
cursor/expand-mrv2-whitelist-9095
branch
from
September 16, 2026 04:05
4fdff6c to
ae37a4e
Compare
cursor
Bot
force-pushed
the
enable-mrv2-whitelist-deae
branch
from
September 16, 2026 04:44
da0c16d to
3538c10
Compare
cursor
Bot
force-pushed
the
cursor/expand-mrv2-whitelist-9095
branch
from
September 16, 2026 04:53
ae37a4e to
df01f7c
Compare
cursor
Bot
force-pushed
the
cursor/expand-mrv2-whitelist-9095
branch
3 times, most recently
from
September 16, 2026 12:58
ba671cc to
ae0186a
Compare
Rebase vllm-project#11692 onto current vllm-project/vllm-ascend main. Replace the env-only use_v2_model_runner override with Ascend-owned whitelist heuristics (Qwen3ForCausalLM; eagle3/mtp/dflash; Triton; non-310P). Keep later unsupported-feature patches for spec-PP and Ascend-supported V1 features (dspark/dflash2). Explicit VLLM_USE_V2_MODEL_RUNNER still wins when set. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Qwen3ForCausalLM now defaults to MRv2. The RLHF sleep/wake suite still depends on the V1 generate path after CuMem remap, so keep the subprocess server on VLLM_USE_V2_MODEL_RUNNER=0. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 stores capturing on the forward-context object. Ascend FIA treats _EXTRA_CTX.capturing as ACL graph capture. When Qwen3 defaults to V2 via the whitelist (env unset), extras leaked onto ctx.capturing and graph_task_group_begin ran on a non-capturing stream (error 107029). Route extras through additional_kwargs whenever VllmConfig.use_v2_model_runner is true, not only when VLLM_USE_V2_MODEL_RUNNER is set. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Whitelist-enabled V2 routed extras through additional_kwargs whenever ctx.vllm_config.use_v2_model_runner was truthy. cpu-ut fixtures pass a bare MagicMock forward context, which auto-creates a truthy flag and hides capturing / max_tokens_across_dp. Require an actual bool True. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 sets forward_context.capturing during piecewise warmup. Ascend FIA treated that as ACL capture and called graph_task_group_begin, which failed with 107029 and hung Qwen3 whitelist-MRv2 e2e. Use the live NPU stream capture state (and skip PIECEWISE) instead of _EXTRA_CTX.capturing when wrapping FIA/PA kernels. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cpu-ut stubs torch.npu as MagicMock, so is_current_stream_capturing() is truthy and FIA entered full_graph_fia. Require an actual True, matching use_v2_model_runner. Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
A2 CI hung on LoRA-only generate and failed dflash acceptance when batch-size-based dynamic K (num_speculative_tokens_per_batch_size) was enabled under the Qwen3 default-V2 whitelist. Drop both from the default feature whitelist; static eagle3/mtp/dflash stay enabled, and VLLM_USE_V2_MODEL_RUNNER still overrides. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
GPU V2 PrefetchOffloader calls torch.cuda.is_current_stream_capturing during load_model. V1 already remaps that CUDA dummy to torch.npu; V2 torch_cuda_wrapper did not, so Qwen3 default-V2 prefetch e2e crashed on NPU. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Qwen3ForCausalLM now defaults to Model Runner V2. GPU V2 prefetch plus NZ graph capture matches the eager baseline, so the strict xfail on test_prefetch_offload_accuracy[NZ-graph] XPASS-fails a2-1 CI. Keep the V1 AscendPrefetchOffloader fail-fast for the unsupported combo. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Route _EXTRA_CTX through additional_kwargs when use_v2_model_runner(get_current_vllm_config()) is true. Whitelist-default V2 leaves VLLM_USE_V2_MODEL_RUNNER unset and GPU ForwardContext has no vllm_config, so env/ctx checks leaked GPU capturing onto FIA. Restore FIA/PA graph_task_group gating to _EXTRA_CTX.capturing. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Replace `_extra_ctx_uses_additional_kwargs` with `if use_v2_model_runner(get_current_vllm_config()) is True`. Guard xlite `index_full_mask` for mypy after merging main. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep getattr/setattr the same as before, only replacing the helper with `if use_v2_model_runner(get_current_vllm_config()) is True`. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Keep getattr/setattr the same as main. Only replace the V2 condition with `use_v2_model_runner(get_current_vllm_config()) is True`. Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
get_current_vllm_config() raises in cpu-ut attention fixtures. Fall back to V1 extra-ctx attrs so FIA can read capturing. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Compiled FIA/MoE read _EXTRA_CTX, which called use_v2_model_runner()
and hit logger.warning_once / info_once. Dynamo cannot trace those
logs, so LoRA and non-whitelist models failed compile, and V2 MoE
baked additional_kwargs.get("moe_comm_method") as None.
Disable Dynamo on the extras helper, proxy getattr/setattr, and
use_v2_model_runner so isolation still uses get_current_vllm_config
at runtime.
Signed-off-by: yjyang62 <yangjinyang5@huawei.com>
Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Decorating __getattr__/__setattr__ with torch._dynamo.disable made mypy treat _EXTRA_CTX as having no dynamic attributes. Keep the dunders undecorated and disable the helpers they call instead. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cpu-ut installs torch 2.10, which marks @torch._dynamo.disable with _torchdynamo_disable instead of _dynamo_disable. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Add MiniMaxM2ForCausalLM, DeepseekV3ForCausalLM, GlmMoeDsaForCausalLM, DeepseekV4ForCausalLM, and Qwen3_5MoeForCausalLM to the default V2 model whitelist, and allow dspark on the feature whitelist. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3MoeForCausalLM now defaults to MRv2, which rejects V1 dynamic EPLB fields and DYNAMIC_EPLB. Keep the two-card EPLB serve tests on V1; MRv2 EPLB remains covered by test_qwen3_mrv2_eplb.py. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3 + dspark now defaults to Model Runner V2. Update the single-block sliding-window acceptance curve to the V2 CI measurement; the no-window dspark baseline is unchanged. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Calling use_v2_model_runner() from compiled _EXTRA_CTX getattr traces logger.warning_once and graph-breaks with @torch._dynamo.disable. Eager-cache the decision as _USE_V2_EXTRA_KWARGS so compiled FIA/MoE constant-folds the extras path like VLLM_USE_V2_MODEL_RUNNER. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
DeepseekV32ForCausalLM now defaults to MRv2, so set_additional_forward_context takes the V2 TP-group path. This kernel-only test never initializes distributed groups. Pin V1 to keep the existing SFA V1 precision coverage. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
DeepseekV32 now defaults to V2, so SFA kernel precision called set_additional_forward_context without TP and crashed. Qwen3.5 MTP hang regression and eagle3 draft_window_size still belong on V1. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Drop the LoRA exclusion from is_supported_v2_model_runner_feature so whitelisted models keep Model Runner V2 when LoRA is enabled. Dynamic speculative decoding remains excluded. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3MoeForCausalLM now defaults to MRv2, which rejects V1 dynamic EPLB fields. Keep the original 235B EPLB nightly case on V1; MRv2 EPLB is covered by Qwen3-235B-A22B-W8A8-MRV2-EPLB.yaml. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
DSpark draft_window_size is not supported on Model Runner V2, so the default V2 path falls back to V1 when it is set. Restore the dspark_sliding_window golden to the V1 CI baseline. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
MRv2 FULL decode ACL graph capture hits a D2H .tolist() on device seq_lens when parallel_drafting is enabled. Keep this acceptance case on V1 until that capture path is safe. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Qwen3Moe now defaults to MRv2, which rejects V1 DYNAMIC_EPLB fields. Switch the W8A8 EPLB case to --enable-eplb + load_collection_phase so it starts on V2. Leave the PIECEWISE and FULL cases unchanged. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Draft VllmConfig uses runner_type="draft" and a draft architecture such
as DeepSeekV4MTPModel, so the generate-only whitelist dropped the draft
worker to V1 while the target stayed on V2. Treat draft configs as
inheriting the target's V2 decision, and normalize hf_overrides to {}
before replace so ModelSlim/pydantic accept the draft copy.
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Spicy-Stick <873805887@qq.com>
Default MRv2 hangs on LoRA-only Qwen3 generate (copy_event.synchronize 107020). Exclude LoRA from the feature whitelist again so those e2e paths stay on V1. Add Qwen3_5ForConditionalGeneration to the model whitelist, including the GDN hybrid case. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Drop the is_hybrid gate so any architecture on the default-V2 whitelist uses Model Runner V2 even when the model is hybrid (e.g. Qwen3.5 GDN). Attention-free models remain on V1. LoRA and dynamic spec still stay on V1 via the feature whitelist. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: ZhangwenTaoHW <zhangwentao101@huawei.com>
Pass optional reasoning_effort through AISBench chat request generation. Set reasoning_effort: low for the DeepSeek-V4 Flash W8A8 A3 nightly GPQA case. Keep thinking: true and the performance case unchanged. Cherry-picked from vllm-project#16721 Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
cursor
Bot
force-pushed
the
cursor/expand-mrv2-whitelist-9095
branch
from
September 17, 2026 13:53
29c01ae to
3c33450
Compare
Qwen3_5ForConditionalGeneration hybrid VL hits MRv2 encoder graph capture (CUDA stream assert) and hybrid KV copy. Keep it on V1 by default. Pin dflash2 PIECEWISE acceptance to V1 after vllm-project#16726 broke dummy propose without set_forward_context; V2 eager remains covered. Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cursor
Bot
force-pushed
the
main
branch
2 times, most recently
from
September 21, 2026 13:29
c82386b to
731afcb
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it?
Stacked on vllm-project#16203.
vllm-project#16203 introduces Ascend-owned default Model Runner V2 selection in
vllm_ascend/mrv2_utils.py(model + feature whitelists, env override, upstream validation no-op). This PR expands that default path to more production models anddspark, and adjusts related platform / spec-decode / CI coverage so default-V2 stays safe on NPU.Scope (范围)
DEFAULT_V2_MODEL_RUNNER_ARCHITECTURESinvllm_ascend/mrv2_utils.py): addQwen3MoeForCausalLM,MiniMaxM2ForCausalLM,DeepseekV3ForCausalLM,DeepseekV32ForCausalLM,GlmMoeDsaForCausalLM,DeepseekV4ForCausalLM, andQwen3_5MoeForCausalLMon top of [Misc][MRv2]enable model runner v2 by default via whitelists vllm-project/vllm-ascend#16203’sQwen3ForCausalLM.Qwen3_5ForConditionalGeneration(Qwen3.5/3.6 dense and VL hybrid) is not on the whitelist. It stays on V1 by default because MRv2 hits hybrid VL encoder-graph capture and hybrid KV copy-with-prefix-cache paths that are not ready for default enablement.dsparkalongsideeagle3/mtp/dflash. Still default V1 when LoRA is enabled, whennum_speculative_tokens_per_batch_size(dynamic spec decode) is set, or when DSpark KV sliding window (additional_config.draft_window_size) is set.is_hybrid=Truedoes not force V1 (e.g.Qwen3_5MoeForCausalLM). Attention-free models and non-generaterunner types (exceptrunner_type="draft", which inherits the target’s V2 decision) remain V1.test_dflash2_acceptancePIECEWISE after [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (#16425) vllm-project/vllm-ascend#16726, and similar) viaVLLM_USE_V2_MODEL_RUNNER=0while moving validated nightly cases to V2 where appropriate.tests/ut/test_mrv2_utils.py; refresh or pin e2e goldens for default-V2 behavior; nightly DeepSeek-V4 Flash GPQA usesreasoning_effort: lowin aisbench config ([CI] Pin DeepSeek-V4 Flash GPQA nightly to low reasoning effort vllm-project/vllm-ascend#16721).Out of scope for the whitelist change itself: general MRv2 feature development on
mainthat is not required to land the expanded defaults (reviewers should still treat this branch as stacked on vllm-project#16203 for the core gating logic).MRv2 switching scheme (切换方案)
Decision order in
use_v2_model_runner()(vllm_ascend/mrv2_utils.py), installed onVllmConfigbyapply_v2_model_runner_config_patch()(vllm-project#16203):VLLM_USE_V2_MODEL_RUNNERis0or1, use that value (explicit opt-out or opt-in for any model/feature).is_default_v2_model_runner_model()is true:runner_type == "generate"and architecture ∈ whitelist, orrunner_type == "draft"(inherit target whitelist; do not re-check draft architecture).is_attention_free, embedding/other non-generate runners, architectures not in the whitelist (includingQwen3_5ForConditionalGeneration)._v2_model_runner_environment_ready()requiresis_supported_v2_model_runner_feature():{eagle3, mtp, dflash, dspark}without the exclusions below.num_speculative_tokens_per_batch_size; DSpark withdraft_window_size; any other speculative method.HAS_TRITONmust be true.Default-V2 model architectures
Qwen3ForCausalLMQwen3MoeForCausalLMMiniMaxM2ForCausalLMDeepseekV3ForCausalLMDeepseekV32ForCausalLMGlmMoeDsaForCausalLMDeepseekV4ForCausalLMQwen3_5MoeForCausalLMDefault-V2 speculative methods:
eagle3,mtp,dflash,dspark(subject to the feature exclusions above).Does this PR introduce any user-facing change?
Yes. The architectures above now default to Model Runner V2 on A2/A3/A5 when Triton is present, the platform is not 310P, and the feature whitelist passes. Whitelisted hybrid variants (for example Qwen3.5 MoE) also default to V2;
Qwen3_5ForConditionalGenerationdense/VL stays on V1. LoRA, dynamic speculative decoding, and DSpark sliding window stay on V1 unlessVLLM_USE_V2_MODEL_RUNNER=1. All other models remain V1 by default. ExplicitVLLM_USE_V2_MODEL_RUNNER=0/1behavior is unchanged.How was this patch tested?
Unit tests added/updated:
tests/ut/test_mrv2_utils.py(expanded architecture list, hybrid whitelisted models,Qwen3_5ForConditionalGenerationexcluded,dsparkand DSpark sliding-window / LoRA / dynamic-spec gates, draft runner inheritance)tests/ut/tools/test_aisbench.py(reasoning_effortpassthrough) where included on this branchE2e: selective
VLLM_USE_V2_MODEL_RUNNERpins and updated goldens for default-V2 paths;test_dflash2_v2_acceptancestill exercises V2 eager whiletest_dflash2_acceptancestays on V1 for PIECEWISE ACL graph.CI: cpu-ut / NPU e2e on this PR.
vLLM main: vllm-project/vllm@84030bb