Revert "[Feature][MRV2] expand default MRv2 architecture whitelist and add dspark" - #16832
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
This pull request simplifies the Model Runner V2 enablement logic on Ascend by removing the custom architecture and feature whitelists (deleting mrv2_utils.py) and instead driving use_v2_model_runner directly via the VLLM_USE_V2_MODEL_RUNNER environment variable. It cleans up associated tests, mocks, and configurations, and restructures EPLB settings. The reviewer feedback highlights inconsistent types for the DYNAMIC_EPLB environment variable (boolean true vs. string "true") across several YAML configuration files, which could cause parsing issues. Additionally, the reviewer advises against replacing getattr with direct attribute access in xlite.py to avoid potential AttributeError exceptions with older configurations or mock objects.
16e1f7f to
8209dbf
Compare
…d add dspark" This reverts commit c7ca0b6. Signed-off-by: yjyang62 <yangjinyang5@huawei.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
2d99020 to
c5c8406
Compare
|
/nightly QWEN3_235B_PD
|
|
/nightly Qwen3-32B-QuaRot
|
|
/nightly Qwen3-30B-A3B-BF16-A2-TP4 |
… offload (#16544)" (#16905) ## What this PR does Reverts #16544, originally authored by @dragondream-chen, by reverting squash-merge commit `200309da4198f8c150d4dd365e55e98cb88b8400`. This PR contains only the revert of #16544. It does not include #16890 or any other feature/fix commit. ## Conflict resolution The revert had one conflict in `tests/ut/test_ascend_forward_context.py` because a later `main` commit (#16832) touched the same test. The resolution restores the pre-#16544 forward-context invocation while preserving #16832's later removal of the V2 cached-flag assertion. ## Validation - `git diff --check upstream/main..HEAD` - Verified the branch is exactly one commit ahead of the latest `main` - DCO sign-off included - Runtime tests not run locally; rely on repository CI / Ascend environment validation - vLLM main: vllm-project/vllm@84030bb Signed-off-by: pgzddxx <184603735+pgzddxx@users.noreply.github.com>
Rebase the PR onto latest main as a linear history so CI's `git rebase $BASE_SHA` no longer replays old commits onto the vllm-project#16832 tree that deleted mrv2_utils.py. Keep default-V2 selection, the Ascend feature blacklist, and the SFA C8 DCP hardware capability check from vllm-project#16656. Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com> Signed-off-by: Cursor Agent <cursoragent@cursor.com>
### What this PR does / why we need it? On vLLM 0.28.0, Ascend MRV2 PCP+DP is rejected during parallel configuration validation. Allowing configuration alone is insufficient: dispatch still synchronizes the global token count before PCP partitioning, so a 44-token prefill with PCP2 reaches `DPMetadata.make` with 44 recorded tokens but a 22-token local batch. - Scope the configuration workaround to vLLM 0.28.0, Ascend MRV2 explicitly enabled, DP > 1, PCP > 1, and DCP = 1. Preserve other validation and restore the real PCP size on every exit. - Reuse the existing rank-segment rules to compute the PCP execution count before dispatch, including replicated decode tokens and uneven prefill partitions. - Pass only the computed count through a call-scoped context to the original dispatch function. Preserve global request state, original DP synchronization, and its consistency checks. Based on upstream main `628fac6d8`. The linear commit preserves #16832's revert of default MRV2 whitelist selection; MRV2 remains explicitly enabled through the environment. The graph replay revert is already upstream in #16726 and is not included. Existing upstream MRV2 selection and KV-cache/preemption changes are retained. ### Does this PR introduce _any_ user-facing change? Enables the scoped PCP+DP configuration and fixes its dispatch token-count mismatch on vLLM 0.28.0. No new user-facing flags. ### How was this patch tested? - On revision `fb0554b7f`: the PR was linearized and successfully rebased onto the refreshed upstream main using the same operation as CI. The PR diff is byte-identical to the resolved merge diff. Ruff lint/format and AST syntax checks passed for all seven PR Python files; `git diff upstream/main --check` passed. The local `format.sh ci` attempt stopped because `pre-commit` is unavailable. Full checks are being rerun by GitHub CI. The previous run passed pre-commit but stopped during CI rebase, before unit tests started. - Before this merge: Ruff lint, Ruff format check, and `git diff --check` passed. The six transplanted adaptation functions and the added unit test/fixture definitions were also compared structurally against the tested local implementation. - On the prior local base (`bddbab4c3` plus the adaptation): 121 unit tests passed, 2 skipped in the Ascend development container, covering PCP partition counts, zero counts, dummy/no-PCP paths, version gating, context isolation/restoration, and configuration validation. The corresponding unit regression tests are included in this PR; the earlier E2E test expansion is excluded. - Earlier service smoke on that local line: MiniMax + Eagle3, TP2/DP2/PCP2, seq16, FlashComm1 and FULL_DECODE_ONLY completed the original 44-token request and four different-length concurrent requests. This preceded the count-only context simplification. The exact main-based PR revision has not yet been rerun in the remote runtime. Import/override compatibility, single-rank and multi-rank model smoke remain pending on this revision. Full accuracy, MTP/MLA coverage and performance regression validation are not claimed; keep this PR in draft. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com> Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Reverts #16626