[Refactor][Device][4/N] Migrate worker and patch selection to hardware profiles - #15256
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request is the fourth in a series of seven aimed at refactoring the Device HAL. It shifts the responsibility of feature activation from hardcoded device identity checks to a flexible hardware profile matrix. This improves maintainability and decouples business logic from specific hardware versions, ensuring that features are enabled based on actual supported capabilities rather than device names. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
745f892 to
5b8eaaf
Compare
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Refactor hardware capability checks to use hardware profile supports APISuggested PR Summary:
### What this PR does / why we need it?
This PR refactors the hardware capability checks across the codebase. Instead of querying specific device types directly (e.g., `is_310p()` or checking for `AscendDeviceType.A5`), the code now utilizes a unified capability-based approach via `get_current_hardware_profile().supports(HardwareCapability...)`. This decouples feature logic from specific hardware models, making it easier to maintain and extend to new hardware profiles.
Specifically, this PR:
- Defines new capabilities in `HardwareCapability` (e.g., `ATB_EXTENSIONS`, `ATB_WARMUP`, `DISTRIBUTED_COMMUNICATION_ADAPTATION`, `FP8_ATTENTION`, `GDN_COMPATIBILITY`, `LOCAL_KV_COMM_RESOURCE`, `STANDARD_MAMBA_PATCH`, `TRITON_BATCH_MEMCPY`).
- Updates hardware profile definitions and standard capabilities.
- Replaces direct device type checks in patches and worker code with capability checks.
- Updates unit tests to mock the new hardware profile retrieval functions.
### Does this PR introduce _any_ user-facing change?
No. This is an internal refactoring of hardware capability checks and does not change any user-facing APIs or behavior.
### How was this patch tested?
The changes were verified by updating and running the existing unit tests in `tests/ut/device/test_hardware_profile.py`, `tests/ut/patch/worker/test_patch_kimi_k25.py`, and `tests/ut/worker/a2/test_worker_v1.py`.I have reviewed the changes and have no additional feedback to provide as the refactoring is clean and the tests have been appropriately updated.
|
CI status update: all selected NPU jobs passed on attempt 2, including the jobs that initially failed while ModelScope was returning HTTP 500. The only remaining test failure is the CPU EPLB job, which reproduced identically on both attempts after CI rebased the PR onto the fixed
The tests expect closed-window load clearing and nonmatching phases to be forwarded as dummy, while the current Rerun job: https://github.com/vllm-project/vllm-ascend/actions/runs/33230211762/job/99061919767 Could a maintainer confirm the upstream baseline fix/handling? I am keeping the unrelated EPLB change out of this worker/patch capability-refactor PR. |
5b8eaaf to
9719afa
Compare
|
Update: upstream PR #15283 ( The PR commit is range-diff equivalent to the previous head, keeps its Signed-off-by trailer, and still changes only the original 12 worker/patch capability-refactor files. Changed-file Ruff check, Ruff format check, compileall, and |
|
The fresh CPU suite after the EPLB baseline fix now has a single, different failure (
This is the known |
…e profiles Signed-off-by: frost_mourne <2906339855@qq.com>
9719afa to
b801633
Compare
|
Upstream fix #15303 has merged as
A fresh E2E run has started: https://github.com/vllm-project/vllm-ascend/actions/runs/33347276298 |
…e profiles (#15376) ## What - Add the attention/quantization capabilities first consumed by this layer to the hardware profile matrix. - Replace device-identity branches in attention, quantization, and the related DeepSeek indexer path with capability checks. - Cover the exact A5-only MLA behaviors for native MLAPO weights and decode prolog without RoPE. - Update focused unit-test mocks and capability expectations. ## Why This is PR 5/7 of the Device HAL / Hardware Profile refactor series, following merged PR #15256. It keeps hardware identity in detection/profile registration while shared attention and quantization logic consumes explicit hardware capabilities. There is no user-visible behavior change; existing per-device behavior is preserved. ## Validation - Rebased onto `main@d93a017` (merged PR #15256). - `git range-diff`: the original two PR5 commits are patch-equivalent after rebase; one signed follow-up migrates two MLA identity branches added by newer upstream code. - Ruff check on all 16 changed Python files: passed. - Ruff format check on all 16 changed Python files: passed. - Python compileall on all 16 changed Python files: passed. - `git diff --check upstream/main...HEAD`: passed. - Added-device-identity check in changed production code: passed. - The development host has no pytest or mypy installation and only Python 3.10, so the focused unit tests and Python 3.10/3.11/3.12 mypy comparison are left to repository CI. ## Series - 5/7: attention/quantization hardware-profile migration - Depends on: #15256 (merged as `d93a0174985d7f926a5375cd823040a7439e34e9`) - Next: model runner/DeepSeek (submitted only after this PR merges) - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: frost_mourne <2906339855@qq.com>
…e profiles (vllm-project#15256) ## What - Add the worker/patch capabilities consumed by this layer to the hardware profile matrix. - Replace worker and monkey-patch device-identity branches with capability checks. - Update focused unit-test mocks and capability expectations. ## Why This is PR 4/7 of the Device HAL / Hardware Profile refactor series, following merged PR vllm-project#14076. It keeps hardware identity in detection/profile registration while shared worker and patch logic consumes capabilities. There is no user-visible behavior change; existing per-device behavior is preserved. ## Validation - Ruff check on all 12 changed Python files: passed. - Ruff format check on all 12 changed Python files: passed. - Python compileall on all changed Python files: passed. - `git diff --check upstream/main...HEAD`: passed. - Targeted unit tests and the Python 3.10/3.11/3.12 mypy matrix were not available on the development host and are left to repository CI. ## Series - 4/7: worker/patch hardware-profile migration - Depends on: vllm-project#14076 (merged) - Next: attention/quantization (submitted only after this PR merges) - vLLM main: vllm-project/vllm@ba07e4a Signed-off-by: frost_mourne <2906339855@qq.com>
…e profiles (vllm-project#15256) ## What - Add the worker/patch capabilities consumed by this layer to the hardware profile matrix. - Replace worker and monkey-patch device-identity branches with capability checks. - Update focused unit-test mocks and capability expectations. ## Why This is PR 4/7 of the Device HAL / Hardware Profile refactor series, following merged PR vllm-project#14076. It keeps hardware identity in detection/profile registration while shared worker and patch logic consumes capabilities. There is no user-visible behavior change; existing per-device behavior is preserved. ## Validation - Ruff check on all 12 changed Python files: passed. - Ruff format check on all 12 changed Python files: passed. - Python compileall on all changed Python files: passed. - `git diff --check upstream/main...HEAD`: passed. - Targeted unit tests and the Python 3.10/3.11/3.12 mypy matrix were not available on the development host and are left to repository CI. ## Series - 4/7: worker/patch hardware-profile migration - Depends on: vllm-project#14076 (merged) - Next: attention/quantization (submitted only after this PR merges) - vLLM main: vllm-project/vllm@ba07e4a Signed-off-by: frost_mourne <2906339855@qq.com>
…e profiles (vllm-project#15376) ## What - Add the attention/quantization capabilities first consumed by this layer to the hardware profile matrix. - Replace device-identity branches in attention, quantization, and the related DeepSeek indexer path with capability checks. - Cover the exact A5-only MLA behaviors for native MLAPO weights and decode prolog without RoPE. - Update focused unit-test mocks and capability expectations. ## Why This is PR 5/7 of the Device HAL / Hardware Profile refactor series, following merged PR vllm-project#15256. It keeps hardware identity in detection/profile registration while shared attention and quantization logic consumes explicit hardware capabilities. There is no user-visible behavior change; existing per-device behavior is preserved. ## Validation - Rebased onto `main@d93a017` (merged PR vllm-project#15256). - `git range-diff`: the original two PR5 commits are patch-equivalent after rebase; one signed follow-up migrates two MLA identity branches added by newer upstream code. - Ruff check on all 16 changed Python files: passed. - Ruff format check on all 16 changed Python files: passed. - Python compileall on all 16 changed Python files: passed. - `git diff --check upstream/main...HEAD`: passed. - Added-device-identity check in changed production code: passed. - The development host has no pytest or mypy installation and only Python 3.10, so the focused unit tests and Python 3.10/3.11/3.12 mypy comparison are left to repository CI. ## Series - 5/7: attention/quantization hardware-profile migration - Depends on: vllm-project#15256 (merged as `d93a0174985d7f926a5375cd823040a7439e34e9`) - Next: model runner/DeepSeek (submitted only after this PR merges) - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: frost_mourne <2906339855@qq.com>
…e profiles (vllm-project#15376) ## What - Add the attention/quantization capabilities first consumed by this layer to the hardware profile matrix. - Replace device-identity branches in attention, quantization, and the related DeepSeek indexer path with capability checks. - Cover the exact A5-only MLA behaviors for native MLAPO weights and decode prolog without RoPE. - Update focused unit-test mocks and capability expectations. ## Why This is PR 5/7 of the Device HAL / Hardware Profile refactor series, following merged PR vllm-project#15256. It keeps hardware identity in detection/profile registration while shared attention and quantization logic consumes explicit hardware capabilities. There is no user-visible behavior change; existing per-device behavior is preserved. ## Validation - Rebased onto `main@d93a017` (merged PR vllm-project#15256). - `git range-diff`: the original two PR5 commits are patch-equivalent after rebase; one signed follow-up migrates two MLA identity branches added by newer upstream code. - Ruff check on all 16 changed Python files: passed. - Ruff format check on all 16 changed Python files: passed. - Python compileall on all 16 changed Python files: passed. - `git diff --check upstream/main...HEAD`: passed. - Added-device-identity check in changed production code: passed. - The development host has no pytest or mypy installation and only Python 3.10, so the focused unit tests and Python 3.10/3.11/3.12 mypy comparison are left to repository CI. ## Series - 5/7: attention/quantization hardware-profile migration - Depends on: vllm-project#15256 (merged as `d93a0174985d7f926a5375cd823040a7439e34e9`) - Next: model runner/DeepSeek (submitted only after this PR merges) - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: frost_mourne <2906339855@qq.com>
#16803) ### What this PR does / why we need it? Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware profiles` series (#14076, #15256, #15376, #15407, #15478). That series migrated most `is_ 310p() ` call sites to semantic hardware-profile capabilities, but three call sites in the v2 worker runtime path were left behind, and the ` is _310p()` compatibility helper itself was kept alive. This PR finishes the migration and removes the helper. Residual call sites removed: - `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects `Ascend310PBlockTables` on 310P - `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the Triton-free 310P `ModelState` (2 sites) - `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P skips the upstream v2 model runner validation ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM main: vllm-project/vllm@84030bb Signed-off-by: spoon1116 <1522707055@qq.com>
vllm-project#16803) ### What this PR does / why we need it? Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478). That series migrated most `is_ 310p() ` call sites to semantic hardware-profile capabilities, but three call sites in the v2 worker runtime path were left behind, and the ` is _310p()` compatibility helper itself was kept alive. This PR finishes the migration and removes the helper. Residual call sites removed: - `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects `Ascend310PBlockTables` on 310P - `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the Triton-free 310P `ModelState` (2 sites) - `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P skips the upstream v2 model runner validation ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM main: vllm-project/vllm@84030bb Signed-off-by: spoon1116 <1522707055@qq.com>
vllm-project#16803) ### What this PR does / why we need it? Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478). That series migrated most `is_ 310p() ` call sites to semantic hardware-profile capabilities, but three call sites in the v2 worker runtime path were left behind, and the ` is _310p()` compatibility helper itself was kept alive. This PR finishes the migration and removes the helper. Residual call sites removed: - `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects `Ascend310PBlockTables` on 310P - `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the Triton-free 310P `ModelState` (2 sites) - `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P skips the upstream v2 model runner validation ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM main: vllm-project/vllm@84030bb Signed-off-by: spoon1116 <1522707055@qq.com>
vllm-project#16803) ### What this PR does / why we need it? Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478). That series migrated most `is_ 310p() ` call sites to semantic hardware-profile capabilities, but three call sites in the v2 worker runtime path were left behind, and the ` is _310p()` compatibility helper itself was kept alive. This PR finishes the migration and removes the helper. Residual call sites removed: - `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects `Ascend310PBlockTables` on 310P - `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the Triton-free 310P `ModelState` (2 sites) - `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P skips the upstream v2 model runner validation ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM main: vllm-project/vllm@84030bb Signed-off-by: spoon1116 <1522707055@qq.com>
What
Why
This is PR 4/7 of the Device HAL / Hardware Profile refactor series, following merged PR #14076. It keeps hardware identity in detection/profile registration while shared worker and patch logic consumes capabilities.
There is no user-visible behavior change; existing per-device behavior is preserved.
Validation
git diff --check upstream/main...HEAD: passed.Series
4/7: worker/patch hardware-profile migration
Depends on: [Refactor][Device][3/N] Migrate core runtime selection to hardware profiles #14076 (merged)
Next: attention/quantization (submitted only after this PR merges)
vLLM main: vllm-project/vllm@ba07e4a