Skip to content

[Refactor][Device][4/N] Migrate worker and patch selection to hardware profiles - #15256

Merged
Tflowers-0129 merged 1 commit into
vllm-project:mainfrom
Tflowers-0129:8-6-device-hal-worker
Aug 31, 2026
Merged

Tflowers-0129 merged 1 commit into
vllm-project:mainfrom
Tflowers-0129:8-6-device-hal-worker

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Aug 29, 2026 •

Copy link
Copy Markdown
Collaborator

What

  • Add the worker/patch capabilities consumed by this layer to the hardware profile matrix.
  • Replace worker and monkey-patch device-identity branches with capability checks.
  • Update focused unit-test mocks and capability expectations.

Why

This is PR 4/7 of the Device HAL / Hardware Profile refactor series, following merged PR #14076. It keeps hardware identity in detection/profile registration while shared worker and patch logic consumes capabilities.

There is no user-visible behavior change; existing per-device behavior is preserved.

Validation

  • Ruff check on all 12 changed Python files: passed.
  • Ruff format check on all 12 changed Python files: passed.
  • Python compileall on all changed Python files: passed.
  • git diff --check upstream/main...HEAD: passed.
  • Targeted unit tests and the Python 3.10/3.11/3.12 mypy matrix were not available on the development host and are left to repository CI.

Series

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request is the fourth in a series of seven aimed at refactoring the Device HAL. It shifts the responsibility of feature activation from hardcoded device identity checks to a flexible hardware profile matrix. This improves maintainability and decouples business logic from specific hardware versions, ensuring that features are enabled based on actual supported capabilities rather than device names.

Highlights

  • Capability-Based Refactoring: Migrated worker and patch selection logic from hardcoded device-specific checks (e.g., is_310p) to a flexible capability-based system using the HardwareCapability enum.
  • Expanded Capability Matrix: Added new capabilities to the hardware profile matrix, including ATB_EXTENSIONS, FP8_ATTENTION, and TRITON_BATCH_MEMCPY, to support a wider range of features.
  • Test Suite Updates: Updated unit tests to align with the new architecture, replacing device-type mocks with hardware profile mocks to ensure correct capability verification.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Refactor hardware capability checks to use hardware profile supports API

Suggested PR Summary:

### What this PR does / why we need it?
This PR refactors the hardware capability checks across the codebase. Instead of querying specific device types directly (e.g., `is_310p()` or checking for `AscendDeviceType.A5`), the code now utilizes a unified capability-based approach via `get_current_hardware_profile().supports(HardwareCapability...)`. This decouples feature logic from specific hardware models, making it easier to maintain and extend to new hardware profiles.

Specifically, this PR:
- Defines new capabilities in `HardwareCapability` (e.g., `ATB_EXTENSIONS`, `ATB_WARMUP`, `DISTRIBUTED_COMMUNICATION_ADAPTATION`, `FP8_ATTENTION`, `GDN_COMPATIBILITY`, `LOCAL_KV_COMM_RESOURCE`, `STANDARD_MAMBA_PATCH`, `TRITON_BATCH_MEMCPY`).
- Updates hardware profile definitions and standard capabilities.
- Replaces direct device type checks in patches and worker code with capability checks.
- Updates unit tests to mock the new hardware profile retrieval functions.

### Does this PR introduce _any_ user-facing change?
No. This is an internal refactoring of hardware capability checks and does not change any user-facing APIs or behavior.

### How was this patch tested?
The changes were verified by updating and running the existing unit tests in `tests/ut/device/test_hardware_profile.py`, `tests/ut/patch/worker/test_patch_kimi_k25.py`, and `tests/ut/worker/a2/test_worker_v1.py`.

I have reviewed the changes and have no additional feedback to provide as the refactoring is clean and the tests have been appropriately updated.

@Tflowers-0129 Tflowers-0129 added the ready-precise run selected e2e test for pr label Aug 29, 2026
@Tflowers-0129
Tflowers-0129 disabled auto-merge August 29, 2026 03:00
@Tflowers-0129
Tflowers-0129 enabled auto-merge (squash) August 29, 2026 03:09
@Tflowers-0129

Copy link
Copy Markdown
Collaborator Author

CI status update: all selected NPU jobs passed on attempt 2, including the jobs that initially failed while ModelScope was returning HTTP 500.

The only remaining test failure is the CPU EPLB job, which reproduced identically on both attempts after CI rebased the PR onto the fixed main snapshot bb6d5b82ecb97da8fa0f8fa1905af4a717c642a7:

  • TestAscendEPLBController::test_closed_window_clears_every_registered_model
  • TestAscendEPLBController::test_nonmatching_phase_is_forwarded_as_dummy

The tests expect closed-window load clearing and nonmatching phases to be forwarded as dummy, while the current AscendEPLBController inherits the upstream step() behavior and does neither. This PR does not modify vllm_ascend/worker/v2/eplb.py or tests/ut/worker/v2/test_eplb_controller.py; those files have the same relevant behavior on current main.

Rerun job: https://github.com/vllm-project/vllm-ascend/actions/runs/33230211762/job/99061919767
Resulting gate failure: https://github.com/vllm-project/vllm-ascend/actions/runs/33230211762/job/99071307879

Could a maintainer confirm the upstream baseline fix/handling? I am keeping the unrelated EPLB change out of this worker/patch capability-refactor PR.

@Tflowers-0129
Tflowers-0129 force-pushed the 8-6-device-hal-worker branch from 5b8eaaf to 9719afa Compare August 29, 2026 09:23
@Tflowers-0129

Copy link
Copy Markdown
Collaborator Author

Update: upstream PR #15283 ([Test] Fix UT for EPLBController) has now landed on main. I rebased this PR onto main@a507ee4ce5990b1e36470da02a31ed7fe9be9abf; the new head is 9719afa8925f94d2b96af037ceeefb07f9d30de4.

The PR commit is range-diff equivalent to the previous head, keeps its Signed-off-by trailer, and still changes only the original 12 worker/patch capability-refactor files. Changed-file Ruff check, Ruff format check, compileall, and git diff --check passed. A fresh E2E run is in progress: https://github.com/vllm-project/vllm-ascend/actions/runs/33245366935

@Tflowers-0129

Copy link
Copy Markdown
Collaborator Author

The fresh CPU suite after the EPLB baseline fix now has a single, different failure (2818 passed, 12 skipped, 1 failed):

This is the known main regression tracked by #15308 after #15267 changed mode 1 to roll back to dispatch_ffn_combine. The existing fix is #15303, whose CI is in progress. This Device HAL PR does not modify ascend_config.py, this test, or MegaMoE selection, so I will keep that unrelated fix out of PR #15256 and rebase once #15303 lands.

@Tflowers-0129 Tflowers-0129 added ready-all run all e2e test for pr and removed ready-precise run selected e2e test for pr labels Aug 31, 2026
…e profiles

Signed-off-by: frost_mourne <2906339855@qq.com>
@Tflowers-0129
Tflowers-0129 force-pushed the 8-6-device-hal-worker branch from 9719afa to b801633 Compare August 31, 2026 01:20

Copy link
Copy Markdown
Collaborator Author

Upstream fix #15303 has merged as 813f7374b. I rebased this PR onto that exact main head and force-pushed only after confirming the remote branch was still at 9719afa89.

  • New head: b801633b6
  • git range-diff: patch-equivalent (9719afa89 = b801633b6)
  • Diff boundary unchanged: 12 files, 59 insertions / 33 deletions
  • Signed-off-by preserved
  • Ruff check: passed
  • Ruff format --check: passed
  • Python compileall: passed
  • git diff --check: passed
  • Shared production-code device-identity addition check: passed

A fresh E2E run has started: https://github.com/vllm-project/vllm-ascend/actions/runs/33347276298

@Tflowers-0129
Tflowers-0129 merged commit d93a017 into vllm-project:main Aug 31, 2026
55 of 59 checks passed
Tflowers-0129 added a commit that referenced this pull request Aug 31, 2026
…e profiles (#15376)

## What

- Add the attention/quantization capabilities first consumed by this
layer to the hardware profile matrix.
- Replace device-identity branches in attention, quantization, and the
related DeepSeek indexer path with capability checks.
- Cover the exact A5-only MLA behaviors for native MLAPO weights and
decode prolog without RoPE.
- Update focused unit-test mocks and capability expectations.

## Why

This is PR 5/7 of the Device HAL / Hardware Profile refactor series,
following merged PR #15256. It keeps hardware identity in
detection/profile registration while shared attention and quantization
logic consumes explicit hardware capabilities.

There is no user-visible behavior change; existing per-device behavior
is preserved.

## Validation

- Rebased onto `main@d93a017` (merged
PR #15256).
- `git range-diff`: the original two PR5 commits are patch-equivalent
after rebase; one signed follow-up migrates two MLA identity branches
added by newer upstream code.
- Ruff check on all 16 changed Python files: passed.
- Ruff format check on all 16 changed Python files: passed.
- Python compileall on all 16 changed Python files: passed.
- `git diff --check upstream/main...HEAD`: passed.
- Added-device-identity check in changed production code: passed.
- The development host has no pytest or mypy installation and only
Python 3.10, so the focused unit tests and Python 3.10/3.11/3.12 mypy
comparison are left to repository CI.

## Series

- 5/7: attention/quantization hardware-profile migration
- Depends on: #15256 (merged as
`d93a0174985d7f926a5375cd823040a7439e34e9`)
- Next: model runner/DeepSeek (submitted only after this PR merges)

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: frost_mourne <2906339855@qq.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Aug 31, 2026
…e profiles (vllm-project#15256)

## What

- Add the worker/patch capabilities consumed by this layer to the
hardware profile matrix.
- Replace worker and monkey-patch device-identity branches with
capability checks.
- Update focused unit-test mocks and capability expectations.

## Why

This is PR 4/7 of the Device HAL / Hardware Profile refactor series,
following merged PR vllm-project#14076. It keeps hardware identity in
detection/profile registration while shared worker and patch logic
consumes capabilities.

There is no user-visible behavior change; existing per-device behavior
is preserved.

## Validation

- Ruff check on all 12 changed Python files: passed.
- Ruff format check on all 12 changed Python files: passed.
- Python compileall on all changed Python files: passed.
- `git diff --check upstream/main...HEAD`: passed.
- Targeted unit tests and the Python 3.10/3.11/3.12 mypy matrix were not
available on the development host and are left to repository CI.

## Series

- 4/7: worker/patch hardware-profile migration
- Depends on: vllm-project#14076 (merged)
- Next: attention/quantization (submitted only after this PR merges)

- vLLM main:
vllm-project/vllm@ba07e4a

Signed-off-by: frost_mourne <2906339855@qq.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
…e profiles (vllm-project#15256)

## What

- Add the worker/patch capabilities consumed by this layer to the
hardware profile matrix.
- Replace worker and monkey-patch device-identity branches with
capability checks.
- Update focused unit-test mocks and capability expectations.

## Why

This is PR 4/7 of the Device HAL / Hardware Profile refactor series,
following merged PR vllm-project#14076. It keeps hardware identity in
detection/profile registration while shared worker and patch logic
consumes capabilities.

There is no user-visible behavior change; existing per-device behavior
is preserved.

## Validation

- Ruff check on all 12 changed Python files: passed.
- Ruff format check on all 12 changed Python files: passed.
- Python compileall on all changed Python files: passed.
- `git diff --check upstream/main...HEAD`: passed.
- Targeted unit tests and the Python 3.10/3.11/3.12 mypy matrix were not
available on the development host and are left to repository CI.

## Series

- 4/7: worker/patch hardware-profile migration
- Depends on: vllm-project#14076 (merged)
- Next: attention/quantization (submitted only after this PR merges)

- vLLM main:
vllm-project/vllm@ba07e4a

Signed-off-by: frost_mourne <2906339855@qq.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
…e profiles (vllm-project#15376)

## What

- Add the attention/quantization capabilities first consumed by this
layer to the hardware profile matrix.
- Replace device-identity branches in attention, quantization, and the
related DeepSeek indexer path with capability checks.
- Cover the exact A5-only MLA behaviors for native MLAPO weights and
decode prolog without RoPE.
- Update focused unit-test mocks and capability expectations.

## Why

This is PR 5/7 of the Device HAL / Hardware Profile refactor series,
following merged PR vllm-project#15256. It keeps hardware identity in
detection/profile registration while shared attention and quantization
logic consumes explicit hardware capabilities.

There is no user-visible behavior change; existing per-device behavior
is preserved.

## Validation

- Rebased onto `main@d93a017` (merged
PR vllm-project#15256).
- `git range-diff`: the original two PR5 commits are patch-equivalent
after rebase; one signed follow-up migrates two MLA identity branches
added by newer upstream code.
- Ruff check on all 16 changed Python files: passed.
- Ruff format check on all 16 changed Python files: passed.
- Python compileall on all 16 changed Python files: passed.
- `git diff --check upstream/main...HEAD`: passed.
- Added-device-identity check in changed production code: passed.
- The development host has no pytest or mypy installation and only
Python 3.10, so the focused unit tests and Python 3.10/3.11/3.12 mypy
comparison are left to repository CI.

## Series

- 5/7: attention/quantization hardware-profile migration
- Depends on: vllm-project#15256 (merged as
`d93a0174985d7f926a5375cd823040a7439e34e9`)
- Next: model runner/DeepSeek (submitted only after this PR merges)

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: frost_mourne <2906339855@qq.com>
LQDLove pushed a commit to LQDLove/vllm-ascend that referenced this pull request Sep 5, 2026
…e profiles (vllm-project#15376)

## What

- Add the attention/quantization capabilities first consumed by this
layer to the hardware profile matrix.
- Replace device-identity branches in attention, quantization, and the
related DeepSeek indexer path with capability checks.
- Cover the exact A5-only MLA behaviors for native MLAPO weights and
decode prolog without RoPE.
- Update focused unit-test mocks and capability expectations.

## Why

This is PR 5/7 of the Device HAL / Hardware Profile refactor series,
following merged PR vllm-project#15256. It keeps hardware identity in
detection/profile registration while shared attention and quantization
logic consumes explicit hardware capabilities.

There is no user-visible behavior change; existing per-device behavior
is preserved.

## Validation

- Rebased onto `main@d93a017` (merged
PR vllm-project#15256).
- `git range-diff`: the original two PR5 commits are patch-equivalent
after rebase; one signed follow-up migrates two MLA identity branches
added by newer upstream code.
- Ruff check on all 16 changed Python files: passed.
- Ruff format check on all 16 changed Python files: passed.
- Python compileall on all 16 changed Python files: passed.
- `git diff --check upstream/main...HEAD`: passed.
- Added-device-identity check in changed production code: passed.
- The development host has no pytest or mypy installation and only
Python 3.10, so the focused unit tests and Python 3.10/3.11/3.12 mypy
comparison are left to repository CI.

## Series

- 5/7: attention/quantization hardware-profile migration
- Depends on: vllm-project#15256 (merged as
`d93a0174985d7f926a5375cd823040a7439e34e9`)
- Next: model runner/DeepSeek (submitted only after this PR merges)

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: frost_mourne <2906339855@qq.com>
Tflowers-0129 pushed a commit that referenced this pull request Sep 20, 2026
#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (#14076, #15256, #15376, #15407, #15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
vllm-project#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
vllm-project#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
vllm-project#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-all run all e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants