Skip to content

[Refactor][Device][7/N] Migrate MoE compilation and sampling to hardware profiles - #15478

Merged
Tflowers-0129 merged 3 commits into
vllm-project:mainfrom
Tflowers-0129:8-6-device-hal-moe
Sep 3, 2026
Merged

Tflowers-0129 merged 3 commits into
vllm-project:mainfrom
Tflowers-0129:8-6-device-hal-moe

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Sep 1, 2026 •

Copy link
Copy Markdown
Collaborator

What

  • add the final MoE/compilation/sampling hardware capabilities and a MoECommPolicy
  • migrate MoE communication selection, graph fusion gates, routing, shared experts, token dispatch, SwiGLU-OAI MX quant, and top-k/top-p sampling away from direct device identity checks
  • preserve the latest upstream Situ, hierarchical communication, SFA, DSV4, sampler, and CANN MegaMoe paths
  • model the new A3-only CANN MegaMoe consumer with HardwareCapability.CANN_MEGAMOE and update the exact capability matrix

Why

Shared business logic should consume hardware capabilities and implementation policies instead of branching on Ascend device identities. This keeps hardware detection/profile registration as the single identity boundary and makes the final consumers explicit.

This is a refactor only; there is no user-visible behavior change.

Series

Device HAL / Hardware Profile refactor 7/N.

Depends on merged #15407 (d4ebe8a0e0fdf56135507259f90f80a45f736cab) and is rebased onto current main (30f54b5c341f015bcc124022e62156d3e5ef5919). This base includes the GLM5.3 documentation-link fix from #15508 and the MegaMoe prefill update from #14439.

Testing

  • ruff check on all 14 changed Python files
  • ruff format --check on all 14 changed Python files
  • python3 -m compileall -q on all 14 changed Python files
  • git diff --check upstream/main..HEAD
  • production-code scan for newly introduced direct device-identity branches
  • E2E run https://github.com/vllm-project/vllm-ascend/actions/runs/33579196584: 23 successful jobs, 4 skipped, zero failures; CPU, all selected NPU jobs, 310P, upstream, pre-commit, DCO, Read the Docs, and ci-gate succeeded

The local static checks above passed. The development server does not provide the complete pytest/mypy/vLLM environment, so targeted unit tests and Python 3.10/3.11/3.12 mypy comparison are covered by the successful CI run and are not claimed as local results.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request serves as the final installment (7/7) of the Device HAL refactor series. It completes the decoupling of shared business logic from specific Ascend device identities by transitioning MoE compilation, routing, and sampling to a hardware capability and policy-based architecture. This change improves maintainability and simplifies the integration of future hardware targets by ensuring that consumers of shared logic rely on explicit capability checks rather than branching on device identity.

Highlights

  • MoE Communication Policy: Introduced the MoECommPolicy enum to define and select communication strategies based on hardware profiles rather than hardcoded device types.
  • Hardware Capabilities: Added new HardwareCapability flags for MoE dispatch, sampling, and graph fusion features to support more granular hardware feature detection.
  • Logic Migration: Migrated core MoE compilation, routing, and sampling logic away from direct Ascend device identity checks to a capability-based system.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Misc] Refactor hardware capability and MoE communication policy checks

Suggested PR Summary:

### What this PR does / why we need it?
This PR refactors the hardware capability and MoE communication policy checks across the codebase. Instead of directly checking `AscendDeviceType` (e.g., `A2`, `A3`, `A5`, `310P`), the code now queries the active `HardwareProfile` for specific `HardwareCapability` flags and `MoECommPolicy` configurations. This decouples device-specific logic from operator and communication implementations, making the codebase more maintainable and extensible for future hardware generations.

Additionally, a review comment suggests optimizing the hot path in `select_moe_comm_method` by moving the `selector_by_policy` dictionary to the module level as a constant to avoid recreation overhead on every forward pass.

### Does this PR introduce _any_ user-facing change?
No. This is an internal refactoring of hardware profile and capability checks.

### How was this patch tested?
Existing unit tests in `tests/ut/device/test_hardware_profile.py`, `tests/ut/ops/a2/test_token_dispatcher.py`, `tests/ut/ops/test_moe_mlp.py`, and `tests/ut/test_ascend_forward_context.py` were updated to mock the hardware profile and passed successfully.

Comment thread vllm_ascend/ascend_forward_context.py Outdated
@Tflowers-0129 Tflowers-0129 changed the title [Refactor][Device][7/7] Migrate MoE compilation and sampling to hardware profiles [Refactor][Device][7/N] Migrate MoE compilation and sampling to hardware profiles Sep 1, 2026
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@Tflowers-0129 Tflowers-0129 added the ready-all run all e2e test for pr label Sep 2, 2026
…are profiles

Signed-off-by: frost_mourne <2906339855@qq.com>
Signed-off-by: frost_mourne <2906339855@qq.com>
Signed-off-by: frost_mourne <2906339855@qq.com>
@Tflowers-0129
Tflowers-0129 merged commit 83f9ef3 into vllm-project:main Sep 3, 2026
32 checks passed
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
…are profiles (vllm-project#15478)

## What

- add the final MoE/compilation/sampling hardware capabilities and a
`MoECommPolicy`
- migrate MoE communication selection, graph fusion gates, routing,
shared experts, token dispatch, SwiGLU-OAI MX quant, and top-k/top-p
sampling away from direct device identity checks
- preserve the latest upstream Situ, hierarchical communication, SFA,
DSV4, sampler, and CANN MegaMoe paths
- model the new A3-only CANN MegaMoe consumer with
`HardwareCapability.CANN_MEGAMOE` and update the exact capability matrix

## Why

Shared business logic should consume hardware capabilities and
implementation policies instead of branching on Ascend device
identities. This keeps hardware detection/profile registration as the
single identity boundary and makes the final consumers explicit.

This is a refactor only; there is no user-visible behavior change.

## Series

Device HAL / Hardware Profile refactor 7/N.

Depends on merged vllm-project#15407 (`d4ebe8a0e0fdf56135507259f90f80a45f736cab`)
and is rebased onto current `main`
(`30f54b5c341f015bcc124022e62156d3e5ef5919`). This base includes the
GLM5.3 documentation-link fix from vllm-project#15508 and the MegaMoe prefill update
from vllm-project#14439.

## Testing

- `ruff check` on all 14 changed Python files
- `ruff format --check` on all 14 changed Python files
- `python3 -m compileall -q` on all 14 changed Python files
- `git diff --check upstream/main..HEAD`
- production-code scan for newly introduced direct device-identity
branches
- E2E run
https://github.com/vllm-project/vllm-ascend/actions/runs/33579196584: 23
successful jobs, 4 skipped, zero failures; CPU, all selected NPU jobs,
310P, upstream, pre-commit, DCO, Read the Docs, and `ci-gate` succeeded

The local static checks above passed. The development server does not
provide the complete pytest/mypy/vLLM environment, so targeted unit
tests and Python 3.10/3.11/3.12 mypy comparison are covered by the
successful CI run and are not claimed as local results.

---------

Signed-off-by: frost_mourne <2906339855@qq.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…are profiles (vllm-project#15478)

## What

- add the final MoE/compilation/sampling hardware capabilities and a
`MoECommPolicy`
- migrate MoE communication selection, graph fusion gates, routing,
shared experts, token dispatch, SwiGLU-OAI MX quant, and top-k/top-p
sampling away from direct device identity checks
- preserve the latest upstream Situ, hierarchical communication, SFA,
DSV4, sampler, and CANN MegaMoe paths
- model the new A3-only CANN MegaMoe consumer with
`HardwareCapability.CANN_MEGAMOE` and update the exact capability matrix

## Why

Shared business logic should consume hardware capabilities and
implementation policies instead of branching on Ascend device
identities. This keeps hardware detection/profile registration as the
single identity boundary and makes the final consumers explicit.

This is a refactor only; there is no user-visible behavior change.

## Series

Device HAL / Hardware Profile refactor 7/N.

Depends on merged vllm-project#15407 (`d4ebe8a0e0fdf56135507259f90f80a45f736cab`)
and is rebased onto current `main`
(`30f54b5c341f015bcc124022e62156d3e5ef5919`). This base includes the
GLM5.3 documentation-link fix from vllm-project#15508 and the MegaMoe prefill update
from vllm-project#14439.

## Testing

- `ruff check` on all 14 changed Python files
- `ruff format --check` on all 14 changed Python files
- `python3 -m compileall -q` on all 14 changed Python files
- `git diff --check upstream/main..HEAD`
- production-code scan for newly introduced direct device-identity
branches
- E2E run
https://github.com/vllm-project/vllm-ascend/actions/runs/33579196584: 23
successful jobs, 4 skipped, zero failures; CPU, all selected NPU jobs,
310P, upstream, pre-commit, DCO, Read the Docs, and `ci-gate` succeeded

The local static checks above passed. The development server does not
provide the complete pytest/mypy/vLLM environment, so targeted unit
tests and Python 3.10/3.11/3.12 mypy comparison are covered by the
successful CI run and are not claimed as local results.

---------

Signed-off-by: frost_mourne <2906339855@qq.com>
Tflowers-0129 pushed a commit that referenced this pull request Sep 20, 2026
#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (#14076, #15256, #15376, #15407, #15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
xqchen7 pushed a commit to xqchen7/vllm-ascend that referenced this pull request Sep 22, 2026
vllm-project#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
zhaochuang001 pushed a commit to zhaochuang001/vllm-ascend that referenced this pull request Sep 22, 2026
vllm-project#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
vllm-project#16803)

### What this PR does / why we need it?
Follow-up to the `[ Refactor ][ Device ][ x/N ] ... to hardware
profiles` series (vllm-project#14076, vllm-project#15256, vllm-project#15376, vllm-project#15407, vllm-project#15478).

That series migrated most `is_ 310p() ` call sites to semantic
hardware-profile capabilities, but three call sites in the v2 worker
runtime path were left behind, and the ` is _310p()` compatibility
helper itself was kept alive. This PR finishes the migration and removes
the helper.

Residual call sites removed:

- `vllm_ ascend/patch/worker/patch _v2/patch_ block _table.py` — selects
`Ascend310PBlockTables` on 310P
- `vllm_ ascend/worker/v2/model _states/ __init__ .py` — selects the
Triton-free 310P `ModelState` (2 sites)
- `vllm_ ascend/patch/platform/patch _use_ v2 _model_ runner.py` — 310P
skips the upstream v2 model runner validation

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: spoon1116 <1522707055@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:core module:ops module:tests ready-all run all e2e test for pr ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants