Skip to content

[Misc]feat: adapt to vLLM main (ba94a3b9) - #10807

Closed
wjunLu wants to merge 1 commit into
vllm-project:mainfrom
wjunLu:main2main_auto_2026-06-22_14-01
Closed

wjunLu wants to merge 1 commit into
vllm-project:mainfrom
wjunLu:main2main_auto_2026-06-22_14-01

Conversation

@wjunLu

@wjunLu wjunLu commented Jun 22, 2026 •

Copy link
Copy Markdown
Collaborator

Cumulative Adaptation Summary

Step step-1

Verdict: No-op (no vllm-ascend changes required)

Upstream changes:

  • vllm/model_executor/models/gemma4_mm.py — Model-specific get_served_model_name usage
  • vllm/v1/sample/ops/topk_topp_triton.py — XPU block-size tuning (internal)
  • vllm/v1/worker/cpu_worker.py — CPU worker library check improvements

Why no adaptation:

  • Gemma4: No vllm-ascend model, patch, or reference
  • Triton kernel: vllm-ascend does not call or override apply_top_k_top_p_triton; imports TopKTopPSampler wrapper only
  • CPU worker: No vllm-ascend CPUWorker subclass or reference

Step step-2

Verdict: No-op (no vllm-ascend changes required)

Upstream changes:

  • vllm/config/kv_transfer.py — Docstring-only change removing P2pNcclConnector reference
  • vllm/distributed/kv_transfer/kv_connector/factory.py — Removed P2pNcclConnector registration
  • vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_common.py — Comment-only change
  • vllm/distributed/kv_transfer/kv_connector/v1/p2p/ — Entire p2p directory deleted (P2pNcclConnector, P2pNcclEngine, TensorMemoryPool)

Why no adaptation:

  • vllm-ascend has zero references to P2pNcclConnector, P2pNcclEngine, TensorMemoryPool, or the v1.p2p module
  • vllm-ascend's KV connector registrations in vllm_ascend/distributed/kv_transfer/__init__.py are entirely Ascend-native (Mooncake, AscendStore, UCM, LMCacheAscend) and unaffected by the upstream removal

Step step-3

Verdict: No-op (no vllm-ascend changes required)

Upstream changes:

  • vllm/_custom_ops.py — Internal refactor: cache release_dnnl_matmul_handler in self.dtor
  • vllm/benchmarks/serve.py — New _align_prompts_to_server_tokenizer() prompt alignment function
  • vllm/config/quantization.py — Add kFp8StaticChannelSym import, "fp8_per_channel_static" key, "fp8_per_channel" online shorthand
  • vllm/model_executor/layers/quantization/__init__.py — Add "fp8_per_channel" to QuantizationMethods Literal
  • vllm/model_executor/layers/quantization/online/base.py — Import/register Fp8PtpcOnlineLinearMethod and Fp8PtpcOnlineMoEMethod
  • vllm/model_executor/layers/quantization/online/fp8.py — New Fp8PtpcOnlineLinearMethod and Fp8PtpcOnlineMoEMethod classes; extend _Fp8OnlineMoEBase.__init__ with optional params (weight_key, activation_key, allow_vllm_cutlass, all with defaults preserving previous behavior); add per_act_token_quant/per_out_ch_quant class attrs

Why no adaptation:

  • vllm-ascend does not import from vllm.model_executor.layers.quantization.online (base or fp8) — the online dispatch system is entirely separate from Ascend's quantization configs

  • vllm-ascend does not import from vllm/config/quantization.py — uses register_quantization_config from the layers.quantization package instead

  • vllm-ascend does not subclass _Fp8OnlineMoEBase or any online quantization method — Ascend has its own AscendLinearMethod/AscendFusedMoEMethod wrappers

  • vllm-ascend has zero references to kFp8StaticChannelSym, _ONLINE_LINEAR_METHODS, _ONLINE_MOE_METHODS, _ONLINE_SHORTHANDS, or CPUDNNLGEMMHandler

  • The _Fp8OnlineMoEBase.__init__ signature changes are backward-compatible (all new params have defaults matching previous behavior)

  • All other changes are purely additive (new classes, new dict/Literal entries)

  • vLLM version: v0.22.1

  • vLLM main: vllm-project/vllm@967c5c3

Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request adapts the vllm-ascend codebase to significant architectural changes in the upstream vLLM repository. The primary focus is the refactoring of the FusedMoE layer into the new RoutedExperts and MoERunner components. To maintain support for Ascend-specific features, the PR implements targeted monkey-patches for the MoERunner and the FusedMoE factory, ensuring that existing Ascend optimizations are preserved while remaining compatible with the updated upstream API.

Highlights

  • Upstream Adaptation: Migrated the MoE layer architecture to align with the upstream vLLM refactor, transitioning from FusedMoE to the new RoutedExperts and MoERunner structure.
  • Compatibility Patching: Introduced monkey-patching utilities in vllm_ascend/utils.py to ensure Ascend-specific optimizations remain functional within the new upstream factory and runner patterns.
  • Type System Updates: Updated isinstance checks and import references across quantization and model files to correctly target RoutedExperts instead of the legacy FusedMoE class.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Misc] Adapt Ascend FusedMoE to upstream RoutedExperts refactoring

Suggested PR Summary:

### What this PR does / why we need it?
This PR adapts the Ascend FusedMoE implementation to the upstream refactoring that introduced `RoutedExperts` and `MoERunner`. It updates `AscendFusedMoE` and `AscendFusedMoE310` to inherit from `RoutedExperts` and monkey-patches the `FusedMoE` factory and `MoERunner` to ensure compatibility.

Feedback has been provided to address potential issues:
- Initialize `self.vllm_config` in `AscendFusedMoE310` to prevent `AttributeError` on 310P.
- Add the `forward` method override to `AscendFusedMoE310` to support legacy callers.
- Filter out duplicate parameters in `_patched_named_parameters` to avoid yielding duplicates.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
CI testing.

Comment on lines +182 to 184
# Replace quant_method with Ascend 310P version
assert self.quant_method is not None
# Keep base_quant_method aligned with the Ascend-replaced quant_method
# so FusedMoE.maybe_init_modular_kernel doesn't dispatch into the
# upstream UnquantizedFusedMoEMethod.maybe_make_prepare_finalize.
self.base_quant_method = self.quant_method

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

self.vllm_config is not initialized in AscendFusedMoE310.init (unlike AscendFusedMoE which retrieves it via get_current_vllm_config()). Since RoutedExperts.init does not set self.vllm_config, any downstream access to self.vllm_config on 310P will raise an AttributeError.\n\nWe should initialize self.vllm_config using get_current_vllm_config().

Suggested change
# Replace quant_method with Ascend 310P version
assert self.quant_method is not None
# Keep base_quant_method aligned with the Ascend-replaced quant_method
# so FusedMoE.maybe_init_modular_kernel doesn't dispatch into the
# upstream UnquantizedFusedMoEMethod.maybe_make_prepare_finalize.
self.base_quant_method = self.quant_method
# Replace quant_method with Ascend 310P version\n assert self.quant_method is not None\n self.base_quant_method = self.quant_method\n\n from vllm.config import get_current_vllm_config\n self.vllm_config = get_current_vllm_config()

Comment on lines 231 to 234
@property
def is_internal_router(self) -> bool:
# 310P Ascend path expects router logits from the model forward path.
return False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Unlike AscendFusedMoE, AscendFusedMoE310 is missing the forward method override. Legacy callers that still call forward() directly on the routed experts layer instead of the runner on 310P will bypass the custom _runner / forward_impl path or fail.\n\nWe should add the forward method override to AscendFusedMoE310 to match AscendFusedMoE.

    @property\n    def is_internal_router(self) -> bool:\n        # 310P Ascend path expects router logits from the model forward path.\n        return False\n\n    def forward(\n        self,\n        hidden_states: torch.Tensor,\n        router_logits: torch.Tensor,\n    ) -> torch.Tensor | tuple[torch.Tensor, torch.Tensor]:\n        if hasattr(self, \"_runner\"):\n            return self._runner.forward(hidden_states, router_logits)\n        self._ensure_moe_quant_config_init()\n        return self.forward_impl(hidden_states, router_logits)

Comment thread vllm_ascend/utils.py
Comment on lines +668 to +677
def _patched_named_parameters(self, prefix="", recurse=True, remove_duplicate=True):
if recurse and "_modules" in self.__dict__ and "routed_experts" in self._modules:
routed_experts = self._modules["routed_experts"]
for name, param in routed_experts.named_parameters():
yield prefix + name, param
yield from _original_named_params(
self, prefix=prefix, recurse=recurse, remove_duplicate=remove_duplicate
)

moe_runner_mod.MoERunner.named_parameters = _patched_named_parameters

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The current implementation of _patched_named_parameters yields parameters from routed_experts first, and then calls _original_named_params which will yield the same parameters again with the routed_experts. prefix. Because the two generator calls do not share a parameter memoization set, duplicate parameter tensors will be yielded. This can cause issues (such as ValueError in PyTorch optimizers or duplicate counting in parameter logging/profiling).\n\nWe should filter out the duplicate routed_experts. parameters from the original generator's output.

    def _patched_named_parameters(self, prefix=\"\", recurse=True, remove_duplicate=True):\n        if recurse and \"_modules\" in self.__dict__ and \"routed_experts\" in self._modules:\n            routed_experts = self._modules[\"routed_experts\"]\n            for name, param in routed_experts.named_parameters():\n                yield prefix + name, param\n        for name, param in _original_named_params(\n            self, prefix=prefix, recurse=recurse, remove_duplicate=remove_duplicate\n        ):\n            if recurse and \"routed_experts.\" in name:\n                continue\n            yield name, param

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

def _forward_impl(
self,
layer: torch.nn.Module,
hidden_states: torch.Tensor,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FusedMoE -> RoutedExperts 重构正确性确认

将 FusedMoE 替换为 RoutedExperts 是跟随上游 vLLM API 变迁的必要适配。代码变更量大,需要关注以下几点:

  1. AscendFusedMoE 和 AscendFusedMoE310 的 __init__ 签名:从 *args, **kwargs 改为显式参数列表,这是好的改进(提高可读性)。但新增了大量 Ascend-specific kwargs(如 tid2eid, gate, shared_experts 等),这些是通过 routed_experts_args 传入的。请确认上游 RoutedExperts.__init__ 是否会正确忽略这些 Ascend-specific kwargs,避免 unexpected keyword argument 错误。

  2. _get_quant_method 方法:AscendFusedMoE 新增了 _get_quant_method 方法覆盖父类的 quant method 初始化逻辑。请确认这与上游 RoutedExperts 的 _ensure_moe_quant_config_init 机制是否兼容,避免 quant_method 被重复初始化或覆盖。

@vllm-ascend-ci

Copy link
Copy Markdown
Collaborator

Superseded by #12420.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants