Skip to content

[310p] Qwen3.5 moe support - #7346

Closed
Tflowers-0129 wants to merge 4 commits into
vllm-project:mainfrom
Tflowers-0129:next-new
Closed

Tflowers-0129 wants to merge 4 commits into
vllm-project:mainfrom
Tflowers-0129:next-new

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Mar 16, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Does this PR introduce any user-facing change?

How was this patch tested?

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands the capabilities of the vLLM-Ascend project by integrating Qwen3.5 models, particularly their Gated Delta Network attention, onto the Ascend 310P architecture. The changes focus on providing robust PyTorch-based fallback implementations for key operations, optimizing KV cache handling, and enhancing the stability of paged attention through improved diagnostics. These updates ensure that Qwen3.5 models can leverage the 310P platform effectively, despite the absence of Triton support.

Highlights

  • Qwen3.5 MoE Support for Ascend 310P: Introduced comprehensive support for Qwen3.5 models, specifically the Gated Delta Network (GDN) attention mechanism, on the Ascend 310P platform. This includes PyTorch fallback implementations for various GDN-related operations, as Triton is not supported on 310P.
  • Optimized KV Cache Management: Enhanced the 310P model runner with a new function to select logical kernel block sizes for paged attention, improving KV cache initialization and overall memory efficiency.
  • Improved Paged Attention Robustness: Added detailed logging and preflight checks for paged attention operations, which will aid in debugging and ensure more robust execution on 310P.
  • MoE Runner Initialization Fix: Corrected an issue in the fused MoE module where the runner was initialized too early, ensuring proper selection of the moe_forward_shared function for shared experts.
  • New Unit Tests: Added a new unit test file to validate the functionality of the 310P attention kernel block size selection.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Changelog
  • tests/ut/_310p/test_model_runner_310p.py
    • Added unit tests for 310P attention kernel block size selection.
  • vllm_ascend/_310p/attention/attention_v1.py
    • Imported necessary modules like os, torch, and init_logger.
    • Added logging utilities (_is_rank0_process, _pa310_debug_enabled, _tensor_meta) for better debugging.
    • Integrated preflight checks and error handling into forward_paged_attention and forward_chunked_prefill_310.
  • vllm_ascend/_310p/fused_moe/fused_moe.py
    • Recreated the runner after shared experts are set to correctly select moe_forward_shared.
    • Modified _forward_shared_experts to align with the vLLM interface for shared experts.
  • vllm_ascend/_310p/model_runner_310p.py
    • Imported math, get_dtype_size, and NPUInputBatch.
    • Defined constants PAGED_ATTENTION_SPLIT_BLOCK_SIZE_310P and PAGED_ATTENTION_HEAD_BLOCK_PRODUCT_LIMIT_310P.
    • Implemented get_310p_attention_kernel_block_sizes for logical kernel block size selection.
    • Refactored KV cache initialization into _allocate_kv_cache_tensors and _reshape_kv_cache_tensors.
    • Added may_reinitialize_input_batch for dynamic input batch reinitialization.
    • Removed comments and simplified logic in _determine_batch_execution_and_padding and _pad_query_start_loc_for_fia.
    • Added _update_hybrid_attention_mamba_layout to skip hybrid layout on 310P.
  • vllm_ascend/_310p/ops/init.py
    • Added initialization for 310P-specific operations, including GDN attention ops.
  • vllm_ascend/_310p/ops/causal_conv1d.py
    • Added PyTorch reference implementations for causal_conv1d_ref_pytorch, causal_conv1d_fn_pytorch, and causal_conv1d_update_pytorch for causal convolution.
  • vllm_ascend/_310p/ops/delta_rule.py
    • Added PyTorch reference implementations for _run_recurrent_gated_delta_rule, chunk_gated_delta_rule_pytorch, fused_recurrent_gated_delta_rule_pytorch, and fused_sigmoid_gating_delta_rule_update_pytorch for Gated Delta Network operations.
  • vllm_ascend/_310p/ops/gdn_attention.py
    • Added gdn_attention_core_impl for the PyTorch implementation of GDN attention core.
    • Added register_gdn_attention_ops to register the custom GDN attention ops.
  • vllm_ascend/_310p/ops/gdn_gating.py
    • Added fused_gdn_gating_pytorch for PyTorch implementation of fused GDN gating.
  • vllm_ascend/_310p/ops/layernorm.py
    • Imported RMSNormGated and added AscendRMSNormGated310 to reuse the upstream native implementation for 310P.
  • vllm_ascend/_310p/ops/qkvzba_split_reshape.py
    • Added fused_qkvzba_split_reshape_cat_pytorch for PyTorch implementation of QKVZBA split, reshape, and concatenate.
  • vllm_ascend/_310p/patch/init.py
    • Imported vllm_ascend._310p.patch.patch_qwen3_next.
  • vllm_ascend/_310p/patch/patch_qwen3_next.py
    • Defined Ascend310Qwen3Next_GatedDeltaNet to override Qwen3NextGatedDeltaNet.forward and _forward_core with PyTorch fallback implementations for 310P.
  • vllm_ascend/_310p/worker_310p.py
    • Imported HAS_TRITON and added a logger info message for 310P Worker initialization.
    • Added a placeholder _load_310p_patches method.
  • vllm_ascend/patch/worker/init.py
    • Imported is_310p and conditionally imported patch_qwen3_next_310 for 310P.
    • Removed direct imports of patch_qwen3_next, patch_qwen3_next_mtp, patch_qwen3_5.
  • vllm_ascend/patch/worker/patch_qwen3_next_310.py
    • Added a new patch file to reuse the 310P fallback implementation for Qwen3Next/Qwen3.5.
  • vllm_ascend/utils.py
    • Imported AscendRMSNormGated310 and registered it in _310p_custom_ops.
Activity
  • Tflowers-0129 created this pull request to add Qwen3.5 MoE support for the Ascend 310P platform.
  • The pull request introduces new PyTorch-based implementations for Gated Delta Network (GDN) attention operations, as Triton is not available on 310P.
  • The KV cache management logic has been updated to include optimized kernel block size selection for 310P.
  • Logging and preflight checks have been added to paged attention for improved debugging and stability.
  • A fix was implemented in the fused MoE module to ensure correct runner initialization for shared experts.
  • New unit tests were added to cover the attention kernel block size selection logic.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces 310P-specific implementations and fallbacks for various operations, including attention, MoE, and layernorm, to address the lack of Triton support on 310P. The changes are well-structured, with clear comments explaining the rationale behind the adaptations. New test cases for get_310p_attention_kernel_block_sizes enhance coverage, and improved logging and preflight checks contribute to robustness. However, a critical issue was identified in the causal_conv1d_fn_pytorch function regarding inconsistent tensor dimension handling, which could lead to incorrect calculations or runtime errors.

Comment on lines +119 to +129
if x.dim() == 3:
if x.shape[0] == 1:
x = x.squeeze(0)
elif x.shape[1] == 1:
x = x.squeeze(1).transpose(0, 1)
else:
raise RuntimeError(f"Unsupported x shape for causal_conv1d_fn_pytorch: {tuple(x.shape)}")
if x.dim() != 2:
raise RuntimeError(f"Unsupported x ndim for causal_conv1d_fn_pytorch: {x.dim()}")

feature_dim = x.shape[0]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The causal_conv1d_fn_pytorch function's docstring and internal checks indicate that the input x should always be 2D, specifically (dim, cu_seq_len). However, the code block from lines 119-125 attempts to handle a 3D input, which is inconsistent with the function's expected behavior and the query_start_loc check at line 105. Furthermore, the feature_dim calculation at line 129 (x.shape[1]) is incorrect if x is (dim, cu_seq_len), as it would yield cu_seq_len instead of dim. This discrepancy can lead to incorrect dimension matching with weight.shape[0] and runtime errors.

    # Normalize x to [dim, total_tokens]
    # The docstring and query_start_loc check imply x is always 2D (dim, cu_seq_len).
    if x.dim() != 2:
        raise RuntimeError(f"Unsupported x ndim for causal_conv1d_fn_pytorch: {x.dim()}, expected 2D (dim, total_tokens).")

    feature_dim = x.shape[0]

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant