Skip to content

[BugFix][310p]Fix Triton kernel in block_table crash that caused failures - #8085

Merged
MengqingCao merged 4 commits into
vllm-project:mainfrom
Tflowers-0129:fix_triton_kernel_bugs
Apr 9, 2026
Merged

MengqingCao merged 4 commits into
vllm-project:mainfrom
Tflowers-0129:fix_triton_kernel_bugs

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Apr 9, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

310P devices do not support Triton kernels. Due to incomplete CI coverage, the 310P execution path can incorrectly enter Triton kernels, causing runtime crashes on 310P. This PR fixes the issue.

Does this PR introduce any user-facing change?

NA

How was this patch tested?

ut and e2e tests

…ures running on 310p

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a critical runtime crash occurring on 310P devices caused by the incorrect invocation of Triton kernels. By introducing a dedicated block table management system for 310P, the PR ensures that the execution path remains compatible with the hardware's limitations, effectively bypassing unsupported kernel calls while maintaining correct KV cache slot mapping.

Highlights

  • Triton Kernel Avoidance: Implemented a custom BlockTable for 310P devices to prevent the execution path from incorrectly invoking Triton kernels, which are not supported on this hardware.
  • Infrastructure Update: Added a new MultiGroupBlockTable specifically for 310P to manage KV cache block mappings correctly without relying on Triton-based implementations.
  • Testing: Added comprehensive unit tests to verify the slot mapping logic for the 310P block table implementation.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/pr_test_light.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

Signed-off-by: Tflowers-0129 <2906339855@qq.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new block table management system for 310P hardware, comprising BlockTable and MultiGroupBlockTable classes to handle physical/logical block mapping, distributed processing, and KV cache interleaving. The NPUModelRunner310 is updated to integrate this new system, and comprehensive unit tests are added for the BlockTable's compute_slot_mapping method. Feedback includes a high-severity suggestion to optimize the _convert_physical_to_logical_blocks method using NumPy broadcasting for improved efficiency, and another high-severity comment addressing a potential bug in MultiGroupBlockTable's constructor due to shallow copying of mutable lists.

I am having trouble creating individual review comments. Click here to see my feedback.

vllm_ascend/_310p/block_table.py (203-208)

high

The current implementation for converting physical to logical blocks uses a Python loop and list appends, which can be inefficient for large arrays. You can achieve the same result more efficiently using NumPy broadcasting, which is significantly faster as it operates in compiled C code.

        # Reshape for broadcasting
        physical_blocks_col = physical_blocks.reshape(-1, 1)
        # Create offsets for each physical block
        offsets = np.arange(self.blocks_per_phys_block, dtype=np.int32)
        # Broadcast and flatten
        return (physical_blocks_col * self.blocks_per_phys_block + offsets).ravel()

vllm_ascend/_310p/block_table.py (239-246)

high

Using * to multiply a list containing mutable objects (like lists) creates shallow copies. This means that all elements in the resulting list will reference the same object.

For example:

  • [[0]] * len(block_sizes) on line 240 creates a list where all elements are the same [0] list object.
  • kernel_sizes * len(block_sizes) on line 242 does the same if kernel_sizes is [[...]].

If any of these inner lists are modified later, it will affect all of them, leading to unexpected bugs.

To prevent this, you should use list comprehensions to create new list objects for each element.

        if kernel_sizes is None:
            kernel_sizes = [[0] for _ in range(len(block_sizes))]
        elif len(kernel_sizes) == 1 and len(block_sizes) > 1:
            kernel_sizes = [list(kernel_sizes[0]) for _ in range(len(block_sizes))]
        elif len(kernel_sizes) != len(block_sizes):
            raise ValueError(
                f"kernel_sizes length ({len(kernel_sizes)}) must match block_sizes length ({len(block_sizes)})"
            )

@github-actions

github-actions Bot commented Apr 9, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@Tflowers-0129
Tflowers-0129 force-pushed the fix_triton_kernel_bugs branch from adb2bcb to d97cd1e Compare April 9, 2026 08:44
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@Tflowers-0129 Tflowers-0129 changed the title [BugFix][310p]Fix Triton kernel in block_table crash that caused fail… [BugFix][310p]Fix Triton kernel in block_table crash that caused failures Apr 9, 2026
@MengqingCao
MengqingCao merged commit 4aacaf8 into vllm-project:main Apr 9, 2026
40 checks passed
paulyu12 pushed a commit to paulyu12/vllm-ascend that referenced this pull request Apr 14, 2026
…ures (vllm-project#8085)

### What this PR does / why we need it?
310P devices do not support Triton kernels. Due to incomplete CI
coverage, the 310P execution path can incorrectly enter Triton kernels,
causing runtime crashes on 310P. This PR fixes the issue.

- vLLM version: 
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.19.0
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
guxin108 pushed a commit to guxin108/vllm-ascend that referenced this pull request Apr 24, 2026
…ures (vllm-project#8085)

### What this PR does / why we need it?
310P devices do not support Triton kernels. Due to incomplete CI
coverage, the 310P execution path can incorrectly enter Triton kernels,
causing runtime crashes on 310P. This PR fixes the issue.

- vLLM version:
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.19.0
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: guxin108 <1252896542@qq.com>
zouyida2052 pushed a commit to zouyida2052/vllm-ascend that referenced this pull request Apr 28, 2026
…ures (vllm-project#8085)

### What this PR does / why we need it?
310P devices do not support Triton kernels. Due to incomplete CI
coverage, the 310P execution path can incorrectly enter Triton kernels,
causing runtime crashes on 310P. This PR fixes the issue.

- vLLM version:
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.19.0
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
yangzhe-2026 pushed a commit to yangzhe-2026/vllm-ascend that referenced this pull request May 6, 2026
…ures (vllm-project#8085)

### What this PR does / why we need it?
310P devices do not support Triton kernels. Due to incomplete CI
coverage, the 310P execution path can incorrectly enter Triton kernels,
causing runtime crashes on 310P. This PR fixes the issue.

- vLLM version: 
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.19.0
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
nanxingMy pushed a commit to nanxingMy/vllm-ascend that referenced this pull request May 15, 2026
…ures (vllm-project#8085)

### What this PR does / why we need it?
310P devices do not support Triton kernels. Due to incomplete CI
coverage, the 310P execution path can incorrectly enter Triton kernels,
causing runtime crashes on 310P. This PR fixes the issue.

- vLLM version: 
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.19.0
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: nanxing <1014662416@qq.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
…ures (vllm-project#8085)

### What this PR does / why we need it?
310P devices do not support Triton kernels. Due to incomplete CI
coverage, the 310P execution path can incorrectly enter Triton kernels,
causing runtime crashes on 310P. This PR fixes the issue.

- vLLM version: 
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.19.0
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
…ures (vllm-project#8085)

### What this PR does / why we need it?
310P devices do not support Triton kernels. Due to incomplete CI
coverage, the 310P execution path can incorrectly enter Triton kernels,
causing runtime crashes on 310P. This PR fixes the issue.

- vLLM version: 
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.19.0
---------
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants