Skip to content

[MRV2][Bugfix] Fix two triton ops errors in model_runner_v2 num_nans_kernel and apply_panalties - #13159

Merged
yiz-liu merged 6 commits into
vllm-project:mainfrom
AuroraEmiya:mrv2_ops
Aug 5, 2026
Merged

yiz-liu merged 6 commits into
vllm-project:mainfrom
AuroraEmiya:mrv2_ops

Conversation

@AuroraEmiya

@AuroraEmiya AuroraEmiya commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

This PR has done two things:

  1. This PR fix num_nans wrong import way of libdevice in Ascend env with a right way with CANN
    therefore we add new patch in vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
    metrics_logits.libdevice = triton.language.extra.cann.libdevice
  2. This PR also fix over-grid limit with large input in ops apply_palties

Does this PR introduce any user-facing change?

No

How was this patch tested?

uts are added in

  1. tests/ut/patch/worker/test_patch_v2_num_nans.py
  2. tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces significant improvements to the Triton-based penalty operations within the model runner. By splitting the monolithic penalty kernel into specialized kernels for vocabulary and draft tokens, the implementation achieves better efficiency and robustness. These changes ensure the system can handle larger shapes without crashing and align the penalty logic more closely with reference implementations, while also improving the underlying hardware-specific configuration.

Highlights

  • Triton Kernel Refactoring: Refactored the penalty application logic into two distinct kernels: a vocabulary-wide kernel for general penalty application and a draft-specific kernel for handling speculative token penalties.
  • Performance and Stability: Improved performance by optimizing memory access patterns and added support for larger shapes to prevent kernel crashes.
  • Infrastructure Updates: Added a new test case for verifying CANN libdevice usage and updated the penalty module to include hardware-aware vector core detection.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix]/[Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@AuroraEmiya AuroraEmiya changed the title [Feature][Bugfix] Fix two triton ops in model_runner [MRV2][Bugfix] Fix two triton ops in model_runner_v2 num_nans_kernel and apply_panalties and impove performance of apply_panalties as well Jul 30, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Implement optimized Triton penalty kernels

Suggested PR Summary:

### What this PR does / why we need it?
This pull request refactors and optimizes the penalty application logic by introducing two specialized Triton kernels: `_apply_penalties_vocab_kernel` for full-vocabulary penalties and `_apply_penalties_draft_kernel` for draft token corrections. It also updates the bincount kernel and integrates the CANN libdevice for Triton-Ascend.

Feedback:
Two critical issues were identified in the Triton kernels where out-of-bounds token IDs (such as negative padding tokens or values exceeding the vocabulary size) could lead to illegal memory access or out-of-bounds writes on the GPU. Guard conditions should be added to ensure `draft_token` and `prompt_tokens` are within valid bounds before indexing.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
The changes are tested using the newly added unit test `test_patch_v2_num_nans_uses_cann_libdevice_only` and the updated e2e test suite in `test_penality.py` including large shape handling.

earlier_count += (is_same & (scan_pos < candidate_pos)).to(tl.int32)

# Only the first occurrence owns this logits element.
if earlier_count == 0:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The draft_token loaded from token_ids_ptr is used directly as an index to load from output_counts_row_ptr and logits_row_ptr, and to store back to logits_row_ptr. If draft_token is out of bounds (e.g., negative padding tokens like -1 or values exceeding the vocabulary size), this will cause an illegal memory access on the GPU, leading to a kernel crash. We should guard this by checking that draft_token is within [0, vocab_size). Since logits_stride is equal to vocab_size for the 2D logits tensor, we can use logits_stride as the upper bound.

Suggested change
if earlier_count == 0:
if (earlier_count == 0) & (draft_token >= 0) & (draft_token < logits_stride):

Comment on lines 308 to 321
if block_idx * BLOCK_SIZE < prompt_len:
mask = block < prompt_len
prompt_tokens = tl.load(all_token_ids_ptr + req_state_idx * all_token_ids_stride + block, mask=mask)
idx = prompt_tokens // 32

bit_idx = prompt_tokens % 32
bit = tl.full((BLOCK_SIZE,), 1, tl.int32) << bit_idx

prompt_tokens = tl.load(
all_token_ids_ptr + req_state_idx * all_token_ids_stride + block,
mask=mask,
)
packed_indices = prompt_tokens // 32
bit_indices = prompt_tokens % 32
bits = tl.full((BLOCK_SIZE,), 1, tl.int32) << bit_indices
tl.atomic_or(
prompt_bin_mask_ptr + req_state_idx * prompt_bin_mask_stride + idx,
bit,
prompt_bin_mask_ptr + req_state_idx * prompt_bin_mask_stride + packed_indices,
bits,
mask=mask,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

In _bincount_kernel, prompt_tokens loaded from all_token_ids_ptr can contain out-of-bounds values (such as negative padding tokens like -1 or values exceeding the vocabulary size). Using these directly to compute packed_indices and performing tl.atomic_or will result in out-of-bounds memory writes and GPU crashes. We should guard the mask with a check that prompt_tokens is within [0, vocab_size). Since output_bin_counts_stride is equal to vocab_size, we can use it as the upper bound.

Suggested change
if block_idx * BLOCK_SIZE < prompt_len:
mask = block < prompt_len
prompt_tokens = tl.load(all_token_ids_ptr + req_state_idx * all_token_ids_stride + block, mask=mask)
idx = prompt_tokens // 32
bit_idx = prompt_tokens % 32
bit = tl.full((BLOCK_SIZE,), 1, tl.int32) << bit_idx
prompt_tokens = tl.load(
all_token_ids_ptr + req_state_idx * all_token_ids_stride + block,
mask=mask,
)
packed_indices = prompt_tokens // 32
bit_indices = prompt_tokens % 32
bits = tl.full((BLOCK_SIZE,), 1, tl.int32) << bit_indices
tl.atomic_or(
prompt_bin_mask_ptr + req_state_idx * prompt_bin_mask_stride + idx,
bit,
prompt_bin_mask_ptr + req_state_idx * prompt_bin_mask_stride + packed_indices,
bits,
mask=mask,
)
if block_idx * BLOCK_SIZE < prompt_len:
mask = block < prompt_len
prompt_tokens = tl.load(
all_token_ids_ptr + req_state_idx * all_token_ids_stride + block,
mask=mask,
)
mask = mask & (prompt_tokens >= 0) & (prompt_tokens < output_bin_counts_stride)
packed_indices = prompt_tokens // 32
bit_indices = prompt_tokens % 32
bits = tl.full((BLOCK_SIZE,), 1, tl.int32) << bit_indices
tl.atomic_or(
prompt_bin_mask_ptr + req_state_idx * prompt_bin_mask_stride + packed_indices,
bits,
mask=mask,
)

@yiz-liu yiz-liu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please see if you need add explanation for the new patch.

@AuroraEmiya

Copy link
Copy Markdown
Contributor Author

Please see if you need add explanation for the new patch.

That 's fair. We have added here

What this PR does / why we need it?

This PR has done two things:

  1. This PR fix num_nans wrong import way of libdevice in Ascend env with a right way with CANN
    therefore we add new patch in vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
    metrics_logits.libdevice = triton.language.extra.cann.libdevice
  2. This PR also fix over-grid limit with large input in ops apply_palties and refactor the kernel gaining a 80x speedup

Does this PR introduce any user-facing change?

No

How was this patch tested?

uts are added in

  1. tests/ut/patch/worker/test_patch_v2_num_nans.py
  2. tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py

@github-actions

github-actions Bot commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Liam and others added 3 commits August 3, 2026 10:04
Signed-off-by: Liam <ml646@duke.edu>
…e shape and introduce a new kernel to speedup 80x

Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
…o only focus on known failure. Grid over-65535 error remains. The improvement of performance should be added later when the version and implementation are stable.

Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
@AuroraEmiya AuroraEmiya changed the title [MRV2][Bugfix] Fix two triton ops in model_runner_v2 num_nans_kernel and apply_panalties and impove performance of apply_panalties as well [MRV2][Bugfix] Fix two triton ops errors in model_runner_v2 num_nans_kernel and apply_panalties Aug 3, 2026
@yiz-liu yiz-liu added the ready label Aug 4, 2026
Liam and others added 2 commits August 5, 2026 12:00
Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
@yiz-liu
yiz-liu enabled auto-merge (squash) August 5, 2026 06:31
@yiz-liu

yiz-liu commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@yiz-liu
yiz-liu merged commit 61cfd1f into vllm-project:main Aug 5, 2026
45 checks passed
shenhui-cli added a commit to shenhui-cli/vllm-ascend that referenced this pull request Aug 5, 2026
…into vllm-new

# By shenhui-cli (8) and others
# Via GitHub (1) and shenhui-cli (1)
* 'vllm-new' of https://github.com/shenhui-cli/vllm-ascend: (34 commits)
  When the PR modifies any files in the CSRC folder, skip the test case filtering logic and execute all test cases by default.
  [MRV2][Bugfix] Fix two triton ops errors in model_runner_v2 num_nans_kernel and apply_panalties (vllm-project#13159)
  [Cherry-pick][main][Doc][BugFix] Update proxy script name in DeepSeek-V3.2 tutorial (from vllm-project#13537) (vllm-project#13575)
  [Refactor][quantization]remove mxfp_compat compatibility shim and inline dtype references (vllm-project#13447)
  [MTP][BugFix] Preserve MoE weight loaders during online updates (vllm-project#13337)
  [Feature][xlite] Support MLA and DSA (`DeepseekV3ForCausalLM`, `DeepseekV32ForCausalLM` and `GlmMoeDsaForCausalLM`) in the xlite adapter (vllm-project#13378)
  [BugFix][xlite] Fix DP metadata handling in XliteWrapper and pass down the max number of tokens across all DP ranks for collective communications. (vllm-project#11213)
  [Refactor][Ops] Move expert routing into router classes (vllm-project#13417)
  [Platform][Refactor] Refactor NPUPlatform for better organization and clarity (vllm-project#13484)
  [BugFix] fix fiaV2 contiguous err in GQA (vllm-project#13456)
  [feature][KV Offload] Support Sparse KV Cache Offload (vllm-project#13026)
  [Bugfix][MRV2]Skip D2H copy and synchronize when spec decode is not active (vllm-project#13382)
  [Bugfix] Change AscendSFAIndexerCacheSpec to inherit MLAAttentionSpec (vllm-project#12849)
  [BugFix][FusedMoE] Restore BF16 quant method initialization (vllm-project#13412)
  [Doc] Fix link errors and add section anchors (vllm-project#13485)
  [TEST]Revise the A3 case (vllm-project#13495)
  [CI] modify default cann_version and add build_type support in nightly (vllm-project#13480)
  [CI]Improve logging, optimize function-level recommendation algorithm, and add early exit for no product code changes (vllm-project#13472)
  This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/
  This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/
  ...

# Conflicts:
#	.github/workflows/scripts/test_selector.py
Wyz-134 pushed a commit to Wyz-134/vllm-ascend that referenced this pull request Aug 6, 2026
…kernel and apply_panalties (vllm-project#13159)

### What this PR does / why we need it?
This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
### Does this PR introduce _any_ user-facing change?
No
### How was this patch tested?
uts are added in 
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py


- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>
Wyz-134 pushed a commit to Wyz-134/vllm-ascend that referenced this pull request Aug 6, 2026
…kernel and apply_panalties (vllm-project#13159)

### What this PR does / why we need it?
This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
### Does this PR introduce _any_ user-facing change?
No
### How was this patch tested?
uts are added in 
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py


- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>

Signed-off-by: AuroraEmiya <92282919+AuroraEmiya@users.noreply.github.com>
HMCCMH pushed a commit to hotTea123/vllm-ascend that referenced this pull request Aug 12, 2026
…kernel and apply_panalties (vllm-project#13159)

### What this PR does / why we need it?
This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
### Does this PR introduce _any_ user-facing change?
No
### How was this patch tested?
uts are added in 
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py


- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>
MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
…kernel and apply_panalties (vllm-project#13159)

### What this PR does / why we need it?
This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
### Does this PR introduce _any_ user-facing change?
No
### How was this patch tested?
uts are added in 
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py


- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
…kernel and apply_panalties (vllm-project#13159)

### What this PR does / why we need it?
This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
### Does this PR introduce _any_ user-facing change?
No
### How was this patch tested?
uts are added in 
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py


- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 20, 2026
…kernel and apply_panalties (vllm-project#13159)

This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
No
uts are added in
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>
(cherry picked from commit 61cfd1f)
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 20, 2026
…kernel and apply_panalties (vllm-project#13159)

This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
No
uts are added in
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>
(cherry picked from commit 61cfd1f)
Signed-off-by: jiaqi-lee <15316070896@163.com>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 23, 2026
…kernel and apply_panalties (vllm-project#13159)

This PR has done two things:
1. This PR fix num_nans wrong import way of libdevice in Ascend env with
a right way with CANN
therefore we add new patch in
vllm_ascend/patch/worker/patch_v2/patch_triton.py‎
`metrics_logits.libdevice = triton.language.extra.cann.libdevice`
3. This PR also fix over-grid limit with large input in ops
apply_palties
No
uts are added in
1. tests/ut/patch/worker/test_patch_v2_num_nans.py
2.
tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_penality.py

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

---------

Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: Liam <ml646@duke.edu>
(cherry picked from commit 61cfd1f)
Signed-off-by: jiaqi-lee <15316070896@163.com>
kunpengW-code pushed a commit that referenced this pull request Aug 24, 2026
…#14687)

### What this PR does / why we need it?

This is the v0.26.0 release-blocker backport rollup. It contains only
the 19 audited high-severity correctness and stability fixes that were
still missing from releases/v0.26.0rc at 1f95052.

The branch has 23 physical commits because #14619 is preserved as its
complete five-commit atomic series. Cross-branch equivalents are
deduplicated, and every logical fix remains independently reviewable and
revertible.

#### Included fixes

| # | Source / target PR | Severity | Problem fixed |
|---:|---|---|---|
| 1 | #13538 | P0 correctness | Qwen3-VL MoE + FlashComm1 + deepstack
used the wrong residual tensor and could silently corrupt output. |
| 2 | #13600 | P0 deadlock | MRV1/MRV2 main and draft update streams
could mutually wait during repeated full-graph execution. |
| 3 | #13498 | P0 data corruption | Float32 Mamba state could overwrite
the shared bf16 hidden-state cache buffer. |
| 4 | #13902 | P0 correctness | RL weight reload left ACL graphs
referencing stale W8A8-MXFP8 weight addresses. |
| 5 | #12359 / #12371 | P0 correctness | Mooncake reformatted KV before
all TP/CP pulls for a request completed, causing TP inequality or
reordered KV. |
| 6 | #13111 / #13113 / #13110 | P1 KV correctness | Multi-KV-group
save/load incorrectly reused group-0 block size for every group. |
| 7 | #13116 / #13117 / #13099 | P1 state consistency | Async KV load
failures were not shared with the scheduler, preventing recompute
recovery. |
| 8 | #13308 / #13310 / #13307 | P1 crash | Memcache batch
query/allocation before lazy initialization could assert in scheduler or
worker. |
| 9 | #13012 | P1 hang/corruption | Level-2 sleep/wake could lose the
MoE loader and leave EPLB tensors pointing at released storage. |
| 10 | #13414 | P1 crash/correctness | Dynamic EPLB initialized W8A8
scales for only the first expert weight. |
| 11 | #14001 | P1 crash | MiniMax-M3 index_q was reshaped using total
size instead of the per-head dimension. |
| 12 | #14394 | P1 crash/hang | MRV2 FULL_DECODE_ONLY dropped graph
padding when runtime mode was FULL. |
| 13 | #13136 | P1 crash | P/D + DP zero-token ranks compared None with
MC2 capacity and raised TypeError. |
| 14 | #13183 | P1 unavailable | ec_both was treated as producer-only
and skipped KV specification/data needed by its consumer role. |
| 15 | #13123 | P1 OOB/device error | MRV2 dummy-token remainder was
concentrated on one request and could exceed max_model_len. |
| 16 | #13159 | P1 crash | MRV2 num_nans used the wrong Triton libdevice
and the penalty kernel could exceed the CANN grid limit. |
| 17 | #12940 | P1 crash | DFlash profiling used total query count
instead of actual input tokens for RoPE/graph capture. |
| 18 | #13394 via #13405 | P0 correctness | RL sampling tensor lifetime
errors could produce Inf/OOV tokens and contaminate later output. |
| 19 | #14142 via #14619 | P1 long-run/state correctness | P/D rejection
left stale KV/accounting and unsafe retry/replay behavior could leak,
duplicate, or return wrong responses. |

#### Backport policy

- Selected the audited v0.26 release-adapted commits where available;
the closed rollup #14337 was not revived wholesale.
- Kept only one canonical copy of fixes duplicated across 0.23, 0.25,
and main.
- Manually adapted #13136, the core #13123 input-batch hunk, and #13159
to preserve current v0.26/MegaMoe/model-runner behavior.
- Used the current v0.26 target change from #13405 and the complete
five-commit #14619 series.
- Intentionally excluded performance-only, UX-only, conditional-support,
low-confidence, and owner-unsettled fixes from this release window.

### Does this PR introduce _any_ user-facing change?

Yes, behavior is corrected for the affected configurations: crashes,
deadlocks, hangs, incorrect output, stale KV state, and data corruption
are prevented. There is no new public API, CLI option, or configuration
requirement.

### How was this patch tested?

Local validation completed:

- Audited manifest: 19/19 logical fixes, 23/23 expected source commits;
missing 0, duplicate 0, unexpected 0.
- All 23 commits retain source provenance and Signed-off-by trailers.
- Ruff lint and format checks passed for all 38 changed Python files.
- Python syntax compilation passed for all 38 changed Python files.
- git diff --check, codespell, forbidden-import, package-init,
context-manager, and filename checks passed.
- The final worktree is clean at
86b2ca8.

The backports retain or add focused tests for AscendStore, Mooncake
rejection cleanup, fused MoE/EPLB, W8A8-MXFP8 reload, worker sleep/wake,
MRV2 graph padding, penalty-grid limits, hidden-state extraction, and
two-card speculative DP.

NPU UT/E2E was not run locally because the available Windows environment
has no vLLM, PyTorch/torch_npu, pytest, or Ascend device. CI and
targeted NPU regression are therefore required before merge, especially:

- MRV1/MRV2 full-graph repeated-iteration deadlock/teardown.
- Mooncake TP2/TP4 out-of-order pull KV equality and P/D rejection
cleanup.
- Qwen3-VL FlashComm1 + deepstack fixed-seed correctness.
- Level-2 sleep/wake, dynamic EPLB, and multi-round RL weight reload.
- P/D + DP zero-token ranks, ec_both, DFlash profile/graph, and RL
Inf/OOV sampling.
- Proxy retry and streaming replay behavior from #14619.

#### Review checklist

- [x] Only the 19 approved release-critical logical fixes are included.
- [x] One logical fix per commit; #14619 remains an atomic five-commit
series.
- [x] No performance-only backports are included.
- [x] Source provenance and sign-offs are retained.
- [ ] Repository CI passes.
- [ ] Targeted Ascend NPU correctness and long-run tests pass.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: kyle-zhangchi <chiiiiiizhang@gmail.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
Signed-off-by: jiaqi-lee <15316070896@163.com>
Signed-off-by: tyy0829 <1455207791@qq.com>
Signed-off-by: yejj710 <abyss1999@163.com>
Signed-off-by: jiajinzhu2 <jiajinzhu@huawei.com>
Signed-off-by: chenyue1122 <oyoy7102@163.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: muziyuhui666 <lijianfu9@huawei.com>
Signed-off-by: Pz1116 <zpbzpb123123@gmail.com>
Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
Signed-off-by: likailong <likailong5@huawei.com>
Signed-off-by: hanxi-java <634498162@qq.com>
Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Signed-off-by: HF-001 <1670186653@qq.com>
Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com>
Signed-off-by: Hcm03 <chengminhua1@huawei.com>
Signed-off-by: zhuyixiang <zhuyixiang2014@163.com>
Signed-off-by: moonseeker <2290166829@qq.com>
Co-authored-by: kyle-zhangchi <chiiiiiizhang@gmail.com>
Co-authored-by: tyy0829 <87685049+tyy0829@users.noreply.github.com>
Co-authored-by: yejj <abyss1999@163.com>
Co-authored-by: jiajinzhu2 <jiajinzhu@huawei.com>
Co-authored-by: CHENYUE <56943221+PHOEBEMOON0802@users.noreply.github.com>
Co-authored-by: Xu Rongsheng <73730571+MmMmaru@users.noreply.github.com>
Co-authored-by: yjyang62 <yangjinyang5@huawei.com>
Co-authored-by: muziyuhui666 <lijianfu9@huawei.com>
Co-authored-by: CXY-Katrina <katrina.cxy@gmail.com>
Co-authored-by: cywang250805 <wangchaoyu7@huawei.com>
Co-authored-by: Bill845514379 <huangjianbao2@huawei.com>
Co-authored-by: yejj710 <yejj710@gmail.com>
Co-authored-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: HaoxinZong <116423146+HaoxinZong@users.noreply.github.com>
Co-authored-by: pz1116 <47019764+Pz1116@users.noreply.github.com>
Co-authored-by: zouyida2052 <zouyida2002@gmail.com>
Co-authored-by: iKeybot <92210799+iKeybot-code@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: 韩熙 <63780107+hanxi-java@users.noreply.github.com>
Co-authored-by: zouzy <38661932+zouzy5137@users.noreply.github.com>
Co-authored-by: AuroraEmiya <92282919+AuroraEmiya@users.noreply.github.com>
Co-authored-by: Liam <ml646@duke.edu>
Co-authored-by: kx <1670186653@qq.com>
Co-authored-by: wangxiaoteng888 <56506195+wangxiaoteng888@users.noreply.github.com>
Co-authored-by: Hcm03 <chengminhua1@huawei.com>
Co-authored-by: zhuyixiang <zhuyixiang2014@163.com>
Co-authored-by: moonseeker <2290166829@qq.com>
Leetrytry pushed a commit to Leetrytry/vllm-ascend that referenced this pull request Sep 11, 2026
…vllm-project#14687)

### What this PR does / why we need it?

This is the v0.26.0 release-blocker backport rollup. It contains only
the 19 audited high-severity correctness and stability fixes that were
still missing from releases/v0.26.0rc at 1f95052.

The branch has 23 physical commits because vllm-project#14619 is preserved as its
complete five-commit atomic series. Cross-branch equivalents are
deduplicated, and every logical fix remains independently reviewable and
revertible.

#### Included fixes

| # | Source / target PR | Severity | Problem fixed |
|---:|---|---|---|
| 1 | vllm-project#13538 | P0 correctness | Qwen3-VL MoE + FlashComm1 + deepstack
used the wrong residual tensor and could silently corrupt output. |
| 2 | vllm-project#13600 | P0 deadlock | MRV1/MRV2 main and draft update streams
could mutually wait during repeated full-graph execution. |
| 3 | vllm-project#13498 | P0 data corruption | Float32 Mamba state could overwrite
the shared bf16 hidden-state cache buffer. |
| 4 | vllm-project#13902 | P0 correctness | RL weight reload left ACL graphs
referencing stale W8A8-MXFP8 weight addresses. |
| 5 | vllm-project#12359 / vllm-project#12371 | P0 correctness | Mooncake reformatted KV before
all TP/CP pulls for a request completed, causing TP inequality or
reordered KV. |
| 6 | vllm-project#13111 / vllm-project#13113 / vllm-project#13110 | P1 KV correctness | Multi-KV-group
save/load incorrectly reused group-0 block size for every group. |
| 7 | vllm-project#13116 / vllm-project#13117 / vllm-project#13099 | P1 state consistency | Async KV load
failures were not shared with the scheduler, preventing recompute
recovery. |
| 8 | vllm-project#13308 / vllm-project#13310 / vllm-project#13307 | P1 crash | Memcache batch
query/allocation before lazy initialization could assert in scheduler or
worker. |
| 9 | vllm-project#13012 | P1 hang/corruption | Level-2 sleep/wake could lose the
MoE loader and leave EPLB tensors pointing at released storage. |
| 10 | vllm-project#13414 | P1 crash/correctness | Dynamic EPLB initialized W8A8
scales for only the first expert weight. |
| 11 | vllm-project#14001 | P1 crash | MiniMax-M3 index_q was reshaped using total
size instead of the per-head dimension. |
| 12 | vllm-project#14394 | P1 crash/hang | MRV2 FULL_DECODE_ONLY dropped graph
padding when runtime mode was FULL. |
| 13 | vllm-project#13136 | P1 crash | P/D + DP zero-token ranks compared None with
MC2 capacity and raised TypeError. |
| 14 | vllm-project#13183 | P1 unavailable | ec_both was treated as producer-only
and skipped KV specification/data needed by its consumer role. |
| 15 | vllm-project#13123 | P1 OOB/device error | MRV2 dummy-token remainder was
concentrated on one request and could exceed max_model_len. |
| 16 | vllm-project#13159 | P1 crash | MRV2 num_nans used the wrong Triton libdevice
and the penalty kernel could exceed the CANN grid limit. |
| 17 | vllm-project#12940 | P1 crash | DFlash profiling used total query count
instead of actual input tokens for RoPE/graph capture. |
| 18 | vllm-project#13394 via vllm-project#13405 | P0 correctness | RL sampling tensor lifetime
errors could produce Inf/OOV tokens and contaminate later output. |
| 19 | vllm-project#14142 via vllm-project#14619 | P1 long-run/state correctness | P/D rejection
left stale KV/accounting and unsafe retry/replay behavior could leak,
duplicate, or return wrong responses. |

#### Backport policy

- Selected the audited v0.26 release-adapted commits where available;
the closed rollup vllm-project#14337 was not revived wholesale.
- Kept only one canonical copy of fixes duplicated across 0.23, 0.25,
and main.
- Manually adapted vllm-project#13136, the core vllm-project#13123 input-batch hunk, and vllm-project#13159
to preserve current v0.26/MegaMoe/model-runner behavior.
- Used the current v0.26 target change from vllm-project#13405 and the complete
five-commit vllm-project#14619 series.
- Intentionally excluded performance-only, UX-only, conditional-support,
low-confidence, and owner-unsettled fixes from this release window.

### Does this PR introduce _any_ user-facing change?

Yes, behavior is corrected for the affected configurations: crashes,
deadlocks, hangs, incorrect output, stale KV state, and data corruption
are prevented. There is no new public API, CLI option, or configuration
requirement.

### How was this patch tested?

Local validation completed:

- Audited manifest: 19/19 logical fixes, 23/23 expected source commits;
missing 0, duplicate 0, unexpected 0.
- All 23 commits retain source provenance and Signed-off-by trailers.
- Ruff lint and format checks passed for all 38 changed Python files.
- Python syntax compilation passed for all 38 changed Python files.
- git diff --check, codespell, forbidden-import, package-init,
context-manager, and filename checks passed.
- The final worktree is clean at
86b2ca8.

The backports retain or add focused tests for AscendStore, Mooncake
rejection cleanup, fused MoE/EPLB, W8A8-MXFP8 reload, worker sleep/wake,
MRV2 graph padding, penalty-grid limits, hidden-state extraction, and
two-card speculative DP.

NPU UT/E2E was not run locally because the available Windows environment
has no vLLM, PyTorch/torch_npu, pytest, or Ascend device. CI and
targeted NPU regression are therefore required before merge, especially:

- MRV1/MRV2 full-graph repeated-iteration deadlock/teardown.
- Mooncake TP2/TP4 out-of-order pull KV equality and P/D rejection
cleanup.
- Qwen3-VL FlashComm1 + deepstack fixed-seed correctness.
- Level-2 sleep/wake, dynamic EPLB, and multi-round RL weight reload.
- P/D + DP zero-token ranks, ec_both, DFlash profile/graph, and RL
Inf/OOV sampling.
- Proxy retry and streaming replay behavior from vllm-project#14619.

#### Review checklist

- [x] Only the 19 approved release-critical logical fixes are included.
- [x] One logical fix per commit; vllm-project#14619 remains an atomic five-commit
series.
- [x] No performance-only backports are included.
- [x] Source provenance and sign-offs are retained.
- [ ] Repository CI passes.
- [ ] Targeted Ascend NPU correctness and long-run tests pass.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: kyle-zhangchi <chiiiiiizhang@gmail.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
Signed-off-by: jiaqi-lee <15316070896@163.com>
Signed-off-by: tyy0829 <1455207791@qq.com>
Signed-off-by: yejj710 <abyss1999@163.com>
Signed-off-by: jiajinzhu2 <jiajinzhu@huawei.com>
Signed-off-by: chenyue1122 <oyoy7102@163.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: muziyuhui666 <lijianfu9@huawei.com>
Signed-off-by: Pz1116 <zpbzpb123123@gmail.com>
Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
Signed-off-by: likailong <likailong5@huawei.com>
Signed-off-by: hanxi-java <634498162@qq.com>
Signed-off-by: Liam <ml646@duke.edu>
Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
Signed-off-by: HF-001 <1670186653@qq.com>
Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com>
Signed-off-by: Hcm03 <chengminhua1@huawei.com>
Signed-off-by: zhuyixiang <zhuyixiang2014@163.com>
Signed-off-by: moonseeker <2290166829@qq.com>
Co-authored-by: kyle-zhangchi <chiiiiiizhang@gmail.com>
Co-authored-by: tyy0829 <87685049+tyy0829@users.noreply.github.com>
Co-authored-by: yejj <abyss1999@163.com>
Co-authored-by: jiajinzhu2 <jiajinzhu@huawei.com>
Co-authored-by: CHENYUE <56943221+PHOEBEMOON0802@users.noreply.github.com>
Co-authored-by: Xu Rongsheng <73730571+MmMmaru@users.noreply.github.com>
Co-authored-by: yjyang62 <yangjinyang5@huawei.com>
Co-authored-by: muziyuhui666 <lijianfu9@huawei.com>
Co-authored-by: CXY-Katrina <katrina.cxy@gmail.com>
Co-authored-by: cywang250805 <wangchaoyu7@huawei.com>
Co-authored-by: Bill845514379 <huangjianbao2@huawei.com>
Co-authored-by: yejj710 <yejj710@gmail.com>
Co-authored-by: AuroraEmiya <Sakura.iostream@gmail.com>
Co-authored-by: HaoxinZong <116423146+HaoxinZong@users.noreply.github.com>
Co-authored-by: pz1116 <47019764+Pz1116@users.noreply.github.com>
Co-authored-by: zouyida2052 <zouyida2002@gmail.com>
Co-authored-by: iKeybot <92210799+iKeybot-code@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: 韩熙 <63780107+hanxi-java@users.noreply.github.com>
Co-authored-by: zouzy <38661932+zouzy5137@users.noreply.github.com>
Co-authored-by: AuroraEmiya <92282919+AuroraEmiya@users.noreply.github.com>
Co-authored-by: Liam <ml646@duke.edu>
Co-authored-by: kx <1670186653@qq.com>
Co-authored-by: wangxiaoteng888 <56506195+wangxiaoteng888@users.noreply.github.com>
Co-authored-by: Hcm03 <chengminhua1@huawei.com>
Co-authored-by: zhuyixiang <zhuyixiang2014@163.com>
Co-authored-by: moonseeker <2290166829@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants