Skip to content

[BugFix][0.18.0][310p] fix post-sampling not working in graph mode on 310p - #8077

Merged
wangxiyuan merged 1 commit into
vllm-project:releases/v0.18.0from
Tflowers-0129:v0.18.0
Apr 9, 2026
Merged

wangxiyuan merged 1 commit into
vllm-project:releases/v0.18.0from
Tflowers-0129:v0.18.0

Conversation

@Tflowers-0129

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the service to stall and eventually hang. We first traced the issue to a timeout where the temperature-related div operator was waiting for results from a sub-stream. After investigating the preceding operators, we finally identified the root cause as the q.exponential_() operator, which is not well supported on 310P and triggers an internal issue in the add kernel.

Does this PR introduce any user-facing change?

NA

How was this patch tested?

This patch was thoroughly tested locally(accuracy-dataset test and stress test). It is not easy to design a proper unit test for this case, and I appreciate your understanding.

… 310p

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request resolves a significant stability issue on 310P devices where post-sampling with temperature enabled would lead to system hangs. The solution involves implementing a dedicated sampling pipeline for 310P that intelligently handles device-specific limitations by moving certain operations to the CPU, thereby preventing the unsupported kernel calls and ensuring robust model execution.

Highlights

  • Bug Fix for 310P Post-Sampling: Addressed a critical bug where enabling temperature in post-processing on 310P devices caused service stalls and hangs due to an unsupported q.exponential_() operator, leading to an internal issue in the add kernel.
  • Dedicated 310P Sampler Implementation: Introduced AscendSampler310 and AscendTopKTopPSampler310 to provide a specialized sampling mechanism for 310P devices. This includes a _random_sample_310p function that offloads the problematic q.exponential_() operation to the CPU to ensure compatibility and stability.
  • Model Runner Integration: Updated the NPUModelRunner310 to utilize the newly implemented AscendSampler310 and integrate it with the RejectionSampler for 310P-specific sampling workflows.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a specialized sampler for 310P devices to resolve a hang issue associated with NPU-based exponential noise generation by moving the operation to the CPU. The changes include the implementation of _random_sample_310p and the AscendSampler310 class. Feedback suggests adhering to the repository's PR title and summary formatting standards, optimizing the CPU-to-NPU data transfer using pinned memory and non-blocking operations, and overriding the do_async_exponential method in AscendSampler310 to prevent the base class from triggering the buggy NPU operator.

MambaSpec,
UniformTypeKVCacheSpecs,
)
from vllm.v1.sample.rejection_sampler import RejectionSampler

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request title and summary should follow the repository's style guide format.

Suggested PR Title:

[0.18.0][310p][BugFix] fix post-sampling not working in graph mode on 310P

Suggested PR Summary:

### What this PR does / why we need it?

This PR fixes a hang issue on 310P devices when temperature is enabled in post-processing. The root cause was identified as the `q.exponential_()` operator on NPU, which is buggy on 310P. The fix involves moving the exponential noise generation to the CPU.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

Tested locally with accuracy-dataset and stress tests.
References
  1. Pull Request titles must follow the [Branch][Module][Action] format, and the summary must follow the specified markdown structure. (link)

Comment on lines +34 to +42
with npu_stream_switch(global_stream()):
q = torch.empty_like(probs)
q = q.cpu()
if len(generators) != q.shape[0]:
q.exponential_()
if generators:
for i, generator in generators.items():
q[i].exponential_(generator=generator)
q = q.npu()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The current implementation is inefficient because it allocates q on the NPU and then copies it to the CPU while it's still uninitialized. It's better to allocate q directly on the CPU. Additionally, using pinned memory (pin_memory=True) and an asynchronous transfer (non_blocking=True) will improve performance when moving the tensor back to the NPU.

Suggested change
with npu_stream_switch(global_stream()):
q = torch.empty_like(probs)
q = q.cpu()
if len(generators) != q.shape[0]:
q.exponential_()
if generators:
for i, generator in generators.items():
q[i].exponential_(generator=generator)
q = q.npu()
q = torch.empty(probs.shape,
dtype=probs.dtype,
device='cpu',
pin_memory=True)
if len(generators) != q.shape[0]:
q.exponential_()
if generators:
for i, generator in generators.items():
q[i].exponential_(generator=generator)
with npu_stream_switch(global_stream()):
q = q.to(device=probs.device, non_blocking=True)

Comment on lines +64 to +66
def __init__(self, logprobs_mode=DEFAULT_LOGPROBS_MODE):
super().__init__(logprobs_mode=logprobs_mode)
self.topk_topp_sampler = AscendTopKTopPSampler310(logprobs_mode=logprobs_mode)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The base class AscendSampler implements do_async_exponential, which performs q.exponential_() on the NPU. If enable_async_exponential is enabled in the configuration, this method will still be called during model execution, which could trigger the 310P hang issue even if the result is not used by the sampler. This method should be overridden to avoid executing the buggy NPU operator.

Suggested change
def __init__(self, logprobs_mode=DEFAULT_LOGPROBS_MODE):
super().__init__(logprobs_mode=logprobs_mode)
self.topk_topp_sampler = AscendTopKTopPSampler310(logprobs_mode=logprobs_mode)
def __init__(self, logprobs_mode=DEFAULT_LOGPROBS_MODE):
super().__init__(logprobs_mode=logprobs_mode)
self.topk_topp_sampler = AscendTopKTopPSampler310(logprobs_mode=logprobs_mode)
def do_async_exponential(self, b_s, head_dim, generators):
# Disable async NPU exponential on 310P to avoid hangs.
pass

@wangxiyuan wangxiyuan changed the title [BugFix][0.18.0][310p] fix post-sampling not working in graph mode on… [BugFix][0.18.0][310p] fix post-sampling not working in graph mode on 310p Apr 9, 2026
@wangxiyuan
wangxiyuan merged commit 82e17f6 into vllm-project:releases/v0.18.0 Apr 9, 2026
12 of 13 checks passed
keyi-zz pushed a commit to keyi-zz/vllm-ascend that referenced this pull request Apr 20, 2026
… 310p (vllm-project#8077)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
immengzi pushed a commit to immengzi/vllm-ascend that referenced this pull request May 21, 2026
… 310p (vllm-project#8077)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Tflowers-0129 pushed a commit that referenced this pull request Sep 15, 2026
…V2 on the 310P (#16503)

### What this PR does / why we need it?

Enable **temperature / top-k / top-p** sampling on Ascend **310P Model
Runner V2**.

- First-version `Ascend310PSampler` only supported greedy (`argmax`) and
rejected non-zero temperature.
- Mainline MRV2 uses Triton Gumbel sampling; 310P has no Triton, and NPU
`exponential_` / large RNG can hang under ACLGraph (same issue fixed in
MRV1).
- This PR reuses the **MRV1 inverse-CDF path** (`_random_sample_310p`:
CPU uniform per request → NPU `softmax` + `cumsum` + `searchsorted`),
while keeping the MRV2 `sampling_states` surface required by MTP draft
`propose()`.
- Still rejects min_p / penalties / logprobs / logit processors
(follow-up).

Refs: MRV1 post-sampling fixes (#8017 / #8077), CDF optimization
(#12966).

RFC: #15577 

### Does this PR introduce _any_ user-facing change?

No. On 310P with `VLLM_USE_V2_MODEL_RUNNER=1`, requests may now use
`temperature`, `top_k`, and `top_p` (previously raised
`NotImplementedError`). Default greedy behavior is unchanged.

### How was this patch tested?

- UT: `tests/ut/_310p/test_model_runner_v2_310p.py` (temperature accept
/ scale / greedy; still rejects penalties)

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: Thiagor2002 <13476117628@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants