Skip to content

[BugFix][310p] Fix post-sampling not working in graph mode on 310P - #8017

Merged
wangxiyuan merged 3 commits into
vllm-project:mainfrom
Tflowers-0129:fix/310p-postprocess-on-error
Apr 8, 2026
Merged

wangxiyuan merged 3 commits into
vllm-project:mainfrom
Tflowers-0129:fix/310p-postprocess-on-error

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Apr 7, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the service to stall and eventually hang. We first traced the issue to a timeout where the temperature-related div operator was waiting for results from a sub-stream. After investigating the preceding operators, we finally identified the root cause as the q.exponential_() operator, which is not well supported on 310P and triggers an internal issue in the add kernel.

Does this PR introduce any user-facing change?

NA

How was this patch tested?

This patch was thoroughly tested locally(accuracy-dataset test and stress test). It is not easy to design a proper unit test for this case, and I appreciate your understanding.

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a critical stability issue where 310P devices would hang during post-sampling due to an incompatible exponential operator kernel. By offloading the exponential distribution generation to the CPU and synchronizing streams appropriately, the service can now perform temperature-based sampling reliably without triggering internal NPU kernel errors.

Highlights

  • Bug Fix: Resolved a service hang issue on 310P devices caused by the unsupported q.exponential_() operator in the temperature-based post-sampling process.
  • Implementation: Introduced AscendSampler310 to handle exponential generation on the CPU, bypassing the problematic NPU kernel.
  • Integration: Updated NPUModelRunner310 to utilize the new sampler and ensure compatibility with existing rejection sampling logic.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

github-actions Bot commented Apr 7, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

Signed-off-by: Tflowers-0129 <2906339855@qq.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[310P][Sampler][Feature] Implement AscendSampler310 and RejectionSampler support

Suggested PR Summary:

### What this PR does / why we need it?
This pull request introduces the `AscendSampler310` and `AscendTopKTopPSampler310` classes to support sampling operations specifically for the Ascend 310P hardware. It also integrates `RejectionSampler` into the `NPUModelRunner310`. 

Feedback was provided regarding inefficient memory management and potential device mismatches in `sampler.py`. Specifically, the `_generate_exponential_q` function should be updated to handle device placement more robustly, and tensors should be initialized directly on the CPU to avoid unnecessary NPU allocations and blocking transfers when performing CPU-based exponential sampling.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
CI passed with existing tests.

Comment thread vllm_ascend/_310p/sample/sampler.py Outdated
Comment thread vllm_ascend/_310p/sample/sampler.py Outdated
Comment thread vllm_ascend/_310p/sample/sampler.py Outdated
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Comment thread vllm_ascend/_310p/model_runner_310p.py
@wangxiyuan
wangxiyuan merged commit b7987f9 into vllm-project:main Apr 8, 2026
39 checks passed
Tflowers-0129 added a commit to pu-zhe/vllm-ascend that referenced this pull request Apr 9, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
(cherry picked from commit b7987f9)
guxin108 pushed a commit to guxin108/vllm-ascend that referenced this pull request Apr 24, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: guxin108 <1252896542@qq.com>
chenchuw886 pushed a commit to chenchuw886/vllm-ascend that referenced this pull request Apr 27, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
zouyida2052 pushed a commit to zouyida2052/vllm-ascend that referenced this pull request Apr 28, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
yangzhe-2026 pushed a commit to yangzhe-2026/vllm-ascend that referenced this pull request May 6, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
nanxingMy pushed a commit to nanxingMy/vllm-ascend that referenced this pull request May 15, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: nanxing <1014662416@qq.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
…llm-project#8017)

### What this PR does / why we need it?

Enabling temperature in post-processing on 310P devices can cause the
service to stall and eventually hang. We first traced the issue to a
timeout where the temperature-related `div` operator was waiting for
results from a sub-stream. After investigating the preceding operators,
we finally identified the root cause as the `q.exponential_()` operator,
which is not well supported on 310P and triggers an internal issue in
the `add` kernel.

### Does this PR introduce _any_ user-facing change?
NA

### How was this patch tested?
This patch was thoroughly tested locally(accuracy-dataset test and
stress test). It is not easy to design a proper unit test for this case,
and I appreciate your understanding.
- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@14acf42

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Tflowers-0129 pushed a commit that referenced this pull request Sep 15, 2026
…V2 on the 310P (#16503)

### What this PR does / why we need it?

Enable **temperature / top-k / top-p** sampling on Ascend **310P Model
Runner V2**.

- First-version `Ascend310PSampler` only supported greedy (`argmax`) and
rejected non-zero temperature.
- Mainline MRV2 uses Triton Gumbel sampling; 310P has no Triton, and NPU
`exponential_` / large RNG can hang under ACLGraph (same issue fixed in
MRV1).
- This PR reuses the **MRV1 inverse-CDF path** (`_random_sample_310p`:
CPU uniform per request → NPU `softmax` + `cumsum` + `searchsorted`),
while keeping the MRV2 `sampling_states` surface required by MTP draft
`propose()`.
- Still rejects min_p / penalties / logprobs / logit processors
(follow-up).

Refs: MRV1 post-sampling fixes (#8017 / #8077), CDF optimization
(#12966).

RFC: #15577 

### Does this PR introduce _any_ user-facing change?

No. On 310P with `VLLM_USE_V2_MODEL_RUNNER=1`, requests may now use
`temperature`, `top_k`, and `top_p` (previously raised
`NotImplementedError`). Default greedy behavior is unchanged.

### How was this patch tested?

- UT: `tests/ut/_310p/test_model_runner_v2_310p.py` (temperature accept
/ scale / greedy; still rejects penalties)

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: Thiagor2002 <13476117628@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants