Skip to content

[BugFix][Spec Decode] Preserve RNG state for zero-draft requests - #13755

Open
yaleyoou wants to merge 1 commit into
vllm-project:mainfrom
yaleyoou:fix/zero-draft-rng
Open

yaleyoou wants to merge 1 commit into
vllm-project:mainfrom
yaleyoou:fix/zero-draft-rng

Conversation

@yaleyoou

@yaleyoou yaleyoou commented Aug 7, 2026 •

Copy link
Copy Markdown

What this PR does / why we need it?

sample_recovered_tokens() consumed per-request RNG for zero-draft
requests and then discarded the generated exponential samples with
torch.where. This made later rejection sampling depend on scheduling
and batch composition.

Use the existing Python num_draft_tokens data to skip generator-aware
draws for zero-draft rows. Active-draft and unseeded sampling behavior
remain unchanged. Add NPU regression coverage for zero, active, mixed,
independent, partial-generator, and skipped-round cases.

Fixes #13608

Does this PR introduce any user-facing change?

No API change. Fixed-seed speculative decoding now preserves request RNG
state during zero-draft rounds.

How was this patch tested?

  • 7 new focused NPU regression tests passed

  • 31 existing NPU rejection-kernel tests passed

  • 19 rejection-sampler unit tests passed

  • 85 spec-decode unit tests passed

  • No scalar synchronization, new H2D copy, or persistent HBM growth was observed

  • git diff --check and relevant pinned formatting/lint hooks passed

  • Full format.sh ci did not complete because actionlint's Go dependency
    download timed out at the proxy

  • Ascend 910B3, CANN 9.0.0, torch 2.10.0,
    torch-npu 2.10.0.post2, triton-ascend 3.2.1
    vllm-project/vllm@0351e9a

  • vLLM version: v0.27.1

  • vLLM main: vllm-project/vllm@cdc4824

@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a bug in the speculative decoding process where zero-draft requests were unnecessarily consuming RNG state. By conditionally skipping exponential sampling for requests without draft tokens, the implementation ensures consistent RNG behavior and improves reproducibility for fixed-seed speculative decoding.

Highlights

  • RNG State Preservation: Modified the rejection sampler to skip RNG draws for zero-draft requests, ensuring the request-specific RNG state remains preserved.
  • Regression Testing: Added a new comprehensive test suite in tests/ut/sample/a2/test_rejection_sampler_rng.py to verify RNG behavior across various draft scenarios.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][BugFix] Avoid advancing RNG state for requests with zero draft tokens

Suggested PR Summary:

### What this PR does / why we need it?
This PR optimizes the rejection sampler by avoiding random number generation for requests that have zero draft tokens. Previously, random numbers were generated for all generators regardless of whether they had draft tokens, which unnecessarily advanced the RNG state and impacted reproducibility. Now, the RNG state is only advanced for active draft requests.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Added unit tests in `tests/ut/sample/a2/test_rejection_sampler_rng.py` to verify generator state preservation and independence.

I have no further feedback as the changes are well-tested and correctly address the reproducibility issue.

@yaleyoou yaleyoou changed the title [BugFix][Spec Decode] Preserve RNG state for zero-draft requests [BugFix][Spec Decode] Preserve RNG state for zero-draft requests Aug 7, 2026
@yaleyoou
yaleyoou force-pushed the fix/zero-draft-rng branch from e011d4f to 05f0855 Compare August 7, 2026 04:50
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: yaleyoou <yaleyoou@gmail.com>
@yaleyoou

Copy link
Copy Markdown
Author

@realliujiaxu @wangxiyuan @Yikun Could one of you please add the ready-precise label? Test selection has completed and cigate is waiting for it.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Zero-draft requests incorrectly advance per-request RNG in rejection sampling

1 participant