Skip to content

[BigFIx][CI] Fix a resample bug and enable mrv2 in Mimimax M2.5 - #16920

Open
AuroraEmiya wants to merge 1 commit into
vllm-project:mainfrom
AuroraEmiya:MRV2/fix_resample
Open

AuroraEmiya wants to merge 1 commit into
vllm-project:mainfrom
AuroraEmiya:MRV2/fix_resample

Conversation

@AuroraEmiya

@AuroraEmiya AuroraEmiya commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

This PR contains two related fixes for the MiniMax-M2.5 + Eagle3 MRV2 nightly path:

  1. Fix residual resampling RNG coupling in MRV2 speculative decoding.

    In the stochastic rejection path, the rejection decision and the residual categorical resample could use the same RNG domain:

    rejection:
      RNG(seed, position)
    
    residual resample:
      RNG(seed, position)
    

    For the first rejected speculative token, both paths can therefore consume the same uniform random value. However, once rejection has happened, that random value is already conditioned by the rejection event and is no longer an independent Uniform(0, 1) draw for residual sampling.

    This PR keeps the rejection decision unchanged and moves only random non-bonus residual resampling into a separate RNG domain:

    residual_position = position + (1 << 29)
    

    The salt is applied only when is_random_residual is true. Greedy paths, bonus-token sampling, residual probability mass computation, ordinary categorical sampling, and draft sampling are unchanged.

  2. Explicitly enable MRV2 for the MiniMax-M2.5 A2 nightly case.

    PR Revert "[Feature][MRV2] expand default MRv2 architecture whitelist and add dspark" #16832 reverted the default MRV2 architecture whitelist introduced by [Feature][MRV2] expand default MRv2 architecture whitelist and add dspark #16626. After that revert, Ascend enables Model Runner V2 only when VLLM_USE_V2_MODEL_RUNNER is explicitly set.

    The MiniMax-M2.5 A2 nightly configuration did not set this environment variable, so the test would otherwise fall back to MRV1 and would no longer exercise the MRV2 speculative-decoding path fixed above.

    This PR adds:

    VLLM_USE_V2_MODEL_RUNNER: "1"

    to the MiniMax-M2.5 A2 nightly configuration.

Related:

Does this PR introduce any user-facing change?

Yes, for the affected MRV2 stochastic speculative-decoding path.

Residual resampling after a random rejection now uses an independent RNG domain, so fixed-seed token trajectories may differ from the previous implementation. The sampling API, probability definition, model inputs/outputs, and configuration schema are unchanged.

The additional VLLM_USE_V2_MODEL_RUNNER=1 change only affects the MiniMax-M2.5 A2 nightly test configuration and ensures that CI continues to exercise MRV2 after #16832.

How was this patch tested?

1. Full MiniMax-M2.5 GPQA validation

The R-FIX implementation was validated with three independent full GPQA Diamond runs using the production-like MiniMax-M2.5 MRV2 + Eagle3 configuration.

Common configuration:

Model: Eco-Tech/MiniMax-M2.5-w8a8-QuaRot
TP: 8
Model Runner: V2
Speculative decoding: Eagle3
num_speculative_tokens: 3
server seed: 1024
request seed: not explicitly set
GPQA Diamond: 198 questions
temperature: 1.0
top_p: 0.95
top_k: 40
max_out_len: 131072
fresh cold start for every run

R-FIX results:

Run Accuracy
Run 1 162 / 198 = 81.82%
Run 2 158 / 198 = 79.80%
Run 3 160 / 198 = 80.81%
Mean 80.81%
Population std 0.82 pp
Range 2.02 pp

All three runs completed with 198 valid results and zero request failures.

For reference, the current single-FP32 categorical baseline was observed at:

78.79% / 85.35% / 80.81%
mean: 81.65%
population std: 2.75 pp
range: 6.57 pp

These runs do not show an accuracy-mean improvement from this patch, and this PR does not claim one. They do show that the fixed MRV2 path executes successfully end-to-end and that the three observed R-FIX runs had substantially lower run-to-run spread. More repetitions would be needed to make a statistical stability claim.

2. Source isolation

The validated R-FIX resample.py used for the three GPQA runs had SHA256:

81af9b95c93cac8e8485f204ac69226a2a52d2a8e5cfbc9e3c4bf8c8a8bf4095

The only intended semantic change in that file is the RNG-domain separation for is_random_residual.

3. CI configuration

The MiniMax-M2.5 A2 nightly config explicitly enables:

VLLM_USE_V2_MODEL_RUNNER=1

so that after #16832 the nightly continues to test the MRV2 + Eagle3 path rather than silently falling back to MRV1.

Signed-off-by: AuroraEmiya <Sakura.iostream@gmail.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a sampling logic issue in the vLLM Ascend backend by adding a noise salt to the resampling process. Additionally, it updates the E2E test configuration for the MiniMax-M2.5 model to utilize the V2 model runner, ensuring better compatibility and performance testing.

Highlights

  • Resample Bug Fix: Introduced a noise salt in the categorical resampling kernel to ensure distinct RNG domains for residual sampling, preventing potential collisions.
  • Model Runner Configuration: Enabled the V2 model runner for the MiniMax-M2.5 model in the nightly E2E test configuration.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enables the V2 model runner in the E2E test configuration for MiniMax M2.5 and fixes a bug in the speculative decoding categorical resampling kernel by ensuring residual sampling uses a separate RNG domain via a noise salt. The review feedback suggests updating the PR title and summary to match the repository style guide, and replacing the use of tl.where with a standard Python ternary operator to avoid potential compilation or performance issues on Triton backends.

Comment thread vllm_ascend/ops/triton/v2/spec_decode/resample.py
@czydyy

czydyy commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly MiniMax-M2.5-w8a8-QuaRot-A2
nightly command triggered.

cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 19, 2026
Port vllm-project#16920: salt the RNG position when resampling a random residual
so residual draws do not reuse the rejection-conditioned noise.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 19, 2026
Port vllm-project#16920: salt the RNG position when resampling a random residual
so residual draws do not reuse the rejection-conditioned noise.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 20, 2026
Port vllm-project#16920: salt the RNG position when resampling a random residual
so residual draws do not reuse the rejection-conditioned noise.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 20, 2026
Port vllm-project#16920: salt the RNG position when resampling a random residual
so residual draws do not reuse the rejection-conditioned noise.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 20, 2026
Port vllm-project#16920: salt the RNG position when resampling a random residual
so residual draws do not reuse the rejection-conditioned noise.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 21, 2026
Port vllm-project#16920: salt the RNG position when resampling a random residual
so residual draws do not reuse the rejection-conditioned noise.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants