Skip to content

[Cherry-pick][releases/v0.23.0][Doc][Misc] Add limitation for reduce sample (from #13468) - #13469

Merged
zzzzwwjj merged 1 commit into
vllm-project:releases/v0.23.0from
vllm-ascend-ci:cherry-pick/pr-13468-to-releases-v0.23.0
Aug 4, 2026
Merged

zzzzwwjj merged 1 commit into
vllm-project:releases/v0.23.0from
vllm-ascend-ci:cherry-pick/pr-13468-to-releases-v0.23.0

Conversation

@vllm-ascend-ci

@vllm-ascend-ci vllm-ascend-ci commented Aug 4, 2026 •

Copy link
Copy Markdown
Collaborator

Cherry-pick of PR #13468 onto releases/v0.23.0.

Original PR: #13468
Original author: @zouzy5137


What this PR does / why we need it?

This PR updates the documentation for the enable_reduce_sample configuration option in vLLM Ascend. It adds explicit limitations for this optimization, specifically noting that:

  1. PD (Prefill-Decode) disaggregation scenarios are not supported.
  2. It must be disabled when sampling logprobs are requested, as logprobs would otherwise be silently computed over partitioned logits instead of the full vocabulary, leading to incorrect logprob values and top-k rankings.

Does this PR introduce any user-facing change?

No, this is a documentation-only update.

How was this patch tested?

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request cherry-picks documentation improvements from the main branch to the v0.23.0 release. It clarifies the constraints of the enable_reduce_sample optimization in vLLM Ascend to ensure users are aware of scenarios where the feature may cause incorrect results or is unsupported.

Highlights

  • Documentation Update: Updated the documentation for the enable_reduce_sample configuration option to include critical usage limitations.
  • Added Limitations: Explicitly noted that the feature is incompatible with PD disaggregation scenarios and must be disabled when requesting sampling logprobs to prevent incorrect output.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the documentation for the enable_reduce_sample configuration option in additional_config.md to outline its limitations regarding PD disaggregation and logprob sampling. The reviewer noted that the PR title and summary do not comply with the repository's style guide and provided a formatted suggestion to correct this.

| `rejection_sampler_config` | dict | `{}` | Configuration options for rejection sampler (block verify and entropy verify). |
| `multistream_dsv4_dsa_overlap` | bool | `True` | Whether to enable dsa multi-stream overlap for DeepSeek V4. |
| `enable_reduce_sample` | bool | `False` | Whether to enable reduce sample optimization to reduce communication and computation overheads in the tensor parallelism scenario. When enabled, logits are kept partitioned across TP ranks and only the small set of top-k candidate values/indices is communicated, instead of performing a full-vocabulary all-to-all/all-gather. |
| `enable_reduce_sample` | bool | `False` | Whether to enable reduce sample optimization to reduce communication and computation overheads in the tensor parallelism scenario. When enabled, logits are kept partitioned across TP ranks and only the small set of top-k candidate values/indices is communicated, instead of performing a full-vocabulary all-to-all/all-gather. **Limitations**: (1) PD disaggregation scenarios are not supported. (2) Must be disabled when sampling logprobs are requested. When reduce sample is enabled, logprobs are silently computed over partitioned logits instead of the full vocabulary, producing incorrect logprob values and top-k rankings.|

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The current Pull Request Title does not fully adhere to the Repository Style Guide. Specifically, it includes an extra [Cherry-pick] prefix and does not strictly follow the [Branch][Module][Action] Pull Request Title format.

According to the Repository Style Guide, please update the PR Title and Summary as suggested below:

Suggested PR Title:

[releases/v0.23.0][Doc][Misc] Add limitation for reduce sample (from #13468)

Suggested PR Summary:

### What this PR does / why we need it?

This PR updates the documentation for the `enable_reduce_sample` configuration option in vLLM Ascend. It adds explicit limitations for this optimization, specifically noting that:
1. PD (Prefill-Decode) disaggregation scenarios are not supported.
2. It must be disabled when sampling logprobs are requested, as logprobs would otherwise be silently computed over partitioned logits instead of the full vocabulary, leading to incorrect logprob values and top-k rankings.

Fixes #13468

### Does this PR introduce _any_ user-facing change?

No, this is a documentation-only update.

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main: https://github.com/vllm-project/vllm/commit/0351e9aa1fdf1a51329d1906881528dfe61fc88e
References
  1. The PR Title must follow the format [Branch][Module][Action] Pull Request Title as specified in the Repository Style Guide. (link)

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 4, 2026
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@zzzzwwjj
zzzzwwjj merged commit a14bc1f into vllm-project:releases/v0.23.0 Aug 4, 2026
6 checks passed
Wyz-134 pushed a commit to Wyz-134/vllm-ascend that referenced this pull request Aug 5, 2026
* releases/v0.23.0: (104 commits)
  [Doc][BugFix] Update proxy script name in DeepSeek-V3.2 tutorial (vllm-project#13537)
  [Doc] Fix link errors and update documentation structure (vllm-project#13483)
  [BugFix][releases/v0.23.0] fix fiaV2 contiguous err in GQA (vllm-project#13458)
  [v0.23.0][BugFix] Isolate layerwise GVA keys by parallel rank (vllm-project#13513)
  [Doc][Feature] Add model support of Ascend 950 (vllm-project#13525)
  [Doc] fix DeepSeek V4 Flash&Pro model tutorial docs link error (vllm-project#13497)
  [Cherry-pick][releases/v0.23.0][Doc][Misc] Add limitation for reduce sample (from vllm-project#13468) (vllm-project#13469)
  [BugFix][v0.23.0][KV Pool] Include MTP KV in layerwise AscendStore transfer (vllm-project#13454)
  [Doc][Misc] Standardize TorchNPU capitalization and update Ascend 950 product terminology (vllm-project#13089)
  [v0.23.0][Doc] Translated Doc files 2026-08-04 (vllm-project#13437)
  [Misc][v0.23.0] Fix translation extraction for tables nested in tabs (vllm-project#13413)
  [Doc] Fix translation and formatting in documentation (vllm-project#13390)
  [releases/v0.23.0][Doc][Misc] Backport Kimi-K2-Thinking tuning docs to v0.23.0 (vllm-project#13361)
  [Doc] Deployment key parameter supplement- vllm-project#13297 (vllm-project#13299)
  [v0.23.0][Doc] Translated Doc files 2026-07-31 (vllm-project#13283)
  [Doc][Misc] Update max-num-seqs configurations in GLM5 tutorial (vllm-project#13203)
  [Cherry-pick][releases/v0.23.0][Doc][Misc] Add deployment reference notice for GLM-5 (from vllm-project#12958) (vllm-project#12960)
  [BugFix][v0.23.0][KV Pool] Guard batch_get_key_info before memcache backend init (vllm-project#13307)
  [DOC]Modify the scope of scenarios supported by CP (vllm-project#13303)
  Revert "[cherry-pick][v0.23.0][Performance] remove D2H sync in QLIMetadata builder for DSA_CP" (vllm-project#13289)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants