Skip to content

[Sampling] Add selected/support sampling logprob modes - #40932

Merged
ByronHsu merged 5 commits into
sgl-project:mainfrom
nanjiangwill:feat/sampling-distribution-metadata
Sep 24, 2026
Merged

ByronHsu merged 5 commits into
sgl-project:mainfrom
nanjiangwill:feat/sampling-distribution-metadata

Conversation

@nanjiangwill

@nanjiangwill nanjiangwill commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add optional sampling_logprobs_mode: Literal["selected", "support"].
  • Preserve the existing default: output_token_sampling_logprobs[t] is the selected output token behavior logprob.
  • In "support" mode, return output_token_sampling_logprobs[t] as list[float], aligned elementwise with output_token_sampling_mask[t].
  • Preserve the existing singular response key and length key.
  • Carry the request-selected representation through raw generation, chat completions, sessions, pipeline parallelism, streaming, multi-tokenizer serving, and P/D disaggregation.

API contract

client.generate(
    ...,
    return_sampling_mask=True,
    sampling_logprobs_mode="support",
)

For output position t, let w_t be the exact post-filter weights captured from the sampler, S_t = {a | w_t(a) > 0}, and

q_t(a) = w_t(a) / sum_{v in S_t} w_t(v).

With return_sampling_mask=false, sampling_logprobs_mode must be omitted.

With return_sampling_mask=true:

  • Omitting sampling_logprobs_mode or setting it to "selected" returns log q_t(output_ids[t]) as a float. This keeps the existing default response shape.
  • sampling_logprobs_mode="support" returns one log q_t(a) per token ID in output_token_sampling_mask[t] as a list. The token-ID and logprob rows have identical lengths and elementwise alignment.
  • Both modes describe the same normalized behavior distribution. The mode changes how much of that distribution is returned, not how it is normalized.
  • Greedy sampling returns a singleton support; selected mode returns 0.0, while support mode returns [0.0].
  • Any explicit sampling_logprobs_mode requires return_sampling_mask=true; when masks are enabled, omission resolves to selected mode at the internal request boundary.

The support row order is not an API guarantee; elementwise alignment is. Overflow and invalid captured support continue to fail closed, so no partial distribution is returned.

Implementation

  • Reuse the exact post-filter weights already captured by the sampling-mask path; there is no second top-k/top-p reconstruction.
  • Preserve the existing producer-captured selected-token path. If token synchronization or grammar changes the selected token, keep the existing fallback that recomputes its logprob from the captured weights.
  • Build the full support-logprob tensor only when at least one request selects support mode.
  • In mixed batches, index-select only support-mode rows. Selected-mode neighbors do not allocate or copy full support-logprob rows.
  • Append the mode as a defaulted tail field in the Python scheduler wire, preserving the existing Rust positional prefix and compatibility with shorter senders.
  • Widen the opt-in P/D sampling-logprob metadata row to the configured sampling-mask capacity so support-mode handoff tokens remain aligned.

Performance boundary

When sampling-mask output is disabled, this adds no sampler tensor work or payload. With masks enabled:

  • Selected mode retains the existing compact result shape: approximately 4RC + 12R bytes for R mask rows and packed capacity C.
  • Support mode adds approximately 4R_sC bytes, where R_s is only the number of support-mode rows.
  • P/D sampling-mask metadata remains behind SGLANG_ENABLE_DISAGG_SAMPLING_MASK; support mode necessarily makes the aligned float row capacity-sized, matching the token-ID row.

A packed/base64 ragged response format is intentionally not part of this change. It is an independent transport optimization and can be added without changing these selected/support semantics.

Validation

  • 314 focused CPU tests passed, 2 skipped, with 142 subtests after the review changes.
  • Coverage includes the public option validation matrix, default selected behavior, explicit support behavior, mixed selected/support batches, support-only row packing, pipeline-parallel transport, multi-tokenizer splitting, P/D handoff, overflow, invalid sampled support, streaming, and custom hard exclusions.
  • All pre-commit hooks passed for the changed files.
  • The registered CUDA sampling-mask suite passed on 2× H100 80 GB GPUs (CUDA 13.0, PyTorch 2.13.0): 34 tests and 28 subtests in 437.53 seconds. Modal run

CI States

Latest PR Test (Base): ⏳ Run #35951430674
Latest PR Test (Extra): ❌ Run #35951430010
Latest PR Test (AMD ROCm 10): ❌ Run #35951430457

@nanjiangwill
nanjiangwill marked this pull request as draft September 23, 2026 15:28
@nanjiangwill
nanjiangwill force-pushed the feat/sampling-distribution-metadata branch from e2058ba to 130295d Compare September 23, 2026 17:13
@nanjiangwill nanjiangwill changed the title Expose sampling distribution metadata in generation responses Expose sampling-support behavior logprobs Sep 23, 2026
@nanjiangwill
nanjiangwill force-pushed the feat/sampling-distribution-metadata branch from 130295d to 3dd6ddc Compare September 23, 2026 20:29
@nanjiangwill nanjiangwill changed the title Expose sampling-support behavior logprobs [Sampling] Align sampling logprobs with returned support Sep 23, 2026
@nanjiangwill
nanjiangwill force-pushed the feat/sampling-distribution-metadata branch from 3dd6ddc to d03166b Compare September 23, 2026 20:35
@nanjiangwill nanjiangwill changed the title [Sampling] Align sampling logprobs with returned support [Sampling] Align sampling logprobs with returned masks Sep 23, 2026
@nanjiangwill
nanjiangwill force-pushed the feat/sampling-distribution-metadata branch from d03166b to 960a136 Compare September 23, 2026 21:37
@nanjiangwill nanjiangwill changed the title [Sampling] Align sampling logprobs with returned masks [Sampling] Add selected/support sampling logprob modes Sep 23, 2026
@nanjiangwill
nanjiangwill force-pushed the feat/sampling-distribution-metadata branch 2 times, most recently from 820a0b7 to 74b08b3 Compare September 23, 2026 23:25
Comment thread python/sglang/srt/entrypoints/openai/protocol.py Outdated
Comment thread python/sglang/srt/managers/scheduler_components/batch_result_processor.py Outdated
Comment thread python/sglang/srt/disaggregation/utils.py Outdated
@nanjiangwill
nanjiangwill force-pushed the feat/sampling-distribution-metadata branch from 74b08b3 to 76443ca Compare September 24, 2026 01:04
@nanjiangwill
nanjiangwill marked this pull request as ready for review September 24, 2026 01:29
@ByronHsu

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 24, 2026
@ByronHsu

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/npu/basic_function/dllm/test_npu_llada2_mini.py

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/npu/basic_function/dllm/test_npu_llada2_mini.py:

⛔ test/registered/npu/basic_function/dllm/test_npu_llada2_mini.py: test/registered/npu/basic_function/dllm/test_npu_llada2_mini.py is registered for NPU (suite base-b-test-4-npu-a3), not for CUDA or CPU; rerun-test.yml has no NPU job. Rerun it with /rerun-failed-ci, or dispatch the NPU workflow manually.

@ByronHsu

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@ByronHsu
ByronHsu merged commit c3685df into sgl-project:main Sep 24, 2026
168 of 204 checks passed
guapisolo pushed a commit that referenced this pull request Sep 24, 2026
@nanjiangwill
nanjiangwill deleted the feat/sampling-distribution-metadata branch September 24, 2026 06:12
nanjiangwill added a commit to modal-projects/sglang that referenced this pull request Sep 29, 2026
jvmncs pushed a commit to modal-projects/sglang that referenced this pull request Oct 2, 2026
…rob modes (sgl-project#40932)

(cherry picked from commit c3685df)

Preserve the v0.5.20 wire prefix and inline device-index construction.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants