Skip to content

[DeepSeek V4.1] Add optional DeepSelect SM90 candidate block Top-K - #39421

Closed
yuyu5333 wants to merge 2 commits into
sgl-project:dsv4.1from
yuyu5333:feature/dsv41-deepselect-sm90
Closed

yuyu5333 wants to merge 2 commits into
sgl-project:dsv4.1from
yuyu5333:feature/dsv41-deepselect-sm90

Conversation

@yuyu5333

@yuyu5333 yuyu5333 commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

DeepSeek V4.1 candidate indexing performs an FP32 block-level Top-K selection before publishing the candidate mask. On NVIDIA Hopper SM90, this selection can use the optimized DeepSelect kernel to reduce the latency of candidate publication for long-context decode workloads.

This integration depends on the SM90 implementation from deepseek-ai/DeepSelect#14.

Modifications

  • Add the SGLANG_OPT_DSV41_DEEPSELECT_CANDIDATE_TOPK environment switch. It is disabled by default, preserving the existing torch.topk path.
  • Use deep_select.topk for candidate-block FP32 Top-K when the switch is enabled on SM90.
  • Pad candidate score rows to the 1024-byte alignment required by DeepSelect.
  • Pass the actual DeepSelect value and index output strides to the Triton candidate-mask publication kernel.
  • Load DeepSelect lazily and report a clear error when the feature is enabled without the package.
  • Add an SM90 regression test for variable sequence lengths, non-aligned score widths, and padded Top-K output strides. The test is skipped when the optional DeepSelect package is unavailable.

Enable the path with:

export SGLANG_OPT_DSV41_DEEPSELECT_CANDIDATE_TOPK=1

Accuracy Tests

The candidate logits and published masks were compared against the existing torch.topk implementation on an NVIDIA H20.

  • All 5 tests in test_dsv4_indexer_postprocess.py passed.
  • All 24 additional SM90 edge cases passed.
  • The edge cases cover zero-length and short rows, variable sequence lengths, non-aligned widths, topk_blocks values of 1, 513, and 2048, padded output strides, and finite tied scores.
  • Random inputs without ties produced exactly matching filtered logits and candidate masks.
  • For finite tied scores, the selected set satisfies the same Top-K threshold and candidate-mask semantics.

Speed Tests and Profiling

Benchmarks were run on an NVIDIA H20 (SM90) with FP32 candidate scores, topk_blocks=2048, cold-cache flushing, 20 warmup iterations, and the median of 100 measured iterations. The measured region includes candidate block score reduction, candidate-mask publication, and the downstream fused paged Top-K.

Final Top-K 512

Rows Sequence length torch.topk DeepSelect Speedup
1 128K 105.184 us 62.784 us 1.68x
1 256K 126.976 us 63.184 us 2.01x
1 512K 127.488 us 80.096 us 1.59x
1 1M 133.936 us 98.480 us 1.36x
6 128K 98.944 us 61.424 us 1.61x
6 256K 124.800 us 63.152 us 1.98x
6 512K 126.864 us 80.032 us 1.59x
6 1M 141.552 us 111.872 us 1.27x

Final Top-K 2048

Rows Sequence length torch.topk DeepSelect Speedup
1 128K 106.000 us 62.928 us 1.68x
1 256K 123.424 us 64.016 us 1.93x
1 512K 123.200 us 81.328 us 1.51x
1 1M 129.840 us 100.256 us 1.30x
6 128K 99.680 us 62.720 us 1.59x
6 256K 121.568 us 63.760 us 1.91x
6 512K 124.896 us 80.912 us 1.54x
6 1M 141.392 us 112.928 us 1.25x

Across the tested shapes, the complete candidate publication and final Top-K path is 1.25x to 2.01x faster. The candidate block Top-K and mask-publication region alone is 1.34x to 2.19x faster.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #35080145279
Latest PR Test (Extra): ❌ Run #35080144980
Latest PR Test (AMD ROCm 10): ❌ Run #35080145194

@hnyls2002
hnyls2002 force-pushed the dsv4.1 branch 21 times, most recently from c81d5c2 to 660722f Compare September 16, 2026 05:47
@yuyu5333
yuyu5333 force-pushed the feature/dsv41-deepselect-sm90 branch from 3fc1ce3 to a0d85cb Compare September 16, 2026 09:33
@hnyls2002
hnyls2002 deleted the branch sgl-project:dsv4.1 September 18, 2026 09:55
@hnyls2002 hnyls2002 closed this Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants