Skip to content

[DSv4.1] Support SM90 candidate indexers - #40860

Open
YJR722 wants to merge 1 commit into
sgl-project:mainfrom
YJR722:YJR722/dsv41-sm90-triton-candidate-indexer
Open

YJR722 wants to merge 1 commit into
sgl-project:mainfrom
YJR722:YJR722/dsv41-sm90-triton-candidate-indexer

Conversation

@YJR722

@YJR722 YJR722 commented Sep 23, 2026 •

Copy link
Copy Markdown

Motivation

Add candidate-only indexer scoring on SM90 for DeepSeek-V4.1, following #40574. The current v41_indexer implementation already provides unified prefill/decode protocols and compact candidate block IDs, but Hopper consumers still score the full context before selecting among those blocks.

DeepGEMM's paged sparse FP4 MQA scorer is not available on SM90. An earlier approach expanded each query's block IDs and gathered its dequantized K rows into a [queries, candidates, head_dim] tensor before Torch scoring. This PR instead uses Triton to read candidate K positions directly, avoiding that per-query K tensor. A future SM90 DeepGEMM sparse scorer could replace this implementation behind the existing protocol.

Modifications

  • Reuse DenseBlocksBackend, BlockIds, Selection, and the existing publish_* / consume_* interfaces from the upstream refactor, without introducing new public input or candidate metadata classes.
  • Enable direct candidate scoring for SM90 prefill consumers using shared BF16 K. Source layers retain upstream scoring and candidate publication. Per-request K dequantization and a separate Top-K operation are still required.
  • Extend fp4_index_logits_decode with optional candidate block addressing for decode/verify consumers. Preserve the upstream dense invisible-tile shortcut. The existing context-sized slot map is still constructed.
  • Restore compact score columns to request-local positions before the existing selection writes, preserving causal masking, padding, and per-request tail slicing.
  • Enable the existing length-based CUDA graph variants on SM90: four variants for ordinary decode and two for DSPARK target verify. Reuse upstream dispatch and skip handling.
  • Retain an independent Torch reference and the Torch fallback for unsupported prefill shapes. SM100 backend selection is unchanged.

Accuracy Tests

The regression tests added in this PR passed on H20 at 528619e4b2: 37 tests passed. They cover BF16/FP4 candidate scoring against reference implementations, prefill protocol equivalence, causal masking, padding, index mapping, tail slicing, and CUDA graph replay/selection.

GPQA Diamond: 198 questions, one sample per question, evaluated with sgl-eval. Both arms use the same 1P1D H20-3e deployment (TP8/EP8/DP1/CP1), DSPARK block5 and decode full CUDA graphs, with both P and D running the corresponding revision. Parameters: concurrency=16, thinking enabled, reasoning_effort=max, temperature=1.0, top_p=0.95, max_tokens=65536, seed=1.

Metric Main (cc012abd21) This PR (528619e4b2)
Correct / total 179 / 198 178 / 198
Accuracy 90.40% 89.90%
Truncated at max_tokens 3 / 198 (1.52%) 6 / 198 (3.03%)
Request errors 0 0

Speed Tests and Profiling

Baseline A: cc012abd21. This PR B: 528619e4b2.

Both experiments use DeepSeek-V4.1-Flash in a real 1P1D deployment, with two nodes of 8×H20-3e, TP8/EP8/DP1/CP1.

Prefill TTFT

D is fixed; only P switches between A/B. FlashInfer MXFP4 with FP8 compute, chunk=1024, concurrency=1, OSL=1, speculation and CUDA graphs disabled. Each length uses identical input IDs, cache flushes, one warmup, and three measured requests. Values are median streaming TTFT.

ISL Main TTFT (s) This PR TTFT (s) TTFT reduction
16K 3.9603 3.9465 0.35%
32K 7.4733 7.0634 5.48%
64K 15.7998 14.2035 10.10%
128K 38.9008 31.8004 18.25%
256K 108.1726 80.0343 26.01%

16K is effectively unchanged; gains increase with context length.

DSPARK decode/verify throughput and effective TPOT

P is fixed; only D switches between A/B. DSPARK block5, static verify, full CUDA graphs, steady-state BS8, and OSL limit=40,000. Each configuration runs three times with identical inputs; measurements are taken while the active batch size remains stable at 8.

ISL Main throughput (tok/s) This PR throughput (tok/s) Throughput gain Main effective TPOT (ms) This PR effective TPOT (ms) TPOT reduction
16K 1189.47 1288.78 8.35% 6.726 6.209 7.69%
32K 1155.42 1207.89 4.54% 6.924 6.628 4.28%
64K 1085.60 1202.75 10.79% 7.371 6.652 9.75%
128K 508.30 621.08 22.19% 16.417 13.283 19.09%

Results are three-run averages.

Checklist

  • Format the changed code according to the SGLang pre-commit configuration.
  • Add kernel and protocol regression tests.
  • Update user-facing documentation, if required.
  • Provide targeted correctness tests, GPQA Diamond accuracy results, and performance benchmarks.
  • Follow the SGLang code style guidance.

CI States

Latest PR Test (Base): ❌ Run #37717868631
Latest PR Test (Extra): ❌ Run #37717868167
Latest PR Test (AMD ROCm 10): ❌ Run #37717868527

@yuan-luo

Copy link
Copy Markdown
Collaborator

Please fix the lint first.

@YJR722
YJR722 force-pushed the YJR722/dsv41-sm90-triton-candidate-indexer branch from 22056a5 to c2c2e2d Compare September 23, 2026 06:47
@Frank-whw

Copy link
Copy Markdown

Thanks for this work. I prepared two small follow-up commits on top of the current PR head:

  1. 775ab53e1a54d5574b4106476951944f5ead5465 fixes compact top-k metadata so valid candidate positions are ordered before -1 padding.
  2. 170214f81ae37fa95dcfcdadb9f5377680d4b009 adds a regression test for compact int32 prefill candidate-block metadata across score tiles, guarding against a return to full-width mask concatenation.

Branch: https://github.com/Frank-whw/sglang/tree/test-dsv41-h200-prefill-memory

Would you be willing to cherry-pick them in this order?

git fetch https://github.com/Frank-whw/sglang.git test-dsv41-h200-prefill-memory
git cherry-pick 775ab53e1a54d5574b4106476951944f5ead5465 170214f81ae37fa95dcfcdadb9f5377680d4b009

I can provide the detailed validation notes privately if useful.

@YJR722

YJR722 commented Sep 24, 2026

Copy link
Copy Markdown
Author

@Frank-whw Thanks for preparing these follow-up patches!
For 775ab53, the CandidateIndexer prefill interface explicitly permits unordered output, and _low_ratio_index_topk_torch already sorts valid positions ahead of padding before mapping them to KV slots. Adding another sort in _write_topk appears redundant for the current path, so I’d prefer not to cherry-pick this commit as-is. If you have a failing case that exposes a gap in the existing handling, I’d be happy to take a closer look.
For 170214f, the goal of protecting compact int32 candidate metadata across score tiles makes sense. I’d like to review the test coverage a bit further, particularly whether the blanket torch.cat prohibition is the right guard against full-width mask accumulation. I’ll consider this test independently of the sorting change.
Any reproducer or validation notes you can share would be helpful. Thanks again for reviewing the PR!

@Frank-whw

Copy link
Copy Markdown

@YJR722 Thanks for the careful review. I agree: publish_prefill permits unordered positions, and the backend sorts valid positions before KV-slot mapping. The ordering expectation behind 775ab53 was too strong, so please disregard that commit.

I prepared a clean standalone replacement test on top of the current PR head, without 775ab53:

It removes the blanket torch.cat prohibition. The test now checks the API-level contract only: score rows are processed tile by tile under a constrained score budget, and the source publishes compact [rows, topk_blocks] int32 block metadata with the documented 16K-row bound.

If this coverage looks useful, the standalone commit can be fetched with:

git fetch https://github.com/Frank-whw/sglang.git test-dsv41-prefill-candidate-metadata
git cherry-pick b00d103a28537e8570c9c66327650397ec05c629

Thanks again!

@YJR722
YJR722 force-pushed the YJR722/dsv41-sm90-triton-candidate-indexer branch from b670c59 to 528619e Compare September 28, 2026 07:17
@YJR722

YJR722 commented Sep 28, 2026

Copy link
Copy Markdown
Author

Rebased onto the latest main and adapted the SM90 implementation to the new v41_indexer interfaces. Updated the PR description with H20 correctness tests, GPQA Diamond results, and prefill TTFT / DSPARK decode performance benchmarks. Ready for another review—thanks!

@Oasis-Git Oasis-Git mentioned this pull request Oct 2, 2026
11 of 41 tasks
@yuan-luo

yuan-luo commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Thanks for implementing SM90 candidate indexers. Could you resolve the conflicts after rebasing main?

Reuse the v41_indexer publish/consume protocols and compact block metadata. Score prefill candidates directly from shared BF16 K, and add candidate block addressing to the paged FP4 decode/verify scorer. Preserve the dense source path and independent Torch reference.

Enable length-based candidate CUDA graph variants on SM90 and add scorer, protocol, and graph-selection regression coverage.
@YJR722
YJR722 force-pushed the YJR722/dsv41-sm90-triton-candidate-indexer branch from 528619e to 1875c52 Compare October 8, 2026 02:26

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants