Repository navigation
Conversation
|
Please fix the lint first. |
22056a5 to
c2c2e2d
Compare
|
Thanks for this work. I prepared two small follow-up commits on top of the current PR head:
Branch: https://github.com/Frank-whw/sglang/tree/test-dsv41-h200-prefill-memory Would you be willing to cherry-pick them in this order? git fetch https://github.com/Frank-whw/sglang.git test-dsv41-h200-prefill-memory
git cherry-pick 775ab53e1a54d5574b4106476951944f5ead5465 170214f81ae37fa95dcfcdadb9f5377680d4b009I can provide the detailed validation notes privately if useful. |
|
@Frank-whw Thanks for preparing these follow-up patches! |
|
@YJR722 Thanks for the careful review. I agree: I prepared a clean standalone replacement test on top of the current PR head, without
It removes the blanket If this coverage looks useful, the standalone commit can be fetched with: git fetch https://github.com/Frank-whw/sglang.git test-dsv41-prefill-candidate-metadata
git cherry-pick b00d103a28537e8570c9c66327650397ec05c629Thanks again! |
b670c59 to
528619e
Compare
|
Rebased onto the latest main and adapted the SM90 implementation to the new |
|
Thanks for implementing SM90 candidate indexers. Could you resolve the conflicts after rebasing main? |
Reuse the v41_indexer publish/consume protocols and compact block metadata. Score prefill candidates directly from shared BF16 K, and add candidate block addressing to the paged FP4 decode/verify scorer. Preserve the dense source path and independent Torch reference. Enable length-based candidate CUDA graph variants on SM90 and add scorer, protocol, and graph-selection regression coverage.
528619e to
1875c52
Compare
Motivation
Add candidate-only indexer scoring on SM90 for DeepSeek-V4.1, following #40574. The current
v41_indexerimplementation already provides unified prefill/decode protocols and compact candidate block IDs, but Hopper consumers still score the full context before selecting among those blocks.DeepGEMM's paged sparse FP4 MQA scorer is not available on SM90. An earlier approach expanded each query's block IDs and gathered its dequantized K rows into a
[queries, candidates, head_dim]tensor before Torch scoring. This PR instead uses Triton to read candidate K positions directly, avoiding that per-query K tensor. A future SM90 DeepGEMM sparse scorer could replace this implementation behind the existing protocol.Modifications
DenseBlocksBackend,BlockIds,Selection, and the existingpublish_*/consume_*interfaces from the upstream refactor, without introducing new public input or candidate metadata classes.fp4_index_logits_decodewith optional candidate block addressing for decode/verify consumers. Preserve the upstream dense invisible-tile shortcut. The existing context-sized slot map is still constructed.Accuracy Tests
The regression tests added in this PR passed on H20 at
528619e4b2: 37 tests passed. They cover BF16/FP4 candidate scoring against reference implementations, prefill protocol equivalence, causal masking, padding, index mapping, tail slicing, and CUDA graph replay/selection.GPQA Diamond: 198 questions, one sample per question, evaluated with
sgl-eval. Both arms use the same 1P1D H20-3e deployment (TP8/EP8/DP1/CP1), DSPARK block5 and decode full CUDA graphs, with both P and D running the corresponding revision. Parameters: concurrency=16, thinking enabled,reasoning_effort=max, temperature=1.0, top_p=0.95,max_tokens=65536, seed=1.cc012abd21)528619e4b2)max_tokensSpeed Tests and Profiling
Baseline A:
cc012abd21. This PR B:528619e4b2.Both experiments use DeepSeek-V4.1-Flash in a real 1P1D deployment, with two nodes of 8×H20-3e, TP8/EP8/DP1/CP1.
Prefill TTFT
D is fixed; only P switches between A/B. FlashInfer MXFP4 with FP8 compute, chunk=1024, concurrency=1, OSL=1, speculation and CUDA graphs disabled. Each length uses identical input IDs, cache flushes, one warmup, and three measured requests. Values are median streaming TTFT.
16K is effectively unchanged; gains increase with context length.
DSPARK decode/verify throughput and effective TPOT
P is fixed; only D switches between A/B. DSPARK block5, static verify, full CUDA graphs, steady-state BS8, and OSL limit=40,000. Each configuration runs three times with identical inputs; measurements are taken while the active batch size remains stable at 8.
Results are three-run averages.
Checklist
CI States
Latest PR Test (Base): ❌ Run #37717868631
Latest PR Test (Extra): ❌ Run #37717868167
Latest PR Test (AMD ROCm 10): ❌ Run #37717868527