Repository navigation
Conversation
DSpark verifies gamma + 1 tokens per request at every batch size. Past batch 8 the verify step cost grows with verify tokens faster than the tail draft positions add accepted tokens, so a narrower verify commits more tokens per second. With SGLANG_DSPARK_VERIFY_WIDTHS (e.g. "3,4") the worker builds a target attention backend, decode graphs and verify epilogue for each extra width after the full-width graphs, and swaps them onto the target runner around the verify forward. Each step verifies the first w positions of the proposal, where w maximizes predicted committed tokens per second: the expected accept length from running per-position acceptance, times the SPS table steps per second at bs * w verify tokens. Acceptance counts are read with a fixed lag so every TP rank picks the same width; the full width is taken until every position is measured and periodically after. Grammar and logprob steps keep the full width. Off by default.
2 of 4 tasks
4 of 5 tasks
Collaborator
Author
|
Closing here. It rebases cleanly onto |
Collaborator
Author
|
Reopened to track the verify-width work until it goes to the cookbook and sgl-project/sglang. It rebases cleanly onto |
Collaborator
Author
|
Superseded by sgl-project#41994, which carries the per-step verify width onto main together with #18, re-measured on main (decode +9.6% / +10.4% / +14.6% at batch 64 / 128 / 256 on the High-Throughput cell, unchanged at batch 1 and 8). |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
@kevin-mii DSpark verifies
gamma + 1tokens per request at every batch size (6 on DeepSeek-V4.1-Flash). Past batch 8 the verify step cost grows with verify tokens faster than the tail draft positions add accepted tokens, so a single width is right only for small batches. A static-verify sweep on MI355X (draft block shrunk with the width) puts the best width at 6 up to batch 8, 4 at batch 64 and 3 at batch 128-256, where it decodes up to 18% faster than width 6.The confidence-driven ragged verify (
SGLANG_RAGGED_VERIFY_MODE=compact) does not run on V4.1 yet. The engram hasher and the low-ratio compressor state paths assume equal verify blocks per request. This PR adapts the width per step instead, with every request of the step sharing it, so those paths are untouched.Modifications
SGLANG_DSPARK_VERIFY_WIDTHS(e.g.3,4, off by default) plus--speculative-dspark-sps-table-pathenables it.SGLANG_DSPARK_FORCE_VERIFY_WIDTHpins one captured width for A/B.model_runner.pyis unchanged.wverifies the anchor plus the firstw - 1drafts of the unchanged gamma-token proposal. Accept, commit and the folded epilogue run at widthw. The result carrieswas its output row stride.bs * wverify tokens, with a 2% switch margin.Accuracy Tests
GSM8K (64 fixed questions, natural EOS, concurrency 64): 61-63/64 on every server in every arm. Mean accept length at a forced width matches a server whose draft block equals that width (2.80 at width 4 vs 2.77).
Speed Tests and Profiling
Setup:
dba1be0a, this branch ate2e824dc58.--cuda-graph-max-bs-decode 256, AITER built as in this branch's Dockerfile.sglang.benchmark.dspark_sps_profiler(static verify, 4K prompts).Decode tok/s while every request is decoding, with mean accept length:
Checklist
CI States
Latest PR Test (Base): ❌ Missing
run-cilabel -- add it to run CI tests.Latest PR Test (Extra): ❌ Blocked --
run-ciis required first.Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.