Skip to content

feat(spec): pick the DSpark static verify width per step - #16

Closed
jhinpan wants to merge 1 commit into
kevin-mii:dsv41-amd-mainfrom
jhinpan:feat/dspark-verify-width
Closed

jhinpan wants to merge 1 commit into
kevin-mii:dsv41-amd-mainfrom
jhinpan:feat/dspark-verify-width

Conversation

@jhinpan

@jhinpan jhinpan commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

@kevin-mii DSpark verifies gamma + 1 tokens per request at every batch size (6 on DeepSeek-V4.1-Flash). Past batch 8 the verify step cost grows with verify tokens faster than the tail draft positions add accepted tokens, so a single width is right only for small batches. A static-verify sweep on MI355X (draft block shrunk with the width) puts the best width at 6 up to batch 8, 4 at batch 64 and 3 at batch 128-256, where it decodes up to 18% faster than width 6.

The confidence-driven ragged verify (SGLANG_RAGGED_VERIFY_MODE=compact) does not run on V4.1 yet. The engram hasher and the low-ratio compressor state paths assume equal verify blocks per request. This PR adapts the width per step instead, with every request of the step sharing it, so those paths are untouched.

Modifications

  • SGLANG_DSPARK_VERIFY_WIDTHS (e.g. 3,4, off by default) plus --speculative-dspark-sps-table-path enables it. SGLANG_DSPARK_FORCE_VERIFY_WIDTH pins one captured width for A/B.
  • After the full-width graphs, the worker builds a target attention backend, decode graphs and a verify epilogue for each extra width. It swaps them onto the target model runner only around the verify forward. model_runner.py is unchanged.
  • A step at width w verifies the anchor plus the first w - 1 drafts of the unchanged gamma-token proposal. Accept, commit and the folded epilogue run at width w. The result carries w as its output row stride.
  • The width maximizes predicted committed tokens per second: the expected accept length from running per-position acceptance, times the SPS table's steps per second at bs * w verify tokens, with a 2% switch margin.
    • Acceptance counts come from a device histogram read with a fixed lag, so every TP rank picks the same width.
    • The full width runs until every draft position is measured, and every 64th step after, so the positions a narrow width never verifies stay current.
  • Grammar and logprob steps keep the full width. Unsupported setups fail at startup instead of falling back silently: compact/cap-accept, DP attention, PP, mamba targets, simulated acceptance, and the DSpark debug observers.
  • CPU unit tests cover the policy:
    • hazard estimation from widths that cannot observe the deeper positions;
    • full width until every position is measured;
    • width shrinking with batch size on the measured MI355X table;
    • the switch margin, forced width, and rejection of out-of-range widths.

Accuracy Tests

GSM8K (64 fixed questions, natural EOS, concurrency 64): 61-63/64 on every server in every arm. Mean accept length at a forced width matches a server whose draft block equals that width (2.80 at width 4 vs 2.77).

Speed Tests and Profiling

Setup:

  • MI355X x4, TP4/EP4, DeepSeek-V4.1-Flash dba1be0a, this branch at e2e824dc58.
  • Cookbook MI350X Low-Latency cell (DSpark block 5) with --cuda-graph-max-bs-decode 256, AITER built as in this branch's Dockerfile.
  • Both arms also carry [AMD] DeepSeek-V4.1 gfx950 kernel campaign #8 (int64 FlashMLA store) and fix(scheduler): stop the min-free-slots delay once every waiting request fits #14 (min-free-slots fix, without which DSpark stops at 254 running requests).
  • MI355X SPS table built with sglang.benchmark.dspark_sps_profiler (static verify, 4K prompts).
  • Workload: B distinct real-text 4096-token prompts, 1024 output tokens, temperature 0, fresh servers in mirrored order (f6, auto, f4, f3, f3, f4, auto, f6), four scored bursts per arm.

Decode tok/s while every request is decoding, with mean accept length:

B width 6 (today) width 4 width 3 auto auto vs width 6 (95% CI)
1 333 / 3.07 317 269 330 / 3.04 -1.0% [-3.4, +1.5]
8 1730 / 3.27 1690 1525 1724 / 3.27 -0.3% [-2.6, +1.8]
64 6462 / 3.22 6744 6115 6744 / 2.81 +4.4% [+3.4, +5.4]
128 8671 / 3.25 9375 9352 9322 / 2.44 +7.5% [+7.1, +7.9]
256 10759 / 3.25 11837 12117 12092 / 2.42 +12.4% [+12.3, +12.5]
  • Auto lands within 0.6% of the best fixed width at every batch size.
  • End to end including prefill: +3.6% (B=64), +5.9% (B=128), +6.4% (B=256).
  • Cost: the two extra widths add 0.67 GB of graph memory and about one minute of startup.
  • The draft still proposes 5 tokens. A sweep that shrinks the draft block to the same width is faster again at large batches: -1.0% (B=64), -1.5% (B=128), -4.6% (B=256) for keeping the full draft. Shrinking the draft per width is a follow-up.

Checklist


CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.

DSpark verifies gamma + 1 tokens per request at every batch size. Past
batch 8 the verify step cost grows with verify tokens faster than the
tail draft positions add accepted tokens, so a narrower verify commits
more tokens per second.

With SGLANG_DSPARK_VERIFY_WIDTHS (e.g. "3,4") the worker builds a target
attention backend, decode graphs and verify epilogue for each extra width
after the full-width graphs, and swaps them onto the target runner around
the verify forward. Each step verifies the first w positions of the
proposal, where w maximizes predicted committed tokens per second: the
expected accept length from running per-position acceptance, times the
SPS table steps per second at bs * w verify tokens. Acceptance counts are
read with a fixed lag so every TP rank picks the same width; the full
width is taken until every position is measured and periodically after.
Grammar and logprob steps keep the full width. Off by default.
@jhinpan

jhinpan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Closing here. It rebases cleanly onto dsv41-amd-5-integration; we will send it to sgl-project/sglang with measurements once sgl-project#41308 is merged.

@jhinpan jhinpan closed this Sep 29, 2026
@jhinpan

jhinpan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Reopened to track the verify-width work until it goes to the cookbook and sgl-project/sglang. It rebases cleanly onto dsv41-amd-5-integration + #18; with an MI355X SPS table built there, auto widths decode +9.1% / +10.2% / +14.9% faster at batch 64 / 128 / 256 on the High-Throughput cell (ABBA), unchanged at batch 1 and 8.

@jhinpan jhinpan reopened this Sep 29, 2026
@jhinpan

jhinpan commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by sgl-project#41994, which carries the per-step verify width onto main together with #18, re-measured on main (decode +9.6% / +10.4% / +14.6% at batch 64 / 128 / 256 on the High-Throughput cell, unchanged at batch 1 and 8).

@jhinpan jhinpan closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant