Skip to content

perf(amd): faster DSpark scheduling for DeepSeek-V4.1 on gfx950 - #18

Closed
jhinpan wants to merge 2 commits into
kevin-mii:dsv41-amd-5-integrationfrom
jhinpan:perf/dsv41-amd-dspark-scheduler
Closed

jhinpan wants to merge 2 commits into
kevin-mii:dsv41-amd-5-integrationfrom
jhinpan:perf/dsv41-amd-dspark-scheduler

Conversation

@jhinpan

@jhinpan jhinpan commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Two of our changes to the previous DSv4.1 AMD branch (dsv41-amd-main) still apply to this branch and still make DeepSeek-V4.1-Flash with DSpark faster on MI355X:

  • Batch 1 to 64 decode. With overlap scheduling and speculative decoding on HIP, FutureMap.resolve_seq_lens_cpu and resolve_mixed_spec_tails block the host on the previous forward's publish event (publish_ready.synchronize()), so the GPU idles between the DSpark verify and the next draft. Replacing the sync with Event.wait() on the separate schedule stream is slower still (1.2 ms more per verify on dsv41-amd-main): a device wait pending in another HIP queue while a graph runs slows every dispatch of that graph.
  • 256 concurrent requests. MinFreeSlotsDelayer, auto-enabled for DSpark, holds fresh prefills until min(4, max(2, (max_running_requests + 5) // 6)) slots are free. On the High-Throughput cell a burst of 256 settles at 253 running with 3 queued, and those 3 wait for the whole decode of the running batch: 75 s maximum TTFT.

Our other dsv41-amd-main PRs are not needed here. #8's int64 KV-store offsets are upstream as sgl-project#41159. #11's per-pool sparse PA loads need an AITER pin that includes ROCm/aiter#4919. #13 (draft-block metadata in the draft graph) is 2.9% faster alone at batch 1 but adds nothing measurable once the host no longer blocks: +0.14% [-0.03, +0.31] per verify at batch 1, -0.10% to +0.16% at batch 8-64. #15's SWA retractions do not occur on this branch: the SWA pool holds 724K tokens (230K on dsv41-amd-main), and a 256 x 4K burst peaks at 28%. #16 (per-step verify width) is opt-in and needs an SPS table, so it is not part of this PR.

Modifications

  • perf(amd): keep speculative overlap scheduling on the forward stream: on HIP with overlap scheduling, speculative decoding and no pipeline parallelism, the scheduler runs on the forward stream and the FutureMap publish waits become in-queue Event.wait(). Stream order carries the forward-to-schedule dependencies, so the host no longer blocks and nothing waits in another queue while the forward graph runs. Without speculation there is no publish wait, so the separate schedule stream stays.
  • fix(scheduler): stop the min-free-slots delay once every waiting request fits: MinFreeSlotsDelayer.should_delay takes the number of waiting requests and delays only while the free slots are fewer than both the threshold and the waiting requests. With at least threshold requests waiting the behavior is unchanged. The unit test covers 254 running with 2 free: 2 waiting are admitted, 3 still wait.

Accuracy Tests

  • GSM8K, full test set (1319 questions, natural EOS, concurrency 128), two fresh High-Throughput servers per arm: 1280 and 1280 for the base, 1280 and 1280 with this PR (97.04% both, 2 truncated per server). The failing questions differ between servers: two base servers agree on 1307 answers, a base and a PR server on 1298-1305.
  • Batch 1, 64 fixed GSM8K questions and 8 retrieval probes per server: 61-63/64 and 8/8 in both arms, no truncation. On this branch batch-1 outputs already differ between fresh servers of the same arm (11-17 of 64 identical GSM8K sequences between two base servers), so output identity is not used.
  • The FutureMap consume-once checks (the SGLANG_IS_IN_CI debug path, forced on in a local build) pass with this change at batch 1, 8 and 64 plus GSM8K, and fault at warmup when publish() skips its write.

Speed Tests and Profiling

4x MI355X (gfx950), TP4/EP4, DeepSeek-V4.1-Flash dba1be0a, this branch at fdf14605e5, the MI350X cookbook cells of this branch (Low-Latency: DSpark block 5, decode graphs up to 128; High-Throughput: graphs up to 256, 256 running), docker/rocm.Dockerfile's runtime environment plus the cells' AITER_FLYDSL_FORCE_REDUCE=1 AITER_BF16_FP8_MOE_BOUND=0, sgl-kernel built from kernels/aot, AITER acf8fdf9 as the Dockerfile builds it. Fresh server per arm, real acceptance, radix cache on and flushed before every measured request or burst, no profiler during timing.

Workload Metric Base This PR Change (95% CI)
Batch 1, 4 real-text 4K prompts x 24 per server, 1024 tokens, Low-Latency decode tok/s 351.7 375.2 +6.69% [+5.35, +8.07]
wall per verify 9.19 ms 8.60 ms
Batch 8, distinct real-text 4K prompts, 1024 tokens, Low-Latency verify cycles/s 66.78 69.57 +4.18% [+3.51, +4.84]
Batch 32 verify cycles/s 41.10 42.21 +2.72% [+2.32, +3.17]
Batch 64 verify cycles/s 30.62 31.14 +1.72% [+1.48, +1.94]
256 distinct 4K prompts at once, 3072 tokens, High-Throughput output tok/s 7682 8347 +8.67% [+8.38, +8.96]
TTFT max 74.9 s 15.7 s -79.1% [-79.4, -78.8]
TTFT median 8.06 s 8.10 s +0.46% [+0.23, +0.69]
TPOT median 26.4 ms 26.1 ms -0.97% [-1.83, -0.24]

Batched DSpark rows compare verify cycles per second (decode tok/s / (batch x acceptance length), inside the window where every request is decoding) because batched outputs, and with them acceptance, vary between runs in both arms; decode tok/s moves by +4.25%, +2.19% and +1.92%. Batch 1 and 8-64 are ABCCBA sessions (the middle arm was this PR plus #13); the 256-request rows are ABBA.

Each change against the same base: the forward-stream scheduling is +6.86% [+6.01, +7.65] decode at batch 1 and +5.56%, +2.76%, +2.05% cycles at batch 8, 32, 64; the delayer fix is +7.68% [+7.08, +8.12] output tok/s on the 256-request burst, maximum TTFT 75 s -> 15.6 s, median TTFT unchanged (+0.07% [-0.09, +0.34]), and changes nothing while at least 4 request slots are free.

Checklist


CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.

With overlap scheduling and speculative decoding on ROCm, the FutureMap
publish waits block the host (publish_ready.synchronize()): an Event.wait()
on the separate schedule stream leaves a cross-queue wait pending while
the forward graph runs, and that slows every dispatch of the graph.

Schedule on the forward stream instead, so stream order carries the
forward-to-schedule dependencies and the publish waits stay in one queue.
Without speculation there is no publish wait and the separate schedule
stream keeps overlapping the next batch's preparation, so it stays.
…est fits

The DFlash-family delayer (auto-enabled for DSpark, threshold 4 at 256
running requests) held fresh prefills until four request slots were free.
When fewer requests than that were waiting and all of them fit, nothing more
could be batched, yet they waited until running requests finished: at 254
running with 2 queued, those 2 waited for the whole decode of the batch.

Delay only while the free slots are fewer than the threshold and fewer than
the waiting requests.
@jhinpan

jhinpan commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by sgl-project#41994: DeepSeek-V4.1 AMD support is on main now (sgl-project#41308), so both commits go there directly, re-measured on main (batch 1 decode +6.1%, 256-request burst output +8.6%, max TTFT 75 s -> 15.7 s).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant