Repository navigation
Conversation
With overlap scheduling and speculative decoding on ROCm, the FutureMap publish waits block the host (publish_ready.synchronize()): an Event.wait() on the separate schedule stream leaves a cross-queue wait pending while the forward graph runs, and that slows every dispatch of the graph. Schedule on the forward stream instead, so stream order carries the forward-to-schedule dependencies and the publish waits stay in one queue. Without speculation there is no publish wait and the separate schedule stream keeps overlapping the next batch's preparation, so it stays.
…est fits The DFlash-family delayer (auto-enabled for DSpark, threshold 4 at 256 running requests) held fresh prefills until four request slots were free. When fewer requests than that were waiting and all of them fit, nothing more could be batched, yet they waited until running requests finished: at 254 running with 2 queued, those 2 waited for the whole decode of the batch. Delay only while the free slots are fewer than the threshold and fewer than the waiting requests.
This was referenced Sep 29, 2026
Collaborator
Author
|
Superseded by sgl-project#41994: DeepSeek-V4.1 AMD support is on main now (sgl-project#41308), so both commits go there directly, re-measured on main (batch 1 decode +6.1%, 256-request burst output +8.6%, max TTFT 75 s -> 15.7 s). |
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Two of our changes to the previous DSv4.1 AMD branch (
dsv41-amd-main) still apply to this branch and still make DeepSeek-V4.1-Flash with DSpark faster on MI355X:FutureMap.resolve_seq_lens_cpuandresolve_mixed_spec_tailsblock the host on the previous forward's publish event (publish_ready.synchronize()), so the GPU idles between the DSpark verify and the next draft. Replacing the sync withEvent.wait()on the separate schedule stream is slower still (1.2 ms more per verify ondsv41-amd-main): a device wait pending in another HIP queue while a graph runs slows every dispatch of that graph.MinFreeSlotsDelayer, auto-enabled for DSpark, holds fresh prefills untilmin(4, max(2, (max_running_requests + 5) // 6))slots are free. On the High-Throughput cell a burst of 256 settles at 253 running with 3 queued, and those 3 wait for the whole decode of the running batch: 75 s maximum TTFT.Our other
dsv41-amd-mainPRs are not needed here. #8's int64 KV-store offsets are upstream as sgl-project#41159. #11's per-pool sparse PA loads need an AITER pin that includes ROCm/aiter#4919. #13 (draft-block metadata in the draft graph) is 2.9% faster alone at batch 1 but adds nothing measurable once the host no longer blocks: +0.14% [-0.03, +0.31] per verify at batch 1, -0.10% to +0.16% at batch 8-64. #15's SWA retractions do not occur on this branch: the SWA pool holds 724K tokens (230K ondsv41-amd-main), and a 256 x 4K burst peaks at 28%. #16 (per-step verify width) is opt-in and needs an SPS table, so it is not part of this PR.Modifications
perf(amd): keep speculative overlap scheduling on the forward stream: on HIP with overlap scheduling, speculative decoding and no pipeline parallelism, the scheduler runs on the forward stream and the FutureMap publish waits become in-queueEvent.wait(). Stream order carries the forward-to-schedule dependencies, so the host no longer blocks and nothing waits in another queue while the forward graph runs. Without speculation there is no publish wait, so the separate schedule stream stays.fix(scheduler): stop the min-free-slots delay once every waiting request fits:MinFreeSlotsDelayer.should_delaytakes the number of waiting requests and delays only while the free slots are fewer than both the threshold and the waiting requests. With at leastthresholdrequests waiting the behavior is unchanged. The unit test covers 254 running with 2 free: 2 waiting are admitted, 3 still wait.Accuracy Tests
SGLANG_IS_IN_CIdebug path, forced on in a local build) pass with this change at batch 1, 8 and 64 plus GSM8K, and fault at warmup whenpublish()skips its write.Speed Tests and Profiling
4x MI355X (gfx950), TP4/EP4, DeepSeek-V4.1-Flash
dba1be0a, this branch atfdf14605e5, the MI350X cookbook cells of this branch (Low-Latency: DSpark block 5, decode graphs up to 128; High-Throughput: graphs up to 256, 256 running),docker/rocm.Dockerfile's runtime environment plus the cells'AITER_FLYDSL_FORCE_REDUCE=1 AITER_BF16_FP8_MOE_BOUND=0, sgl-kernel built fromkernels/aot, AITERacf8fdf9as the Dockerfile builds it. Fresh server per arm, real acceptance, radix cache on and flushed before every measured request or burst, no profiler during timing.Batched DSpark rows compare verify cycles per second (decode tok/s / (batch x acceptance length), inside the window where every request is decoding) because batched outputs, and with them acceptance, vary between runs in both arms; decode tok/s moves by +4.25%, +2.19% and +1.92%. Batch 1 and 8-64 are ABCCBA sessions (the middle arm was this PR plus #13); the 256-request rows are ABBA.
Each change against the same base: the forward-stream scheduling is +6.86% [+6.01, +7.65] decode at batch 1 and +5.56%, +2.76%, +2.05% cycles at batch 8, 32, 64; the delayer fix is +7.68% [+7.08, +8.12] output tok/s on the 256-request burst, maximum TTFT 75 s -> 15.6 s, median TTFT unchanged (+0.07% [-0.09, +0.34]), and changes nothing while at least 4 request slots are free.
Checklist
CI States
Latest PR Test (Base): ❌ Missing
run-cilabel -- add it to run CI tests.Latest PR Test (Extra): ❌ Blocked --
run-ciis required first.Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.