Skip to content

perf(amd): keep speculative overlap scheduling on the forward stream - #12

Closed
jhinpan wants to merge 1 commit into
kevin-mii:dsv41-amd-mainfrom
jhinpan:perf/amd-hip-forward-stream-overlap
Closed

jhinpan wants to merge 1 commit into
kevin-mii:dsv41-amd-mainfrom
jhinpan:perf/amd-hip-forward-stream-overlap

Conversation

@jhinpan

@jhinpan jhinpan commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

With overlap scheduling and speculative decoding on ROCm, FutureMap.resolve_seq_lens_cpu and resolve_mixed_spec_tails block the host on the previous forward's publish event (publish_ready.synchronize()). The scheduler cannot prepare the next step while the verify graph runs, so the GPU idles between the DSpark verify and the next draft.

Event.wait() was rejected for regressing TPOT. The cost comes from where that wait sits: the schedule stream is another HIP queue, and a device wait left pending in another queue while a HIP graph runs slows every dispatch of that graph. On MI355X a 512-kernel graph takes 1541 us instead of 853 us with such a wait pending, and swapping the host sync for Event.wait() alone made the DSpark verify graph 1.2 ms slower per cycle.

Modifications

  • On HIP, with overlap scheduling, speculative decoding and no pipeline parallelism, the scheduler runs on the forward stream and the FutureMap publish waits become in-queue Event.wait(). Stream order carries the forward-to-schedule dependencies: the host no longer blocks, and nothing waits in another queue while the forward graph runs.
  • Without speculation there is no publish wait, so the separate schedule stream stays. Moving it onto the forward stream cost 0.2-0.3% at batch 8-256 because the next batch's preparation lost its overlap.

Accuracy Tests

GSM8K (64 fixed questions, natural EOS) and 8 retrieval probes on every server: 62/64 GSM8K and 8/8 retrieval in both arms. On this base batch-1 outputs already differ between fresh servers of the same arm (base vs base: 15/64 identical GSM8K sequences, 62/64 same answer), and the base-vs-PR agreement stays in that range (12-20/64 identical, 62/64 same answer).

Speed Tests and Profiling

MI355X x4, TP4/EP4, DeepSeek-V4.1-Flash dba1be0a, this branch at e2e824dc58, the cookbook MI350X Low-Latency cell (DSpark block 5), AITER built as in this branch's Dockerfile. Both arms also carry #8's int64 FlashMLA store fix. Fresh servers per arm, real acceptance, no profiler during timing.

Workload Metric Base This PR Change (95% CI)
Batch 1, 4 real-text 4K prompts x 24 per server, ABBA decode tok/s 359.3 380.4 +5.87% [+5.05, +6.63]
wall per verify 9.16 ms 8.63 ms
Batch 8, distinct real-text 4K prompts, 1024 tokens, ABCCBA cycle time 15.15 ms 14.52 ms -4.2% [-4.8, -3.5]
Batch 32 cycle time 24.48 ms 24.01 ms -1.9% [-2.4, -1.3]
Batch 64 cycle time 31.75 ms 31.20 ms -1.6% [-1.9, -1.3]

Batched runs report cycle time (batch x acceptance length / decode tok/s inside the window where every request is decoding) because batched outputs, and with them acceptance, vary run to run in both arms.

Checklist


CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 23, 2026
With overlap scheduling and speculative decoding on ROCm, the FutureMap
publish waits block the host (publish_ready.synchronize()): an Event.wait()
on the separate schedule stream leaves a cross-queue wait pending while
the forward graph runs, and that slows every dispatch of the graph.

Schedule on the forward stream instead, so stream order carries the
forward-to-schedule dependencies and the publish waits stay in one queue.
Without speculation there is no publish wait and the separate schedule
stream keeps overlapping the next batch's preparation, so it stays.
@jhinpan
jhinpan force-pushed the perf/amd-hip-forward-stream-overlap branch from e21179c to ff356eb Compare September 24, 2026 04:51
@jhinpan

jhinpan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Carried onto dsv41-amd-5-integration (the sgl-project#41308 head) in #18 and re-measured there. Closing in favor of #18.

@jhinpan jhinpan closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant