Skip to content

fix(scheduler): keep SWA pages for the running batch's next decode - #15

Closed
jhinpan wants to merge 1 commit into
kevin-mii:dsv41-amd-mainfrom
jhinpan:fix/swa-prefill-next-decode-headroom
Closed

jhinpan wants to merge 1 commit into
kevin-mii:dsv41-amd-mainfrom
jhinpan:fix/swa-prefill-next-decode-headroom

Conversation

@jhinpan

@jhinpan jhinpan commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

A request frees its out-of-window SWA pages only when it decodes (maybe_evict_swa, from its second decode on). When the scheduler runs prefills back to back, as it does for a burst of long prompts, every freshly prefilled request still holds its whole prompt in the SWA pool. SWAPrefillBudget reserves nothing for the running batch's next decode (unlike the full pool, whose total_offset charges the running requests' remaining tokens), so admission can fill the SWA pool. The next decode then fails check_decode_mem and retracts requests, which later prefill their 4K prompts again.

On DeepSeek-V4.1-Flash with 256 concurrent 4096-token prompts, this retracts 36-45 requests per burst while full-pool usage stays at 2% (the log shows swa token usage 1.00 followed by KV cache pool is full. Retract requests).

Modifications

  • PrefillAdder charges running_batch.new_tokens_required_next_decode(), the demand check_decode_mem tests, to the prefill budget through reserve_next_decode.
  • SWAPrefillBudget adds it to swa_offset for paged SWA pools. Request rings already hold decode room, and the plain pool's total_offset already covers it, so both are unchanged.
  • Unit test: a prefill that fits the free SWA pages only by taking the running batch's next-decode pages now waits.

Accuracy Tests

GSM8K (64 fixed questions, natural EOS, concurrency 64) on every server: 61-63/64 in both arms.

Speed Tests and Profiling

MI355X x4, TP4/EP4, DeepSeek-V4.1-Flash dba1be0a, this branch at e2e824dc58, AITER built as in this branch's Dockerfile, plus #8's int64 FlashMLA store fix in both arms. 256 distinct real-text 4096-token prompts submitted at once, 3072 forced output tokens, temperature 0; fresh servers, ABBA per cell, two scored bursts per server.

Cell Metric Base This PR Change (95% CI)
Low-Latency (DSpark, graph cap 256) retracted requests per server 36 0
burst output tok/s 6952 7029 +1.10% [+0.80, +1.44]
TTFT median / p99 15.9 / 31.8 s 15.8 / 30.4 s -1.51% [-1.64, -1.38] (median)
High-Throughput (no speculation), two ABBA runs retracted requests per server 45 0
burst output tok/s -0.10% [-2.63, +2.77]; -0.19% [-1.02, +0.68]
TTFT median / p99 15.9 / 32.1 s 15.5 / 30.2 s -2.48% [-2.71, -2.25]; -2.29% [-2.45, -2.10] (median)
decode tok/s at >= 250 running (server log) 10235 / 10264 10240 / 10264 unchanged

The client-side all-active window rate in the High-Throughput cell reads -2.7% / -2.8%. It is not a slower decode: the server's generation throughput at full occupancy is identical in both arms. It is a composition effect, because the base's retracted requests restart later and move the window boundaries.

Checklist


CI States

Latest PR Test (Base): ❌ Missing run-ci label -- add it to run CI tests.
Latest PR Test (Extra): ❌ Blocked -- run-ci is required first.
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.

Out-of-window SWA pages are released only when a request decodes. With
back-to-back prefills (a large burst of long prompts), freshly prefilled
requests keep their whole prompt in the SWA pool until their second decode,
and the prefill budget reserved nothing for the running batch's next decode.
Admission could then fill the pool, so the next decode failed
check_decode_mem and retracted requests that had to be prefilled again.

Charge running_batch.new_tokens_required_next_decode() to the SWA budget
(non-ring pools), the same demand check_decode_mem tests.
@jhinpan

jhinpan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Closing: not needed on dsv41-amd-5-integration. Its SWA pool holds 724K tokens (230K on dsv41-amd-main), and a 256 x 4K burst peaks at 28% SWA usage, so no request is retracted with or without this change.

@jhinpan jhinpan closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant