Repository navigation
Conversation
Out-of-window SWA pages are released only when a request decodes. With back-to-back prefills (a large burst of long prompts), freshly prefilled requests keep their whole prompt in the SWA pool until their second decode, and the prefill budget reserved nothing for the running batch's next decode. Admission could then fill the pool, so the next decode failed check_decode_mem and retracted requests that had to be prefilled again. Charge running_batch.new_tokens_required_next_decode() to the SWA budget (non-ring pools), the same demand check_decode_mem tests.
4 of 5 tasks
Collaborator
Author
|
Closing: not needed on |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
A request frees its out-of-window SWA pages only when it decodes (
maybe_evict_swa, from its second decode on). When the scheduler runs prefills back to back, as it does for a burst of long prompts, every freshly prefilled request still holds its whole prompt in the SWA pool.SWAPrefillBudgetreserves nothing for the running batch's next decode (unlike the full pool, whosetotal_offsetcharges the running requests' remaining tokens), so admission can fill the SWA pool. The next decode then failscheck_decode_memand retracts requests, which later prefill their 4K prompts again.On DeepSeek-V4.1-Flash with 256 concurrent 4096-token prompts, this retracts 36-45 requests per burst while full-pool usage stays at 2% (the log shows
swa token usage1.00 followed byKV cache pool is full. Retract requests).Modifications
PrefillAdderchargesrunning_batch.new_tokens_required_next_decode(), the demandcheck_decode_memtests, to the prefill budget throughreserve_next_decode.SWAPrefillBudgetadds it toswa_offsetfor paged SWA pools. Request rings already hold decode room, and the plain pool'stotal_offsetalready covers it, so both are unchanged.Accuracy Tests
GSM8K (64 fixed questions, natural EOS, concurrency 64) on every server: 61-63/64 in both arms.
Speed Tests and Profiling
MI355X x4, TP4/EP4, DeepSeek-V4.1-Flash
dba1be0a, this branch ate2e824dc58, AITER built as in this branch's Dockerfile, plus #8's int64 FlashMLA store fix in both arms. 256 distinct real-text 4096-token prompts submitted at once, 3072 forced output tokens, temperature 0; fresh servers, ABBA per cell, two scored bursts per server.The client-side all-active window rate in the High-Throughput cell reads -2.7% / -2.8%. It is not a slower decode: the server's generation throughput at full occupancy is identical in both arms. It is a composition effect, because the base's retracted requests restart later and move the window boundaries.
Checklist
CI States
Latest PR Test (Base): ❌ Missing
run-cilabel -- add it to run CI tests.Latest PR Test (Extra): ❌ Blocked --
run-ciis required first.Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.