Repository navigation
[Perf][KV Offload] Publish async lookup results per request group - #57474
Draft
Alex-ai-future wants to merge 1 commit into
Draft
Alex-ai-future wants to merge 1 commit into
Alex-ai-future wants to merge 1 commit into
Conversation
Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: Alex <jihui.huang@daocloud.io>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
[Perf][KV Offload] Publish async lookup results per request group
Summary
AsyncLookupManageralready performs lookup work asynchronously, but itcurrently publishes the results as one batch for the whole worker batch. When
the batch contains multiple request groups, a slow group delays the scheduler
from observing results from groups that have already completed.
Publish each completed request group's results immediately, before processing
the next group. This exposes completed lookup work at finer granularity and
allows the scheduler and tiering manager to make progress while later groups
are still being looked up.
The change preserves the existing worker FIFO order and does not change
backend concurrency, lookup semantics, or promotion policy.
Correctness
shutdown behavior are unchanged.
Validation
Compared with the direct parent
a529c1a748, candidate8b556e27f9publishes and resolves a fast group before a blocked later groupin the controlled test, while preserving FIFO behavior when the slow group is
first. In the tiering harness, this lets a completed hit begin promotion
before the later group's lookup finishes. The current branch contains the same
patch rebased onto the latest
main.No model evaluation was run because this change affects lookup result
publication timing, not model outputs, accuracy, or serving semantics.
I searched the open vLLM PRs for async lookup and KV-offload publication work
and found no PR covering this per-request-group publication change. The older
#23622 is a non-merge PoC for different connector lookup semantics.
AI assistance was used. The human submitter reviewed the change and is
responsible for the design and test results.