Skip to content

[Perf][KV Offload] Publish async lookup results per request group - #57474

Draft
Alex-ai-future wants to merge 1 commit into
vllm-project:mainfrom
Alex-ai-future:fix/kv-offload-eager-lookup-publication
Draft

Alex-ai-future wants to merge 1 commit into
vllm-project:mainfrom
Alex-ai-future:fix/kv-offload-eager-lookup-publication

Conversation

@Alex-ai-future

Copy link
Copy Markdown
Contributor

[Perf][KV Offload] Publish async lookup results per request group

Summary

AsyncLookupManager already performs lookup work asynchronously, but it
currently publishes the results as one batch for the whole worker batch. When
the batch contains multiple request groups, a slow group delays the scheduler
from observing results from groups that have already completed.

Publish each completed request group's results immediately, before processing
the next group. This exposes completed lookup work at finer granularity and
allows the scheduler and tiering manager to make progress while later groups
are still being looked up.

The change preserves the existing worker FIFO order and does not change
backend concurrency, lookup semantics, or promotion policy.

Correctness

  • Lazy backend iterables are fully consumed before their group is published.
  • Backend exceptions are still converted to misses.
  • Request-group ordering, generation checks, duplicate-result handling, and
    shutdown behavior are unchanged.

Validation

Compared with the direct parent a529c1a748, candidate
8b556e27f9 publishes and resolves a fast group before a blocked later group
in the controlled test, while preserving FIFO behavior when the slow group is
first. In the tiering harness, this lets a completed hit begin promotion
before the later group's lookup finishes. The current branch contains the same
patch rebased onto the latest main.

.venv/bin/python -m pytest \
  tests/v1/kv_offload/tiering/test_async_lookup.py -q
# 22 passed

.venv/bin/python -m pytest \
  tests/v1/kv_offload/tiering/test_fs_tier.py \
  tests/v1/kv_offload/tiering/test_obj_tier.py -q
# 82 passed

.venv/bin/python -m pytest \
  tests/v1/kv_offload/tiering/test_tiering_offloading.py -q
# 54 passed

pre-commit run --files \
  vllm/v1/kv_offload/tiering/async_lookup.py \
  tests/v1/kv_offload/tiering/test_async_lookup.py

No model evaluation was run because this change affects lookup result
publication timing, not model outputs, accuracy, or serving semantics.

I searched the open vLLM PRs for async lookup and KV-offload publication work
and found no PR covering this per-request-group publication change. The older
#23622 is a non-merge PoC for different connector lookup semantics.

AI assistance was used. The human submitter reviewed the change and is
responsible for the design and test results.

Co-authored-by: OpenAI Codex <noreply@openai.com>

Signed-off-by: Alex <jihui.huang@daocloud.io>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant