Skip to content

[Bugfix][KV Offload] Reuse in-flight async lookup probes - #55823

Merged
orozery merged 6 commits into
vllm-project:mainfrom
Alex-ai-future:chore/clean
Sep 14, 2026
Merged

orozery merged 6 commits into
vllm-project:mainfrom
Alex-ai-future:chore/clean

Conversation

@Alex-ai-future

Copy link
Copy Markdown
Contributor

Purpose

Fix duplicate same-key async existence probes in AsyncLookupManager.

When request A finishes while its probe is still in flight, cleanup retains the
probe state. A later request B can attach to the same probe and reuse its
verdict, so the backend is accessed once instead of once per request.

The probe lifecycle is explicitly modeled with LookupPhase:

  • PENDING: accumulated but not submitted.
  • IN_FLIGHT: submitted to the worker and awaiting a result.
  • RESOLVED: the result has been applied.

This keeps probe lifecycle management separate from the cached HIT/MISS verdict
and preserves generation checks for stale results. The change is scoped to
same-key existence lookup reuse; it does not duplicate stale-result correctness,
cross-key batching, or promotion/load coalescing work.

The duplicate-probe behavior was identified in this
PR #52103 comment,
which includes a deterministic regression demonstrating the extra backend
probe.

This change does not affect model outputs or accuracy, so model evaluation is
not applicable. AI assistance was used to implement and validate this change.

Test Plan

  • pre-commit run --files vllm/v1/kv_offload/tiering/async_lookup.py tests/v1/kv_offload/tiering/test_async_lookup.py
  • .venv/bin/python -m pytest tests/v1/kv_offload/tiering/test_async_lookup.py -v

Test Result

  • All selected pre-commit hooks passed.
  • 18 passed, 14 warnings.
  • Warnings are existing torch.jit.script_method deprecation warnings.
  • The full test suite was not run.

Alex-ai-future and others added 3 commits September 8, 2026 11:43
Retain submitted unresolved lookup state so replacement requests can share the existing probe. Preserve cleanup before submission and reclaim unclaimed results during flush and shutdown.

Validation: git diff --check passed. Tests and commit hooks were not run at user request while the environment is being prepared.

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Alex <jihui.huang@daocloud.io>
Replace the submitted flag with explicit pending, in-flight, and resolved phases so lookup reuse and cleanup decisions are represented by the probe lifecycle.

Validation: git diff --check passed. Tests and commit hooks were not run at the user's request while the environment is being prepared.

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Alex <jihui.huang@daocloud.io>
Document cached lookup verdicts and update the stale-generation test to submit the replacement probe before injecting its result.

Tests: .venv/bin/python -m pytest tests/v1/kv_offload/tiering/test_async_lookup.py -v

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Alex <jihui.huang@daocloud.io>
@mergify mergify Bot added the bug Something isn't working label Sep 8, 2026
Describe RESOLVED as a final verdict so the lifecycle comment covers both worker results and forced misses.

Tests: not rerun; comment-only change. Prior targeted test run passed (18 tests).

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: Alex <jihui.huang@daocloud.io>
@Alex-ai-future
Alex-ai-future marked this pull request as ready for review September 9, 2026 01:49

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@orozery

orozery commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

/ci run

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #87883 for commit a94c5032a795.

@orozery orozery added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 10, 2026
@github-actions

Copy link
Copy Markdown

✅ @Alex-ai-future, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@orozery

orozery commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #88122 for commit 3cd56accc3d8.

@orozery

orozery commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

/ci retry

@github-actions

Copy link
Copy Markdown

✅ Queued 8 failed job(s) for retry in Buildkite CI #88122.

@orozery
orozery merged commit c612e2b into vllm-project:main Sep 14, 2026
109 checks passed
ItsRoy69 pushed a commit to ItsRoy69/vllm that referenced this pull request Sep 15, 2026
…t#55823)

Signed-off-by: Alex <jihui.huang@daocloud.io>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
keneoneth pushed a commit to keneoneth/vllm that referenced this pull request Sep 16, 2026
…t#55823)

Signed-off-by: Alex <jihui.huang@daocloud.io>
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants