Repository navigation
Conversation
Co-authored-by: Codex <noreply@openai.com> Signed-off-by: Vegetog <110553275+Vegetog@users.noreply.github.com>
Vegetog
requested review from
ApostaC,
WoosukKwon,
alexm-redhat,
heheda12345,
ivanium,
njhill,
orozery,
robertgshaw2-redhat and
ywang96
as code owners
October 8, 2026 07:59
1 task done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Related to #60552 (partial fix). Return
Falsewhen a forced prefix-cache reset still has referenced blocks, so/reset_prefix_cache?reset_running_requests=truereturns the documented retryable200 {"success": false}instead of HTTP 500.Merging this PR alone fixes the reset response, but does not make the full #60552 offloading reproducer pass.
success: falsemeans the cache reset did not complete; referenced blocks remain in use and the caller can retry later. When a preempted request resumes, the existing OffloadingConnector assertion can still crash the engine and interrupt generation.The full reproduced scenario also needs the companion connector fix in #60574, or an equivalent fix for the same bug covered by #57810. This is a requirement for end-to-end recovery, not a required merge order: the two PRs address separate root causes.
A parked streaming-input session (
WAITING_FOR_STREAMING_REQ, #60744) is a different holder: it waits on the client rather than on a transfer, so with this PR alone the reset keeps returningFalsehowever often it is retried; #60764 reclaims those sessions in the same function, so the two changes are complementary.Claims
Validation
Standalone patch on
0f112d180033f3757cf283723f5593a40b02ed82:.venv/bin/python -m pytest \ tests/v1/core/test_scheduler.py \ tests/v1/core/test_async_scheduler.py \ tests/v1/engine/test_engine_core.py \ -k 'reset_prefix_cache or inflight_remote_kv or aux_output_reset or reset_connector_cache or kv_cache_release or pause_synchronizes' -v19 passed. The four new regression cases failed before the fix. Pre-commit passed for both changed files.
The original #60552 reproducer returned HTTP 500 in 5/5 flagged trials on the initial baseline
fb2ac824. With this patch alone, the retryable HTTP response was observed, but repeated runs also exposed the pre-existing offloading preempt/resume crash described in #57810. That crash is addressed separately by #60574 and is not fixed by this PR.Combined serving validation on the same base, with both this patch and #60574 applied: one RTX 3070, Qwen3-0.6B, eager execution with
TRITON_ATTN, 320 GPU blocks and 4 GB CPU offload. The Python checkout reused the official v0.31.0+cu129 CUDA extensions.ignore_eos=Trueand streamed usage checks: 3 flagged and 3 control resets returned200 {"success": false}; all 12 streams completed 64 tokens, all 6 follow-up health/generation checks passed, and every trial recorded positive offload-load bytes.POST /reset_prefix_cache?reset_running_requests=truereturns HTTP 500 while an OffloadingConnector load is in flight #60552 client then completed 5 flagged and 5 control trials on the same server, all HTTP 200 and healthy.These are combined-patch serving results, not a claim that this PR alone fixes the separate connector crash.
No model-quality or performance claim is made. The serving checks validate HTTP responses, stream completion and continued generation; the results above are local validation, not a claim of passing upstream test CI.
Details
Preempting
self.runningdoes not release blocks owned byWAITING_FOR_REMOTE_KVSrequests. The cache manager already reports this condition asFalse; the scheduler currently converts it into an exception only whenreset_running_requests=True.Return
Falseat the same point without changing preemption or transfer ownership. Keep the early return so a failed forced reset does not start a connector reset.EngineCore._reset_caches()already checks the result, so no pause/sleep caller changes are needed.The duplicate-work check before opening this PR found no existing fix for this HTTP return-value contract. Companion #60574 and the overlapping fix in #57810 address the separate same-step offloading resume crash. This PR changes only the scheduler and its existing tests.
Developed and validated with OpenAI Codex assistance.
Pull Request Checklist
I used vLLM's
/pr-checklistskill. (Mandatory for agents, optional for humans).AI assistance was used during the creation of this PR.
Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.
Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.
Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.
Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.