Repository navigation
[Feature][Core] Support draining in-process requests before sleep - #56754
Open
zupengwang wants to merge 1 commit into
Open
zupengwang wants to merge 1 commit into
zupengwang wants to merge 1 commit into
Conversation
Drive running requests and pending model batches to completion before sleeping, preserving waiting requests and buffered generation results. Keep stop-string processing active and reuse per-request output collectors. Cover sleep levels, output kinds, partial wake, and drain failures. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: Wang Zupeng <zupenwang@gmail.com>
Contributor
|
Documentation preview: https://vllm--56754.org.readthedocs.build/en/56754/ |
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
This was referenced Sep 18, 2026
This was referenced Sep 28, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Support
sleep(mode="wait")with the in-process engine (VLLM_ENABLE_V1_MULTIPROCESSING=0). This currently raises aValueError, preventing offline RL callers from draining active rollouts before releasing GPU memory.The client stops admitting waiting requests and drives running requests and queued model batches to completion.
LLMEnginekeeps using its output processor during the drain, so stop strings still abort generation promptly. Existing output collectors preserve cumulative, delta, and final-only results for subsequent delivery, including unfinished-request accounting. The existing pause/sleep path then synchronizes the device and performs cache cleanup and memory release. Waiting requests remain queued for wake-up.This implements the in-process
mode="wait"item in #48311, part of the Q3 RL roadmap #48314. The issue has no assignee or competing claim for this item. Searches of open PRs by issue number and in-process/inproc sleep/drain keywords found no overlapping implementation; #48337 concerns a separate LoRA level-2 allocation issue. In-process data parallelism remains explicitly unsupported because it requires coordinated draining across replicas.Test Plan
Test Result
Validated commit:
e862971f18f61035ae09d53e0ff8506b1f8f5ecd, based on52dd0d7562adb3c2d556ce6fe7ca7c3226e1976b.Commands:
.venv/bin/python -m pytest tests/v1/engine/test_llm_engine.py \ -k 'not test_skip_tokenizer_initialization' -x -vv .venv/bin/python /ch_data/wzp/oss-ai-infra/vllm-inproc-drain-20260914-evidence/validate_tp2.py .venv/bin/python -m pre_commit run --files \ vllm/v1/engine/core_client.py vllm/v1/engine/llm_engine.py \ tests/v1/engine/test_llm_engine.py docs/features/sleep_mode.md .venv/bin/python -m pre_commit run mypy-3.12 --hook-stage manual --files \ vllm/v1/engine/core_client.py vllm/v1/engine/llm_engine.pyGPU tests run on RTX 3090 under the shared host lock. The TP=2 harness runs the same regression tests with a real two-GPU multiprocessing executor. Validation uses this checkout's Python sources with PyTorch 2.13.0+cu130 and existing native extensions; it is not a full native rebuild of the base SHA. One pre-existing Llama-3.2-1B-Instruct test could not start because that model is absent from the offline cache; it is explicitly excluded from the available-model run. No multi-replica DP, PP, quantized, multimodal, or KV-connector coverage is claimed.
AI assistance: Codex assisted with implementation and validation.