Skip to content

[Feature][Core] Support draining in-process requests before sleep - #56754

Open
zupengwang wants to merge 1 commit into
vllm-project:mainfrom
zupengwang:codex/inproc-sleep-drain
Open

zupengwang wants to merge 1 commit into
vllm-project:mainfrom
zupengwang:codex/inproc-sleep-drain

Conversation

@zupengwang

Copy link
Copy Markdown
Contributor

Purpose

Support sleep(mode="wait") with the in-process engine (VLLM_ENABLE_V1_MULTIPROCESSING=0). This currently raises a ValueError, preventing offline RL callers from draining active rollouts before releasing GPU memory.

The client stops admitting waiting requests and drives running requests and queued model batches to completion. LLMEngine keeps using its output processor during the drain, so stop strings still abort generation promptly. Existing output collectors preserve cumulative, delta, and final-only results for subsequent delivery, including unfinished-request accounting. The existing pause/sleep path then synchronizes the device and performs cache cleanup and memory release. Waiting requests remain queued for wake-up.

This implements the in-process mode="wait" item in #48311, part of the Q3 RL roadmap #48314. The issue has no assignee or competing claim for this item. Searches of open PRs by issue number and in-process/inproc sleep/drain keywords found no overlapping implementation; #48337 concerns a separate LoRA level-2 allocation issue. In-process data parallelism remains explicitly unsupported because it requires coordinated draining across replicas.

Test Plan

  • Reproduce the unsupported-mode error with a real OPT-125M model before changing production code.
  • Compare uninterrupted generation against drain/sleep/wake for token IDs, text, finish reasons, and per-token logprobs.
  • Cover all three output kinds and sleep levels 0/1/2; use asynchronous scheduling and full CUDA graphs at level 1, and reload weights after level 2.
  • Verify waiting requests are not admitted during the drain, buffered results are delivered once, partial wake does not resume scheduling, and a subsequent request completes after full wake.
  • Exercise raw client output retention, pending-batch cleanup, unsafe-state rejection, idle draining, and failure propagation before sleep.

Test Result

Validated commit: e862971f18f61035ae09d53e0ff8506b1f8f5ecd, based on 52dd0d7562adb3c2d556ce6fe7ca7c3226e1976b.

  • Baseline: the real-model regression fails with the unsupported in-process wait-mode error.
  • Available adjacent engine suite: 19 passed, 1 deselected.
  • Real TP=2 lifecycle matrix: 9 passed.
  • CPU-only drain regressions: 5 passed.
  • Applicable pre-commit hooks, mypy 3.10, and mypy 3.12: passed.

Commands:

.venv/bin/python -m pytest tests/v1/engine/test_llm_engine.py \
  -k 'not test_skip_tokenizer_initialization' -x -vv
.venv/bin/python /ch_data/wzp/oss-ai-infra/vllm-inproc-drain-20260914-evidence/validate_tp2.py
.venv/bin/python -m pre_commit run --files \
  vllm/v1/engine/core_client.py vllm/v1/engine/llm_engine.py \
  tests/v1/engine/test_llm_engine.py docs/features/sleep_mode.md
.venv/bin/python -m pre_commit run mypy-3.12 --hook-stage manual --files \
  vllm/v1/engine/core_client.py vllm/v1/engine/llm_engine.py

GPU tests run on RTX 3090 under the shared host lock. The TP=2 harness runs the same regression tests with a real two-GPU multiprocessing executor. Validation uses this checkout's Python sources with PyTorch 2.13.0+cu130 and existing native extensions; it is not a full native rebuild of the base SHA. One pre-existing Llama-3.2-1B-Instruct test could not start because that model is absent from the offline cache; it is explicitly excluded from the available-model run. No multi-replica DP, PP, quantized, multimodal, or KV-connector coverage is claimed.

AI assistance: Codex assisted with implementation and validation.

Drive running requests and pending model batches to completion before
sleeping, preserving waiting requests and buffered generation results.
Keep stop-string processing active and reuse per-request output collectors.
Cover sleep levels, output kinds, partial wake, and drain failures.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Wang Zupeng <zupenwang@gmail.com>
@mergify

mergify Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--56754.org.readthedocs.build/en/56754/

@mergify mergify Bot added the documentation Improvements or additions to documentation label Sep 14, 2026
@zupengwang
zupengwang marked this pull request as ready for review September 14, 2026 04:16
@zupengwang
zupengwang requested a review from njhill as a code owner September 14, 2026 04:16

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify

mergify Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @wangzupeng12061.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation needs-rebase

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant