Skip to content

[Core][Frontend] Report committed weight version on each output - #53199

Draft
aoshen02 wants to merge 3 commits into
vllm-project:mainfrom
aoshen02:codex/tito-weight-version-response
Draft

aoshen02 wants to merge 3 commits into
vllm-project:mainfrom
aoshen02:codex/tito-weight-version-response

Conversation

@aoshen02

@aoshen02 aoshen02 commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

Why

Pipeline RL can keep a generation request alive across an in-place weight update. The previous version of this PR bound weight_version at request admission, so every chunk stayed on the old label even after the serving version advanced. On H200, a 1,500-token request remained labeled 1 while /weight_info reported 2.

Design

  • Stamp the current committed EngineCore version on each output, then propagate it through RequestOutput and the Python and Rust /inference/v1/generate responses. There is no per-chunk control RPC.
  • Keep the EngineCore wire field optional and append-only for compatibility. Aggregated/non-streaming responses report the latest output's version.
  • weight_version is serving metadata at output formation, not a guarantee that every token in a long-lived request used one checkpoint. The change is limited to the token-in/token-out API.

This updates the existing #53199 rather than opening a duplicate PR.

Tests

  • H200, Qwen2.5-0.5B, Model Runner V2, same serve flags: after pause(mode=keep) → update_weight_version(2) → resume, the baseline generated 1,500 tokens all labeled 1, despite /weight_info=2; the patched Vime e975 backport returned 1 then 2 on the same request and generated at least 20 tokens after the update. This validates the behavior, not the PR's newer main-based binary.
  • Python token-in/token-out and output tests: 47 passed; scheduler version propagation test: passed. Two focused A/B regressions cover both abort paths: the frontend locally flushing buffered tokens and EngineCore emitting a synthetic abort output on pause(mode=abort). Each failed against Vime e975 before its respective fix and passed after it; the latter was exposed by the four-GPU PipelineRL flush-interval CI.
  • Rust generate tests: 14 passed; Rust LLM tests: 25 passed; Python/Rust MsgPack compatibility fixture: passed.
  • Changed-file pre-commit: passed, including mypy, Ruff and Rust formatting.
  • The four-GPU Vime PipelineRL E2E has not passed yet: its first candidate run could not start the rollout engines because other jobs reduced free GPU memory below the requested budget. It remains a merge gate; this is not a claimed accuracy or throughput result.

AI assisted with implementation and testing. Human submitter review is required before merge.

@mergify

mergify Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @aoshen02.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 16, 2026
@aoshen02
aoshen02 force-pushed the codex/tito-weight-version-response branch from f039ff8 to fc13993 Compare October 4, 2026 01:17
@mergify

mergify Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--53199.org.readthedocs.build/en/53199/

@mergify mergify Bot added the documentation Improvements or additions to documentation label Oct 4, 2026
@aoshen02 aoshen02 changed the title [Core][Frontend] Bind weight versions to generation requests [Core][Frontend] Report committed weight version on each output Oct 4, 2026
@mergify mergify Bot removed the needs-rebase label Oct 4, 2026
Stamp the current EngineCore version on each output and propagate it through Python and Rust token-in/token-out responses. Keep streaming requests observable across in-place weight updates.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02

aoshen02 commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Local validation of the existing Rust weight-version path (AI-assisted):

  • On pinned e975 with the Rust backport, actual H200 Qwen2.5-0.5B TITO streaming/non-streaming requests all expose the committed version on every output, after the native update-version RPC.
  • Four live HTTP cases pass (top-k32 and pure-top-p with the separate Vime capacity extension). This verifies metadata propagation, not changed model weights or an in-place-update performance benchmark.
  • The Rust regression already checks different versions on consecutive chunks; compiled frontend, nullable-response and serializer tests pass. The Vime getter uses the existing /weight_info route; no duplicate query API is needed.
  • No additional implementation change to this PR was necessary. The image-install gap was fixed by compiling the Rust backport instead of installing only Python changes.

Final human review remains required; no automatic merge or latest-image promotion.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation frontend rust scheduler

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant