Repository navigation
Conversation
1 task done
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
aoshen02
force-pushed
the
codex/tito-weight-version-response
branch
from
October 4, 2026 01:17
f039ff8 to
fc13993
Compare
Contributor
|
Documentation preview: https://vllm--53199.org.readthedocs.build/en/53199/ |
Stamp the current EngineCore version on each output and propagate it through Python and Rust token-in/token-out responses. Keep streaming requests observable across in-place weight updates. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: aoshen02 <aoshen@inferact.ai>
aoshen02
force-pushed
the
codex/tito-weight-version-response
branch
from
October 4, 2026 01:26
fc13993 to
3b81d10
Compare
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Contributor
Author
|
Local validation of the existing Rust weight-version path (AI-assisted):
Final human review remains required; no automatic merge or latest-image promotion. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Pipeline RL can keep a generation request alive across an in-place weight update. The previous version of this PR bound
weight_versionat request admission, so every chunk stayed on the old label even after the serving version advanced. On H200, a 1,500-token request remained labeled1while/weight_inforeported2.Design
RequestOutputand the Python and Rust/inference/v1/generateresponses. There is no per-chunk control RPC.weight_versionis serving metadata at output formation, not a guarantee that every token in a long-lived request used one checkpoint. The change is limited to the token-in/token-out API.This updates the existing #53199 rather than opening a duplicate PR.
Tests
pause(mode=keep)→update_weight_version(2)→resume, the baseline generated 1,500 tokens all labeled1, despite/weight_info=2; the patched Vime e975 backport returned1then2on the same request and generated at least 20 tokens after the update. This validates the behavior, not the PR's newer main-based binary.pause(mode=abort). Each failed against Vime e975 before its respective fix and passed after it; the latter was exposed by the four-GPU PipelineRL flush-interval CI.AI assisted with implementation and testing. Human submitter review is required before merge.