chore: Upgrade vLLM from 0.17.1 to 0.20.0 - #2384
Merged
Merged
Conversation
Contributor
Author
|
/ok to test e2a1776 |
Contributor
Author
|
/ok to test e52995e |
kajalj22
force-pushed
the
kajalj/upgrade-vllm-2.11
branch
from
May 6, 2026 20:00
e52995e to
f7ae9ae
Compare
Contributor
Author
|
/ok to test 648be36 |
Contributor
Author
|
/ok to test 21c0e8e |
kajalj22
force-pushed
the
kajalj/upgrade-vllm-2.11
branch
from
May 13, 2026 15:57
21c0e8e to
711a112
Compare
Contributor
Author
|
/ok to test 711a112 |
kajalj22
force-pushed
the
kajalj/upgrade-vllm-2.11
branch
from
May 14, 2026 01:49
711a112 to
a5f29ce
Compare
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
kajalj22
force-pushed
the
kajalj/upgrade-vllm-2.11
branch
from
May 14, 2026 21:21
a5f29ce to
ba3ab1e
Compare
Contributor
Author
|
/ok to test ba3ab1e |
3 tasks
Contributor
Author
|
/ok to test 197eb71 |
Contributor
Author
|
/ok to test 97289f5 |
Contributor
Author
|
/ok to test f19fc23 |
Contributor
Author
|
/ok to test 132152f |
Contributor
Author
|
/ok to test fd73aa2 |
Contributor
Author
|
/ok to test 1af50ac |
- vLLM 0.17.1 → 0.20.0, torch 2.10 → 2.11, torchvision 0.25 → 0.26, flashinfer 0.6.4 → 0.6.8.post1 - Cap requires-python to <3.14 - Adapt to vLLM 0.20 render architecture: move prefix-token override to NeMoRLOpenAIServingRender - Wrap process_weights_after_loading in set_current_vllm_config context - Fix cloudpickle ConfigModuleInstance error (torch 2.11) - Fix Eagle3 draft weight loading: trim padded vocab embeddings Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Kajal Jain <kajalj@nvidia.com>
vLLM 0.20 moved chat preprocessing to the render layer, but create_chat_completion no longer catches errors from that path. Prompts exceeding max_model_len now raise VLLMValidationError as an unhandled exception (500) instead of returning ErrorResponse (400). The Gym proxy only detects context-length overflow on 400, so the 500 crashes rollouts. Catch VLLMValidationError at the endpoint and return HTTP 400 to restore the graceful handling chain. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Kajal Jain <kajalj@nvidia.com>
sglang is incompatible with the vLLM 0.20 / torch 2.11 upgrade. Unconditionally set SKIP_SGLANG_BUILD=1 so the Docker build and tests skip sglang until it is updated separately. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Kajal Jain <kajalj@nvidia.com>
Eagle3 draft models can have draft_vocab_size (32k) different from vocab_size (151k). The old code used a single org_vocab_size from the first VocabParallelEmbedding (embed_tokens), which meant lm_head weights were never trimmed (padded_32k < 151k → condition false), causing an assertion failure in vLLM's weight_loader. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Kajal Jain <kajalj@nvidia.com>
|
Auto-sync is disabled for ready for review pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
Author
|
/ok to test 92e00b7 |
vLLM 0.20 added quant_config arg to ParallelLMHead in Eagle3LlamaForCausalLM.__init__, so the old_snippet no longer matched and the has_own_lm_head patch silently failed to apply. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Kajal Jain <kajalj@nvidia.com>
Contributor
Author
|
/ok to test a0f1d4b |
chtruong814
approved these changes
May 22, 2026
bxyu-nvidia
approved these changes
May 22, 2026
ananthsub
reviewed
May 22, 2026
ananthsub
approved these changes
May 22, 2026
Contributor
Author
|
/ok to test 73e9b54 |
This was referenced May 29, 2026
Closed
4 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Upgrades vLLM from 0.17.1 to 0.20.0, along with dependent version bumps and API adaptation changes.
Dependency changes (
pyproject.toml)<3.14vLLM 0.20 API adaptation (
vllm_worker_async.py)OpenAIServing._preprocess_chattoOpenAIServingRender.preprocess_chatNeMoRLOpenAIServingRender(NeMoRLOpenAIServingMixin, OpenAIServingRender)to carry ourrequired_prefix_token_idsoverride into the new render architecturereasoning_parser,skip_mm_cache,enable_auto_tools,openai_serving_renderWeight loading fix (
vllm_backend.py)process_weights_after_loadingcalls inset_current_vllm_configcontext manager (required by vLLM 0.20)Pickle fix (
megatron_policy_worker.py)self.model.config→getattr(self.model, "config", None)to prevent cloudpickle from capturingtorch.distributed.config(a non-pickleableConfigModuleInstancein torch 2.11)Test plan
🤖 Generated with Claude Code