Skip to content

[ModelRunner V2] Don't pin reused flashinfer tensors - #32799

Merged
WoosukKwon merged 1 commit into
vllm-project:mainfrom
njhill:mrv2-flashinfer
Jan 21, 2026
Merged

WoosukKwon merged 1 commit into
vllm-project:mainfrom
njhill:mrv2-flashinfer

Conversation

@njhill

@njhill njhill commented Jan 21, 2026

Copy link
Copy Markdown
Member

Since we do not have explicit synchronization in ModelRunnerV2, we do not pin reused CPU buffers to avoid a race condition between step N async copies to GPU and step N+1 buffer updates.

Signed-off-by: Nick Hill <nickhill123@gmail.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request correctly addresses a potential race condition in ModelRunnerV2 by conditionally disabling pin_memory for FlashInfer tensors. The added comment clearly explains the rationale behind this change, which is crucial for preventing asynchronous copy issues between CPU and GPU buffers. The implementation aligns with the stated objective of avoiding pinning when ModelRunnerV2 is active.

Comment on lines +609 to +610
self.pin_memory = (
not envs.VLLM_USE_V2_MODEL_RUNNER and is_pin_memory_available()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The decision to disable pin_memory is critical for ModelRunnerV2 due to its lack of explicit synchronization, as noted in the comment. However, relying on the global envs.VLLM_USE_V2_MODEL_RUNNER environment variable introduces a tight coupling between FlashInferMetadataBuilder and the global state. This approach can make the system harder to reason about, test, and debug, especially if the activation of ModelRunnerV2 is not perfectly synchronized with the environment variable's state. For improved modularity and explicitness, consider passing a use_v2_model_runner boolean parameter directly to the FlashInferMetadataBuilder constructor. This would ensure that the pin_memory behavior is directly controlled by the active ModelRunner instance rather than an implicit global flag.

@WoosukKwon WoosukKwon left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@github-project-automation github-project-automation Bot moved this to Ready in NVIDIA Jan 21, 2026
@WoosukKwon
WoosukKwon merged commit 24dc30f into vllm-project:main Jan 21, 2026
10 of 11 checks passed
@github-project-automation github-project-automation Bot moved this from Ready to Done in NVIDIA Jan 21, 2026
monajafi-amd pushed a commit to monajafi-amd/vllm that referenced this pull request Jan 23, 2026
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: mohammad najafi <mohammad.najafi@amd.com>
cwazai pushed a commit to cwazai/vllm that referenced this pull request Jan 25, 2026
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: 陈建华 <1647430658@qq.com>
lapy pushed a commit to lapy/vllm that referenced this pull request Jan 27, 2026
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
mystous pushed a commit to mystous/vllm_hybrid that referenced this pull request May 10, 2026
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
my-other-github-account pushed a commit to my-other-github-account/vllm that referenced this pull request May 15, 2026
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
0826joyce pushed a commit to 0826joyce/vllm-serving-optimization that referenced this pull request May 19, 2026
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
plasticchris pushed a commit to plasticchris/vllm that referenced this pull request Jul 20, 2026
)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants