[BugFix][Worker] Map CUDA stream capturing to NPU in MRv2 - #149
Merged
cursor[bot] merged 1 commit intoSep 15, 2026
Merged
Conversation
GPU V2 PrefetchOffloader calls torch.cuda.is_current_stream_capturing during load_model. V1 already remaps that CUDA dummy to torch.npu; V2 torch_cuda_wrapper did not, so Qwen3 default-V2 prefetch e2e crashed on NPU. Signed-off-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it?
V1 NPU model runner remaps
torch.cuda.is_current_stream_capturingtotorch.npu.is_current_stream_capturing. V2torch_cuda_wrappermapped the other CUDA APIs but missed this one.When Qwen3 defaults to MRv2 (vllm-project#16203), GPU V2
load_modelrunsPrefetchOffloader.start_onload_to_static, which calls the CUDA dummy and crashes:This is the a2-1
test_cpu_weight_offload.pyfailure on PR 16203. Mapping is applied at V2 runner init and is not restored, so laterload_model/ prefetch can query capture state on NPU.Does this PR introduce any user-facing change?
No. Prefetch / CPU weight offload on MRv2 should start instead of crashing.
How was this patch tested?
tests/ut/worker/v2/test_v2_utils.pytest_cpu_weight_offload.py) needs CI / A2 hardwareenable-mrv2-whitelist-deae(PR 16203) so a2-1 prefetch offload can re-run there