Skip to content

[BugFix][Worker] Map CUDA stream capturing to NPU in MRv2 - #149

Merged
cursor[bot] merged 1 commit into
enable-mrv2-whitelist-deaefrom
cursor/v2-cuda-capture-npu-9095
Sep 15, 2026
Merged

cursor[bot] merged 1 commit into
enable-mrv2-whitelist-deaefrom
cursor/v2-cuda-capture-npu-9095

Conversation

@yjyang62

Copy link
Copy Markdown
Owner

What this PR does / why we need it?

V1 NPU model runner remaps torch.cuda.is_current_stream_capturing to torch.npu.is_current_stream_capturing. V2 torch_cuda_wrapper mapped the other CUDA APIs but missed this one.

When Qwen3 defaults to MRv2 (vllm-project#16203), GPU V2 load_model runs PrefetchOffloader.start_onload_to_static, which calls the CUDA dummy and crashes:

RuntimeError: Tried to instantiate dummy base class _cuda_isCurrentStreamCapturing

This is the a2-1 test_cpu_weight_offload.py failure on PR 16203. Mapping is applied at V2 runner init and is not restored, so later load_model / prefetch can query capture state on NPU.

Does this PR introduce any user-facing change?

No. Prefetch / CPU weight offload on MRv2 should start instead of crashing.

How was this patch tested?

  • Unit tests added/updated: tests/ut/worker/v2/test_v2_utils.py
  • Manual testing: NPU e2e (test_cpu_weight_offload.py) needs CI / A2 hardware
  • CI testing: stacked on enable-mrv2-whitelist-deae (PR 16203) so a2-1 prefetch offload can re-run there
Open in Web Open in Cursor 

GPU V2 PrefetchOffloader calls torch.cuda.is_current_stream_capturing
during load_model. V1 already remaps that CUDA dummy to torch.npu;
V2 torch_cuda_wrapper did not, so Qwen3 default-V2 prefetch e2e
crashed on NPU.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
@cursor
cursor Bot merged commit d74ffee into enable-mrv2-whitelist-deae Sep 15, 2026
0 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants