Skip to content

Fix nested torch CUDA header porting - #95

Merged
yeahdongcn merged 2 commits into
mainfrom
fix/torch-cuda-header-porting
Jul 21, 2026
Merged

yeahdongcn merged 2 commits into
mainfrom
fix/torch-cuda-header-porting

Conversation

@yeahdongcn

Copy link
Copy Markdown
Collaborator

Summary

  • keep the broad cuda.h replacement narrowed to exact include directives
  • translate both angle-bracket and quoted torch/cuda.h includes to torch/musa.h
  • add regression coverage that preserves project-local names such as project/decode_jpegs_cuda.h

Motivation

vLLM csrc/cuda_view.cu includes <torch/cuda.h>. The narrowed header rules in torchada 0.1.75 correctly avoided rewriting project-local filenames, but they only restored exact mappings for top-level <cuda.h> and missed this nested PyTorch header.

Validation

  • PYTHONPATH=src python -m pytest tests/test_cpp_extension.py::TestCUDAExtension tests/test_inplace_porting.py -q (19 passed)
  • local torch_musa 2.7.1 BuildExtension._port_directory() regression using the vLLM cuda_view.cu include pattern: torch/cuda.h became torch/musa.h, cuda_runtime.h became musa_runtime.h, and decode_jpegs_cuda.h remained unchanged
  • git diff --check, isort, and Ruff passed for the changed files

Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
@yeahdongcn
yeahdongcn merged commit 86d842c into main Jul 21, 2026
@yeahdongcn
yeahdongcn deleted the fix/torch-cuda-header-porting branch August 13, 2026 12:16
yeahdongcn added a commit that referenced this pull request Oct 10, 2026
The English and Chinese READMEs have not kept up with the May-October work.
This documents what merged, moves the version-gated shims into one table, and
refreshes the measured numbers that had gone stale.

Feature table
- CUDA memory-pool APIs, `torch.cuda.streams`, CUDA-graph executable rotation,
  `torch.cuda._get_device_index`, `get_memory_info()`, and the FlashAttention
  provider shims (#61, #98, #103, #106, #108, #115)
- the "What Works" table goes back to one line per feature; the paragraph-sized
  `log_` / `isfinite` / `out_dtype` cells move into the new section below

New "torch_musa Compatibility" section
- one table of every version-gated shim with the release it is installed on:
  the four `< 2.11.0.post2` patches (#106, #113, #124), the `< 2.13.0`
  `mm`/`bmm` `out_dtype=` backport (#116), the stable-ABI header backport (#86,
  #96), asynchronous `isfinite` (#120), and `torch.cuda.streams` (#98)

New "Environment Variables" section
- the graph-rotation knobs (#72), `TORCHADA_PLATFORM`, the C++ operator-override
  switches (#61, #128), and the two variables that were already documented

Corrected and extended details
- torch.compile: FX `device` builtin (#124), Dynamo's device-index helper (#108),
  `MUSA_VISIBLE_DEVICES` mirroring (#106)
- C++ extensions: nested `<torch/cuda.h>` porting (#95), stable-ABI
  `STABLE_TORCH_LIBRARY_IMPL` rekeying and stream helpers (#100), torch 2.6+
  `include_paths`/`library_paths` signatures (#121), stale JIT build locks (#128)
- MoE tables are generated from checked-in recipes (#115)
- unsupported CUDA runtime APIs as no-ops (#65)
- Performance: replace the 0.1.94 / torch_musa 2.7.1 numbers with the checked-in
  0.1.95 / 2.11.0.post2 entry, and stop claiming every fast path is under 200ns
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant