Skip to content

fix(cpp_extension): accept torch 2.6+ include_paths/library_paths call signatures - #121

Merged
yeahdongcn merged 2 commits into
mainfrom
xd/cpp-extension-path-signature
Oct 8, 2026
Merged

yeahdongcn merged 2 commits into
mainfrom
xd/cpp-extension-path-signature

Conversation

@yeahdongcn

Copy link
Copy Markdown
Collaborator

What / Why

PyTorch 2.10 added torch_include_dirs to torch.utils.cpp_extension.include_paths, and torch_include_dirs and cross_target_platform to library_paths. Its Inductor C++ builder calls include_paths(device_type, torch_include_dirs) positionally and library_paths(device_type, torch_include_dirs=..., cross_target_platform=...).

On MUSA, torchada replaces both helpers with versions that take (cuda, device_type). The positional call therefore hands device_type=False to include_paths, which fails with 'bool' object has no attribute 'lower', and library_paths rejects the two keywords. Every CPU kernel that Inductor compiles with a cold cache goes through these calls, so after import torchada a plain CPU torch.compile with an empty TORCHINDUCTOR_CACHE_DIR fails.

vLLM-Omni's MAGI-2 test suites on MUSA hit this: their CPU torch.compile tests run with a fresh Inductor cache and failed in this call. With this change installed they no longer fail there.

Change

  • include_paths(device_type=None, torch_include_dirs=True, *, cuda=None) and library_paths(device_type=None, torch_include_dirs=True, cross_target_platform=None, *, cuda=None) follow PyTorch's signature, positionally and by keyword.
  • A bool in the device_type position is the PyTorch < 2.6 positional cuda argument, which torch 2.5's builder passes. cuda= stays as a keyword.
  • On MUSA, torch_include_dirs=False leaves out paths under PyTorch's own include and lib directories. torch_musa's directories and torchada's stable-ABI compat headers stay.
  • Off MUSA, torch_include_dirs and cross_target_platform are passed through when the installed torch accepts them.
  • On torch 2.6-2.9 the builder passes only device_type, positionally. With the old signature that string landed in cuda and counted as true, so CPU builds also received the MUSA include and library paths from that device-options call. That call now returns the CPU set (no MUSA library directories).
  • cuda is now keyword-only, so torchada's own two-positional form include_paths(cuda, device_type) is no longer accepted. Nothing in this repository used it.

Verification

New tests in tests/test_cuda_patching.py::TestCppExtensionPaths (MUSA only):

  • test_paths_accept_the_positional_torch_signature: the positional torch 2.10+ calls return the same paths as the keyword calls, and a positional bool matches cuda=;
  • test_paths_without_torch_dirs: with torch_include_dirs=False no returned path lies under the torch package, and the result is a subset of the default;
  • test_inductor_cpu_compile_with_a_cold_cache: a CPU torch.compile(fullgraph=True) in a fresh interpreter with an empty TORCHINDUCTOR_CACHE_DIR builds and returns the expected values.

Results on MTT S5000 (torch 2.11.0.post2, torch_musa 2.11.0.post2+musa5.2.0):

  • A cold-cache CPU torch.compile fails with stock torchada and passes with this change.
  • pytest tests/, with this change applied on top of fix(musa): keep torch.isfinite asynchronous on MUSA floating tensors #120 (the torchada tree used for the vLLM-Omni runs): 634 passed, 19 skipped, 2 failed. The new tests pass. The two failures are TestInductorTemplateHeuristics::test_copies_only_cuda_triton_heuristics_and_clears_cache and ::test_is_idempotent_and_preserves_cache_without_changes (KeyError: ('triton::bmm', 'musa', None) and the same for triton::mm), and they fail the same way on the unmodified source in the same image.

Not covered / notes

AI assistance: Claude Code drafted the change, the tests and this description and ran the MUSA validation listed above.

…l signatures

torch 2.10 added torch_include_dirs to cpp_extension.include_paths and
torch_include_dirs/cross_target_platform to library_paths, and its Inductor
C++ builder calls include_paths(device_type, torch_include_dirs) positionally
and library_paths(device_type, torch_include_dirs=...,
cross_target_platform=...). torchada's replacements took (cuda, device_type),
so the positional call handed device_type=False to include_paths and failed
with "'bool' object has no attribute 'lower'", and library_paths rejected the
keywords. Every CPU kernel Inductor compiles with a cold cache on MUSA hit
this. On torch 2.6-2.9 the builder's positional device_type string landed in
cuda instead, so CPU builds also got the MUSA include and library paths.

Both helpers now take PyTorch's (device_type, torch_include_dirs[,
cross_target_platform]) signature, keep cuda= as a keyword and treat a
positional bool as the PyTorch < 2.6 cuda argument. torch_include_dirs=False
leaves out PyTorch's own include and lib directories; off MUSA the arguments
are passed through when the installed torch accepts them.

Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
@yeahdongcn
yeahdongcn force-pushed the xd/cpp-extension-path-signature branch from 1caedf1 to 0f3cc5d Compare October 8, 2026 09:16
@yeahdongcn
yeahdongcn marked this pull request as ready for review October 8, 2026 09:16
@yeahdongcn
yeahdongcn merged commit 387e503 into main Oct 8, 2026
yeahdongcn added a commit that referenced this pull request Oct 10, 2026
The English and Chinese READMEs have not kept up with the May-October work.
This documents what merged, moves the version-gated shims into one table, and
refreshes the measured numbers that had gone stale.

Feature table
- CUDA memory-pool APIs, `torch.cuda.streams`, CUDA-graph executable rotation,
  `torch.cuda._get_device_index`, `get_memory_info()`, and the FlashAttention
  provider shims (#61, #98, #103, #106, #108, #115)
- the "What Works" table goes back to one line per feature; the paragraph-sized
  `log_` / `isfinite` / `out_dtype` cells move into the new section below

New "torch_musa Compatibility" section
- one table of every version-gated shim with the release it is installed on:
  the four `< 2.11.0.post2` patches (#106, #113, #124), the `< 2.13.0`
  `mm`/`bmm` `out_dtype=` backport (#116), the stable-ABI header backport (#86,
  #96), asynchronous `isfinite` (#120), and `torch.cuda.streams` (#98)

New "Environment Variables" section
- the graph-rotation knobs (#72), `TORCHADA_PLATFORM`, the C++ operator-override
  switches (#61, #128), and the two variables that were already documented

Corrected and extended details
- torch.compile: FX `device` builtin (#124), Dynamo's device-index helper (#108),
  `MUSA_VISIBLE_DEVICES` mirroring (#106)
- C++ extensions: nested `<torch/cuda.h>` porting (#95), stable-ABI
  `STABLE_TORCH_LIBRARY_IMPL` rekeying and stream helpers (#100), torch 2.6+
  `include_paths`/`library_paths` signatures (#121), stale JIT build locks (#128)
- MoE tables are generated from checked-in recipes (#115)
- unsupported CUDA runtime APIs as no-ops (#65)
- Performance: replace the 0.1.94 / torch_musa 2.7.1 numbers with the checked-in
  0.1.95 / 2.11.0.post2 entry, and stop claiming every fast path is under 200ns
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant