Repository navigation
fix(cpp_extension): accept torch 2.6+ include_paths/library_paths call signatures - #121
Merged
Merged
Conversation
This was referenced Oct 7, 2026
[Perf] MAGI-2: run the MUSA mHC Sinkhorn iterations in one Triton kernel
vllm-project/vllm-omni#8594
Open
…l signatures torch 2.10 added torch_include_dirs to cpp_extension.include_paths and torch_include_dirs/cross_target_platform to library_paths, and its Inductor C++ builder calls include_paths(device_type, torch_include_dirs) positionally and library_paths(device_type, torch_include_dirs=..., cross_target_platform=...). torchada's replacements took (cuda, device_type), so the positional call handed device_type=False to include_paths and failed with "'bool' object has no attribute 'lower'", and library_paths rejected the keywords. Every CPU kernel Inductor compiles with a cold cache on MUSA hit this. On torch 2.6-2.9 the builder's positional device_type string landed in cuda instead, so CPU builds also got the MUSA include and library paths. Both helpers now take PyTorch's (device_type, torch_include_dirs[, cross_target_platform]) signature, keep cuda= as a keyword and treat a positional bool as the PyTorch < 2.6 cuda argument. torch_include_dirs=False leaves out PyTorch's own include and lib directories; off MUSA the arguments are passed through when the installed torch accepts them. Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
yeahdongcn
force-pushed
the
xd/cpp-extension-path-signature
branch
from
October 8, 2026 09:16
1caedf1 to
0f3cc5d
Compare
yeahdongcn
marked this pull request as ready for review
October 8, 2026 09:16
This was referenced Oct 9, 2026
yeahdongcn
added a commit
that referenced
this pull request
Oct 10, 2026
The English and Chinese READMEs have not kept up with the May-October work. This documents what merged, moves the version-gated shims into one table, and refreshes the measured numbers that had gone stale. Feature table - CUDA memory-pool APIs, `torch.cuda.streams`, CUDA-graph executable rotation, `torch.cuda._get_device_index`, `get_memory_info()`, and the FlashAttention provider shims (#61, #98, #103, #106, #108, #115) - the "What Works" table goes back to one line per feature; the paragraph-sized `log_` / `isfinite` / `out_dtype` cells move into the new section below New "torch_musa Compatibility" section - one table of every version-gated shim with the release it is installed on: the four `< 2.11.0.post2` patches (#106, #113, #124), the `< 2.13.0` `mm`/`bmm` `out_dtype=` backport (#116), the stable-ABI header backport (#86, #96), asynchronous `isfinite` (#120), and `torch.cuda.streams` (#98) New "Environment Variables" section - the graph-rotation knobs (#72), `TORCHADA_PLATFORM`, the C++ operator-override switches (#61, #128), and the two variables that were already documented Corrected and extended details - torch.compile: FX `device` builtin (#124), Dynamo's device-index helper (#108), `MUSA_VISIBLE_DEVICES` mirroring (#106) - C++ extensions: nested `<torch/cuda.h>` porting (#95), stable-ABI `STABLE_TORCH_LIBRARY_IMPL` rekeying and stream helpers (#100), torch 2.6+ `include_paths`/`library_paths` signatures (#121), stale JIT build locks (#128) - MoE tables are generated from checked-in recipes (#115) - unsupported CUDA runtime APIs as no-ops (#65) - Performance: replace the 0.1.94 / torch_musa 2.7.1 numbers with the checked-in 0.1.95 / 2.11.0.post2 entry, and stop claiming every fast path is under 200ns
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What / Why
PyTorch 2.10 added
torch_include_dirstotorch.utils.cpp_extension.include_paths, andtorch_include_dirsandcross_target_platformtolibrary_paths. Its Inductor C++ builder callsinclude_paths(device_type, torch_include_dirs)positionally andlibrary_paths(device_type, torch_include_dirs=..., cross_target_platform=...).On MUSA, torchada replaces both helpers with versions that take
(cuda, device_type). The positional call therefore handsdevice_type=Falsetoinclude_paths, which fails with'bool' object has no attribute 'lower', andlibrary_pathsrejects the two keywords. Every CPU kernel that Inductor compiles with a cold cache goes through these calls, so afterimport torchadaa plain CPUtorch.compilewith an emptyTORCHINDUCTOR_CACHE_DIRfails.vLLM-Omni's MAGI-2 test suites on MUSA hit this: their CPU
torch.compiletests run with a fresh Inductor cache and failed in this call. With this change installed they no longer fail there.Change
include_paths(device_type=None, torch_include_dirs=True, *, cuda=None)andlibrary_paths(device_type=None, torch_include_dirs=True, cross_target_platform=None, *, cuda=None)follow PyTorch's signature, positionally and by keyword.device_typeposition is the PyTorch < 2.6 positionalcudaargument, which torch 2.5's builder passes.cuda=stays as a keyword.torch_include_dirs=Falseleaves out paths under PyTorch's ownincludeandlibdirectories. torch_musa's directories and torchada's stable-ABI compat headers stay.torch_include_dirsandcross_target_platformare passed through when the installed torch accepts them.device_type, positionally. With the old signature that string landed incudaand counted as true, so CPU builds also received the MUSA include and library paths from that device-options call. That call now returns the CPU set (no MUSA library directories).cudais now keyword-only, so torchada's own two-positional forminclude_paths(cuda, device_type)is no longer accepted. Nothing in this repository used it.Verification
New tests in
tests/test_cuda_patching.py::TestCppExtensionPaths(MUSA only):test_paths_accept_the_positional_torch_signature: the positional torch 2.10+ calls return the same paths as the keyword calls, and a positional bool matchescuda=;test_paths_without_torch_dirs: withtorch_include_dirs=Falseno returned path lies under the torch package, and the result is a subset of the default;test_inductor_cpu_compile_with_a_cold_cache: a CPUtorch.compile(fullgraph=True)in a fresh interpreter with an emptyTORCHINDUCTOR_CACHE_DIRbuilds and returns the expected values.Results on MTT S5000 (torch 2.11.0.post2, torch_musa 2.11.0.post2+musa5.2.0):
torch.compilefails with stock torchada and passes with this change.pytest tests/, with this change applied on top of fix(musa): keep torch.isfinite asynchronous on MUSA floating tensors #120 (the torchada tree used for the vLLM-Omni runs): 634 passed, 19 skipped, 2 failed. The new tests pass. The two failures areTestInductorTemplateHeuristics::test_copies_only_cuda_triton_heuristics_and_clears_cacheand::test_is_idempotent_and_preserves_cache_without_changes(KeyError: ('triton::bmm', 'musa', None)and the same fortriton::mm), and they fail the same way on the unmodified source in the same image.Not covered / notes
library_paths(..., torch_include_dirs=False)is only reached by libtorch-free AOTInductor builds (aot_inductor.link_libtorch=False), which were not exercised on MUSA beyond the unit test.library_pathsstill returns[]for a CPU build, as before. Inductor addsTORCH_LIB_PATHitself.AI assistance: Claude Code drafted the change, the tests and this description and ran the MUSA validation listed above.