Repository navigation
fix(musa): support float64 in-place log - #113
Merged
yeahdongcn merged 9 commits intoSep 18, 2026
Merged
Conversation
This was referenced Sep 14, 2026
yeahdongcn
reviewed
Sep 15, 2026
Add a CPU-runnable unit test for the MUSA float64 Tensor.log_() compatibility patch so the version gate and in-place fallback path are both covered without requiring a MUSA device.
Replace remaining 0.1.86 version strings, including the historical benchmark snapshot, so the package version is consistent across the tree.
yingzhou-bjtu
force-pushed
the
fix/musa-float64-log-inplace
branch
from
September 18, 2026 03:03
99062f6 to
dfc9271
Compare
The float64 log_ backport already requires torch_musa. Read torch.musa.__version__ instead of wrapping the known attributes in getattr.
Read torch.musa.__version__ directly for the float64 log_ gate, document that 2.11.0.post2 also owns the native log_ fix, and restore the dated benchmark snapshot to 0.1.86.
Sign the in-place log_ workaround comment and move the CPU coverage out of TestAcceleratorModuleWrapper into TestTensorLogPatch.
yeahdongcn
approved these changes
Sep 18, 2026
yeahdongcn
added a commit
that referenced
this pull request
Oct 10, 2026
The English and Chinese READMEs have not kept up with the May-October work. This documents what merged, moves the version-gated shims into one table, and refreshes the measured numbers that had gone stale. Feature table - CUDA memory-pool APIs, `torch.cuda.streams`, CUDA-graph executable rotation, `torch.cuda._get_device_index`, `get_memory_info()`, and the FlashAttention provider shims (#61, #98, #103, #106, #108, #115) - the "What Works" table goes back to one line per feature; the paragraph-sized `log_` / `isfinite` / `out_dtype` cells move into the new section below New "torch_musa Compatibility" section - one table of every version-gated shim with the release it is installed on: the four `< 2.11.0.post2` patches (#106, #113, #124), the `< 2.13.0` `mm`/`bmm` `out_dtype=` backport (#116), the stable-ABI header backport (#86, #96), asynchronous `isfinite` (#120), and `torch.cuda.streams` (#98) New "Environment Variables" section - the graph-rotation knobs (#72), `TORCHADA_PLATFORM`, the C++ operator-override switches (#61, #128), and the two variables that were already documented Corrected and extended details - torch.compile: FX `device` builtin (#124), Dynamo's device-index helper (#108), `MUSA_VISIBLE_DEVICES` mirroring (#106) - C++ extensions: nested `<torch/cuda.h>` porting (#95), stable-ABI `STABLE_TORCH_LIBRARY_IMPL` rekeying and stream helpers (#100), torch 2.6+ `include_paths`/`library_paths` signatures (#121), stale JIT build locks (#128) - MoE tables are generated from checked-in recipes (#115) - unsupported CUDA runtime APIs as no-ops (#65) - Performance: replace the 0.1.94 / torch_musa 2.7.1 numbers with the checked-in 0.1.95 / 2.11.0.post2 entry, and stop claiming every fast path is under 200ns
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why this change
On the validated MUSA runtime,
Tensor.log_()still rejectsfloat64in-place execution, even though the equivalent out-of-place
torch.log()operation works. This matters for CUDA-oriented code that uses
log_()and isotherwise expected to run unchanged on MUSA.
What changed
This patch adds a small compatibility handler to torchada's existing patch
registry. For MUSA
float64tensors ontorch_musa < 2.11.0.post2, it computesthe result with the supported out-of-place operation and copies it back. That
keeps the in-place contract while leaving the native path untouched on newer
torch_musa releases.
The patch is deliberately limited to this case. CPU, CUDA, other dtypes, and
the public API are unchanged. The version check reuses torchada's existing
semantic-version helper, and the version-gated behavior is documented in both
README files.
The unit tests now cover both the version gate and the in-place fallback path
in
tests/test_cuda_patching.pyunderTestTensorLogPatch. Those tests areCPU-runnable.
Validation
CPU unit tests for
_patch_tensor_log_():PYTHONPATH=src python3 -m pytest -q tests/test_cuda_patching.py \ -k 'float64_log or torch_musa_version_boundary' --tb=shortResult:
15 passed, 226 deselected in 0.21s.That subset covers:
test_torch_musa_version_boundarytest_float64_log_patch_respects_torch_musa_versiontest_float64_log_patch_uses_out_of_place_copy_on_musaThe focused float64 log tests alone:
Result:
2 passed, 239 deselected in 0.04s.py_compileandgit diff --checkpass for the changed source.MUSA tests on
torch_musa 2.11.0.post1+musa5.2.0:Result:
3 passed in 3.22s.Compatibility
The change is MUSA-specific and keeps the existing CUDA path intact. It is
enabled only for torch_musa versions below
2.11.0.post2; newer versions usetheir native implementation.
Related SGLang-Omni integration:
sgl-project/sglang-omni#2167
Related shared Graph execution-mode change:
sgl-project/sglang-omni#2168