[AtomicOp](fix) Fuse expanded floating-point atomic min/max - #1460
Merged
Merged
Conversation
Contributor
|
🔍 OpenCodeReview found 2 issue(s) in this PR.
|
CHNJZ
force-pushed
the
atomic_fix
branch
2 times, most recently
from
August 10, 2026 03:52
221db43 to
b84d448
Compare
CHNJZ
force-pushed
the
atomic_fix
branch
2 times, most recently
from
August 10, 2026 11:38
a429bd0 to
e7e2046
Compare
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
4 similar comments
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
2 similar comments
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
2 similar comments
Contributor
Author
|
/retry |
Contributor
Author
|
/retry |
hongziqi
approved these changes
Aug 12, 2026
KanuaK
approved these changes
Aug 12, 2026
WuTYSFG
pushed a commit
that referenced
this pull request
Aug 13, 2026
xuedinge233
pushed a commit
to xuedinge233/triton-ascend
that referenced
this pull request
Aug 17, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Background
Triton Ascend previously introduced an intrusive modification in
semantic.pyto directly generate a single floating-pointatomic_minoratomic_max. To keep the Triton frontend aligned with the upstream community implementation, this frontend-specific modification needs to be reverted.For floating-point inputs, upstream Triton implements
atomic_minandatomic_maxby bitcasting the floating-point pointer and value to integer types and splitting one floating-point atomic operation into two integer atomic operations according to the sign bit:For
f32, the value and pointer are bitcast toi32/u32; forf64, they are bitcast toi64/u64. The positive and negative branches are selected using masks derived from the sign bit.Although this implementation is valid for the upstream GPU backends, the decomposed integer atomic form cannot be lowered directly by the current Ascend backend.
A5 issue
On A5 (
Ascend950PR_9579), indirect atomic operations are lowered through the SIMT fast path and eventually converted to__builtin_indirect_atomic.The decomposed form causes two problems:
UMAX/UMINoperations.The observed NPUIR compilation failure is:
A3 issue
On A3 (
Ascend910_9382), the generated unsignedUMAX/UMINoperation cannot use the supported hardware atomic path and falls back to the generic atomic implementation.For example, the
UMAX(i32)branch generated byatomic_min(f32)is decomposed by convertinguint32tof16, performingvmax, and converting the result back. However, the requireduint32 -> f16conversion is not supported.BiShengIR therefore fails in
HIVMDecomposeOpwith:Therefore, neither A3 nor A5 should receive the two integer atomic operations generated by the upstream floating-point expansion.
Solution
This PR keeps
semantic.pyaligned with upstream Triton and handles the decomposed atomic form uniformly in an early backend pass.The existing
AtomicMaxMinCanonicalizeralready restores the integer atomic pair to the original floating-point atomic operation. Previously, however, it was only executed inTritonToLinalg. By that stage,DiscreteMaskAtomicConversionmay already have rewritten the atomic mask and value, preventing the canonicalizer from recognizing the original frontend pattern.This PR reuses
AtomicMaxMinCanonicalizerat the beginning ofDiscreteMaskAccessConversionPass::runOnOperation():A separate
RewritePatternSetis applied before the existing discrete-mask patterns. This allows the split atomic operations to be merged while the original pointer, value and sign-bit mask structure are still available.The canonicalizer is also extended to recognize the sign-bit mask generated by the current upstream frontend:
The updated matching covers:
f32 atomic_minf64 atomic_minf32 atomic_maxf64 atomic_maxAfter canonicalization:
MIN + UMAXis restored to one floating-pointMIN.MAX + UMINis restored to one floating-pointMAX.UMAX/UMINbranch reaches the A3 or A5 lowering path.As a result, A3 can use its existing floating-point atomic lowering path, while A5 can generate one floating-point
__builtin_indirect_atomicwith the original direct GM pointer.