Skip to content

[AtomicOp](fix) Fuse expanded floating-point atomic min/max - #1460

Merged
KanuaK merged 2 commits into
triton-lang:main-devfrom
CHNJZ:atomic_fix
Aug 12, 2026
Merged

KanuaK merged 2 commits into
triton-lang:main-devfrom
CHNJZ:atomic_fix

Conversation

@CHNJZ

@CHNJZ CHNJZ commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Background

Triton Ascend previously introduced an intrusive modification in semantic.py to directly generate a single floating-point atomic_min or atomic_max. To keep the Triton frontend aligned with the upstream community implementation, this frontend-specific modification needs to be reverted.

For floating-point inputs, upstream Triton implements atomic_min and atomic_max by bitcasting the floating-point pointer and value to integer types and splitting one floating-point atomic operation into two integer atomic operations according to the sign bit:

atomic_min:
  non-negative values -> signed MIN
  negative values     -> unsigned UMAX

atomic_max:
  non-negative values -> signed MAX
  negative values     -> unsigned UMIN

For f32, the value and pointer are bitcast to i32/u32; for f64, they are bitcast to i64/u64. The positive and negative branches are selected using masks derived from the sign bit.

Although this implementation is valid for the upstream GPU backends, the decomposed integer atomic form cannot be lowered directly by the current Ascend backend.

A5 issue

On A5 (Ascend950PR_9579), indirect atomic operations are lowered through the SIMT fast path and eventually converted to __builtin_indirect_atomic.

The decomposed form causes two problems:

  • The NPUIR indirect atomic interface does not support the generated unsigned UMAX/UMIN operations.
  • The operation receives a bitcasted integer pointer instead of the original floating-point GM pointer, while NPUIR atomic operations require a direct GM operand.

The observed NPUIR compilation failure is:

bishengir-compile-a5:
FunctionInterfaces.h.inc:1082:
Assertion `index < getNumArguments() && "invalid argument number"` failed.

ERR99999 UNKNOWN application exception

A3 issue

On A3 (Ascend910_9382), the generated unsigned UMAX/UMIN operation cannot use the supported hardware atomic path and falls back to the generic atomic implementation.

For example, the UMAX(i32) branch generated by atomic_min(f32) is decomposed by converting uint32 to f16, performing vmax, and converting the result back. However, the required uint32 -> f16 conversion is not supported.

BiShengIR therefore fails in HIVMDecomposeOp with:

'hivm.hir.vcast' op currently don't support cast
uint32_t_to_half_rintmode

Therefore, neither A3 nor A5 should receive the two integer atomic operations generated by the upstream floating-point expansion.

Solution

This PR keeps semantic.py aligned with upstream Triton and handles the decomposed atomic form uniformly in an early backend pass.

The existing AtomicMaxMinCanonicalizer already restores the integer atomic pair to the original floating-point atomic operation. Previously, however, it was only executed in TritonToLinalg. By that stage, DiscreteMaskAtomicConversion may already have rewritten the atomic mask and value, preventing the canonicalizer from recognizing the original frontend pattern.

This PR reuses AtomicMaxMinCanonicalizer at the beginning of DiscreteMaskAccessConversionPass::runOnOperation():

AtomicMaxMinCanonicalizer
        ↓
DiscreteMaskLoadConversion
DiscreteMaskStoreConversion
DiscreteMaskAtomicConversion

A separate RewritePatternSet is applied before the existing discrete-mask patterns. This allows the split atomic operations to be merged while the original pointer, value and sign-bit mask structure are still available.

The canonicalizer is also extended to recognize the sign-bit mask generated by the current upstream frontend:

shrui(value_bits, 31/63)
  -> cmpi ne 0
  -> xori true

The updated matching covers:

  • f32 atomic_min
  • f64 atomic_min
  • f32 atomic_max
  • f64 atomic_max

After canonicalization:

  • MIN + UMAX is restored to one floating-point MIN.
  • MAX + UMIN is restored to one floating-point MAX.
  • The bitcasted integer pointer is restored to the original floating-point GM pointer.
  • The bitcasted integer value is restored to the original floating-point value.
  • The original explicit mask, or the implicit all-true mask, is restored.
  • No unsigned UMAX/UMIN branch reaches the A3 or A5 lowering path.

As a result, A3 can use its existing floating-point atomic lowering path, while A5 can generate one floating-point __builtin_indirect_atomic with the original direct GM pointer.

@github-actions github-actions Bot added restricted-files Changes include files outside the repository-construction allowlist. compiler Changes to C/C++ compiler backend (lib/, include/) python Changes to Python runtime or bindings ascend-backend Changes to the Ascend NPU backend labels Aug 10, 2026
@github-actions

github-actions Bot commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

🔍 OpenCodeReview found 2 issue(s) in this PR.

  • ✅ Successfully posted inline: 2 comment(s)

Comment thread python/triton/language/semantic.py
@CHNJZ
CHNJZ force-pushed the atomic_fix branch 2 times, most recently from 221db43 to b84d448 Compare August 10, 2026 03:52
Comment thread python/triton/language/semantic.py
@CHNJZ
CHNJZ force-pushed the atomic_fix branch 2 times, most recently from a429bd0 to e7e2046 Compare August 10, 2026 11:38
@CHNJZ

CHNJZ commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

/retry

Comment thread python/triton/language/semantic.py
@CHNJZ

CHNJZ commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

/retry

4 similar comments
@CHNJZ

CHNJZ commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@CHNJZ

CHNJZ commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@CHNJZ

CHNJZ commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

Comment thread python/triton/language/semantic.py
Comment thread python/triton/language/semantic.py
@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

2 similar comments
@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

2 similar comments
@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@CHNJZ

CHNJZ commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

/retry

@KanuaK
KanuaK merged commit 64c7f63 into triton-lang:main-dev Aug 12, 2026
20 of 27 checks passed
xuedinge233 pushed a commit to xuedinge233/triton-ascend that referenced this pull request Aug 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ascend-backend Changes to the Ascend NPU backend compiler Changes to C/C++ compiler backend (lib/, include/) python Changes to Python runtime or bindings restricted-files Changes include files outside the repository-construction allowlist.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants