Skip to content

[Triton/Gluon] gemm_a16w16_atomic accumulate into an existing output buffer - #5895

Merged
nsusanto merged 2 commits into
mainfrom
k3-a16w16-atomic-accumulate
Oct 1, 2026
Merged

nsusanto merged 2 commits into
mainfrom
k3-a16w16-atomic-accumulate

Conversation

@nsusanto

Copy link
Copy Markdown
Contributor

Motivation

Kimi K3 fuses the outputs of shared experts and routed experts. gemm_a16w16_atomic does not have a built in accumulate for an existing accumulator buffer.

Technical Details

  • accumulate=True on gemm_a16w16_atomic, which requires y and never zeroes it:
    • when no split-K, load y, add in fp32, store;
    • when split-K, atomic adds on top of y.
  • Strided y (a column slice) is updated in place.
  • GEMM-A16W16-ATOMIC-N=896-K=3584.json: tuned gfx950 configs for the Kimi-K3 shape.

Test Plan

Added test_gemm_a16_w16_atomic_accumulate
pytest op_tests/triton_tests/gemm/basic/test_gemm_a16w16.py -k atomic

Submission Checklist

@nsusanto
nsusanto requested a review from a team September 27, 2026 22:40
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5895 --add-label <label>

One backend per PR:
A PR changes one kernel backend: [Triton/Gluon] (Triton and Gluon count as one), [HIP], [ASM], [CK], [OPUS] or [FlyDSL]. If the title ends up with two backend tags, split the PR -- as stacked pull requests when one part cannot merge without the other.

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to stop the title rewrites; labels stay in sync either way.

@github-actions github-actions Bot changed the title [Triton] gemm_a16w16_atomic accumulate into an existing output buffer [Triton/Gluon] gemm_a16w16_atomic accumulate into an existing output buffer Sep 27, 2026
@nsusanto
nsusanto force-pushed the k3-a16w16-atomic-accumulate branch from 010ec54 to ce49a45 Compare September 27, 2026 22:43
@zufayu
zufayu requested review from a team and azaidy September 28, 2026 00:31
@Boss2002n
Boss2002n requested a lite review from Copilot September 28, 2026 01:33

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Register the tuning harness and reset the accumulation buffer before each benchmark run.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

Adds accumulate=True support to Triton atomic A16W16 GEMM, with split-K/strided-output tests and gfx950 tuning.

Changes:

  • Adds in-place accumulation API and kernel behavior.
  • Adds contiguous and strided accumulation tests.
  • Adds tuning harness and gfx950 configuration.
File Description
op_tests/​triton_tests/​gemm/​basic/​test_gemm_a16w16.py Adds accumulation tests.
aiter/​ops/​triton/​utils/​_triton/​tuning/​harness_gemm_a16w16_atomic_accumulate.py Adds accumulation tuning harness.
aiter/​ops/​triton/​gemm/​basic/​gemm_a16w16_atomic.py Exposes accumulation support.
aiter/​ops/​triton/​configs/​gfx950/​triton/​gemm/​gemm_a16w16_atomic/​GEMM-A16W16-ATOMIC-N=896-K=3584.json Adds tuned gfx950 configuration.
aiter/​ops/​triton/​_triton_kernels/​gemm/​basic/​gemm_a16w16_atomic.py Implements accumulation in the kernel.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread aiter/ops/triton/utils/_triton/tuning/harness_gemm_a16w16_atomic_accumulate.py Outdated

@Boss2002n Boss2002n left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another Q is - should accumulate not be in the config json? (does it have any perf improvements on other models is another way to ask this question)

what do u think is better? config json or just let the caller pass an argument directly?

Comment thread aiter/ops/triton/utils/_triton/tuning/harness_gemm_a16w16_atomic_accumulate.py Outdated
Comment thread op_tests/triton_tests/gemm/basic/test_gemm_a16w16.py Outdated

@Boss2002n Boss2002n left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@nsusanto
nsusanto force-pushed the k3-a16w16-atomic-accumulate branch from 91dfaf3 to 37b6d77 Compare October 1, 2026 15:34

@azaidy azaidy left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@nsusanto

nsusanto commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

MI35X Test Failures are unrelated. Safe to merge.

@nsusanto
nsusanto merged commit 3412918 into main Oct 1, 2026
67 of 69 checks passed
@nsusanto
nsusanto deleted the k3-a16w16-atomic-accumulate branch October 1, 2026 19:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants