Skip to content

[Triton/Gluon] [gfx950] [dsv4.1-flash] mHC fused kernel - #5824

Merged
vgokhale merged 2 commits into
mainfrom
ahmed/mhc-fusion
Sep 25, 2026
Merged

vgokhale merged 2 commits into
mainfrom
ahmed/mhc-fusion

Conversation

@ahmed-bsod

@ahmed-bsod ahmed-bsod commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Currently in dsv4.1-flash for gfx950,
Prefill (num_tokens >= 1024) uses 5 different kernels for mHC:

  • mhc_post_kernel
  • mhc_pre_gemm_sqrsum
  • mhc_pre_big_fuse
  • _mhc_pre_mix_kernel
  • _hc_head_reduce_store_kernel

and decode (num_tokens < 1024) uses 4 kernels for mHC:

  • mhc_fused_post_pre_gemm_sqrsum
  • mhc_pre_big_fuse
  • _mhc_pre_mix_kernel
  • _hc_head_reduce_store_kernel

and then a kernel for rmsnorm on the input to the attention/moe block add_rmsnorm_quant_kernel

We fuse these kernels and replace them with a main kernel (_mhc_post_pre_delayed_main_kernel) and reduce kernel (_mhc_post_pre_delayed_reduce_kernel).

On a per seam basis (mHC work between two sub-layers) the fused kernel has a speedup of 1.16x for prefill and a speedup of 1.47x for decode compared to the unfused kernels (measured from an e2e trace of ISL 1024 / OSL 32 / CONC 32 / num_prompts 32 / max-num-batched-tokens 16384):


Prefill unfused (T=16316) Time per seam us
mhc_post_kernel 242.1
mhc_pre_gemm_sqrsum 205.9
mhc_pre_big_fuse 125.8
_mhc_pre_mix_kernel 5.8
_hc_head_reduce_store_kernel 155.7
add_rmsnorm_quant_kernel 53.9
Unfused total 789.2
Prefill fused (T=16316) Time per seam us
main kernel 626.4
reduce kernel 56.5
Fused total 682.9
Speedup 1.16x

Decode unfused (T=192) Time per seam us
mhc_fused_post_pre_gemm_sqrsum 11.5
mhc_pre_big_fuse 5.8
_mhc_pre_mix_kernel 4.2
_hc_head_reduce_store_kernel 4.5
add_rmsnorm_quant_kernel 3.9
Unfused total 29.9
Decode fused (T=192) Time per seam us
main kernel 12.1
reduce kernel 8.3
Fused total 20.4
Speedup 1.47x

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5824 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@ahmed-bsod ahmed-bsod changed the title mhc fusion [gfx950] [dsv4.1-flash] mHC fused kernel Sep 24, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Input-contract bugs, inconsistent hardcoded tuning, and the missing benchmark harness should be addressed before approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 3 Medium severity · 1 Low severity

Open (4)
What changed in this PR

Adds a two-launch Triton fusion for delayed mHC seams, reducing intermediate launches in DeepSeek-V4.1 decode and prefill paths.

Changes:

  • Adds fused main and reduction kernels.
  • Exposes a public wrapper with optional post-mixing.
  • Adds numerical and empty-input tests.
File Description
aiter/​ops/​triton/​fusions/​mhc_post_pre_delayed.py Adds the public wrapper and launch logic.
aiter/​ops/​triton/​_triton_kernels/​fusions/​mhc_post_pre_delayed.py Implements fused Triton kernels.
aiter/​ops/​triton/​fusions/​__init__.py Exports the public operation.
aiter/​ops/​triton/​_triton_kernels/​fusions/​__init__.py Exports internal kernels.
op_tests/​triton_tests/​fusions/​test_mhc_post_pre_delayed.py Adds correctness coverage.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread aiter/ops/triton/fusions/mhc_post_pre_delayed.py Outdated
Comment thread aiter/ops/triton/fusions/mhc_fused_post_pre_delayed_rmsnorm.py
Comment thread aiter/ops/triton/fusions/mhc_post_pre_delayed.py Outdated
Comment thread aiter/ops/triton/fusions/mhc_post_pre_delayed.py Outdated
- add bench script
- add tuned configs and config picking function
- the configs for gfx942 are just copies of gfx950. added them since it mentions in the aiter PR guidelines
- address some copilot comments

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The wrapper can corrupt non-contiguous output buffers and silently ignore incomplete post-mix inputs.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity

Open (1)
Resolved since last review (4)

Comment thread aiter/ops/triton/fusions/mhc_fused_post_pre_delayed_rmsnorm.py
@ahmed-bsod
ahmed-bsod marked this pull request as ready for review September 25, 2026 01:48
@ahmed-bsod
ahmed-bsod requested a review from a team September 25, 2026 01:48
@github-actions github-actions Bot changed the title [gfx950] [dsv4.1-flash] mHC fused kernel [Triton/Gluon] [gfx950] [dsv4.1-flash] mHC fused kernel Sep 25, 2026
@azaidy
azaidy requested a review from vgokhale September 25, 2026 14:08
@vgokhale
vgokhale merged commit 62988d5 into main Sep 25, 2026
149 of 162 checks passed
@vgokhale
vgokhale deleted the ahmed/mhc-fusion branch September 25, 2026 15:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants