Skip to content

[Kernel] Add HIP BF16 sparse MLA for GLM-5.3-Flash H64 prefill on gfx950 - #6037

Draft
sumin-hong wants to merge 1 commit into
ROCm:mainfrom
moreh-dev:feat/glm53-bf16-sparse-mla-v1-v2
Draft

sumin-hong wants to merge 1 commit into
ROCm:mainfrom
moreh-dev:feat/glm53-bf16-sparse-mla-v1-v2

Conversation

@sumin-hong

@sumin-hong sumin-hong commented Oct 1, 2026 •

Copy link
Copy Markdown

Draft — full GSM8K OFF/v2 scores are available below.
GPQA-Diamond, RULER and v1 full-model scores remain pending.

Summary

Add an opt-in HIP BF16 sparse-MLA operator for GLM-5.3-Flash, targeting
H64 prefill with large query chunks on gfx950. In the causal-prefill benchmark
at Q2048/warm, v2 is 110.89–160.46% faster than tuned Gluon with
disjoint KV. Paired shared-KV measurements are 50.80–96.56% faster.
These are isolated operator measurements on MI355X.

H64 prefill is the performance target of this PR. H16 and decode/small-Q
measurements, including their regressions, are provided as supplementary
data below. The existing Gluon dispatch is unchanged; selecting this API
is opt-in.

Origin: AITER PR #3459

This work starts from ROCm/aiter PR #3459,
"Introduce the 1st Gen 64 and 128 Heads MLA Decode Kernel for DeepSeek V4 for
MI35x", at e961cb0c0cc3d442fdabb97617d9bb438a4740da, specifically its
persistent HipKittens m16x4 implementation. We reuse its pinned-register/fragment
helpers and BF16 LDS readers. Building on that work, this PR adds the
GLM-5.3-Flash BF16 D512 shared-K/V contract with no appended RoPE, per-query
physical CSR selection, and the Full140 v1 / hybrid v2 implementations below.

H64 prefill performance

The benchmark models causal kpool4 selection with tail entries, covering
both the first chunk (prefix0) and a chunk after a 32K-token prefix. Disjoint
KV is the primary comparison; shared KV is a paired supplementary control
with identical Q/KV values and logical selections. Selection is synthetic
and uniform, rather than generated by the model's learned indexer.

At H64/Q2048/warm, v2 uses S1 for every row below. Each implementation's
split is selected in the first run and held fixed for an independent
reverse-order repeat. All comparison columns use Faster (%):
100 × (baseline time / v2 time − 1). Positive values indicate a gain;
negative values indicate a regression. This is an operator speed ratio,
not an end-to-end serving-throughput measurement.

KV Prefix O-only Faster (%) O+LSE Faster (%)
disjoint 0 +112.04 +110.89
disjoint 32768 +153.97 +160.46
shared 0 +52.01 +50.80
shared 32768 +95.68 +96.56

For these H64 rows, Gluon uses S2 for the first chunk and S1 with a 32K
prefix. The H64 causal-prefill measurements cover seven Q values, paired KV
layouts, two prefixes, warm/cold cache and both output contracts:
4,160 timing records / 249,600 event samples, with zero accuracy exclusions.
They were measured on the submission head/build10_rebase. The supplementary
H16 measurements bring the total to 8,416 records / 504,960 samples.

The timing includes attention main and split reduction. Indexer, KV writes,
CSR conversion and the rest of the model are outside this measurement;
end-to-end prefill latency or throughput gains have not been established.

H64 prefill bandwidth — Q2048, warm

These are the same independently repeated points as the headline table. Cells
show median µs / logical GB/s (split); the final column is Faster (%)
for v2 versus tuned Gluon. V1 is included as the Full140 reference.

KV Prefix Output v1 µs / GB/s (S) v2 µs / GB/s (S) Gluon µs / GB/s (S) Faster (%)
disjoint 0 O-only 1410.727 / 1725.2 (S1) 1021.746 / 2382.0 (S1) 2166.551 / 1123.3 (S2) +112.04
disjoint 0 O+LSE 1408.208 / 1728.7 (S1) 1021.186 / 2383.8 (S1) 2153.592 / 1130.3 (S2) +110.89
disjoint 32768 O-only 1488.608 / 3079.0 (S1) 1140.746 / 4017.9 (S1) 2897.116 / 1582.0 (S1) +153.97
disjoint 32768 O+LSE 1443.548 / 3175.4 (S1) 1103.847 / 4152.6 (S1) 2875.076 / 1594.4 (S1) +160.46
shared 0 O-only 1354.428 / 1796.9 (S1) 1001.646 / 2429.8 (S1) 1522.588 / 1598.4 (S2) +52.01
shared 0 O+LSE 1356.448 / 1794.6 (S1) 1001.845 / 2429.8 (S1) 1510.808 / 1611.3 (S2) +50.80
shared 32768 O-only 1370.268 / 3344.9 (S1) 1054.246 / 4347.5 (S1) 2062.971 / 2221.7 (S1) +95.68
shared 32768 O+LSE 1374.348 / 3335.3 (S1) 1054.586 / 4346.6 (S1) 2072.871 / 2211.4 (S1) +96.56

Logical bytes count BF16 Q and O, one gathered BF16 KV vector per valid
selection, the full fixed-width CSR indices/indptr and optional FP32 LSE.
For E valid selections, D512 and top-k width2051:
B = 4QHD + 2DE + 4Q×2051 + 4(Q+1) [+ 4QH for LSE];
logical GB/s = B / median_us / 1000.
Repeated head reads, split partials and other intermediate traffic are excluded.
The same B is used for every variant at a given input/output contract.
This is useful logical bandwidth, not a hardware-counter measurement of HBM traffic.
Shared KV can reuse cache lines across queries, and warm runs can reuse data
across replays; neither is assigned an HBM-only lower bound.

H64 prefill bandwidth — Q2048, disjoint KV, cold

Cold uses a 512 MiB flush outside the timing event. Useful QK/PV work is
4EHD; padding, softmax, address arithmetic and split reduction FLOPs are excluded.

Prefix Output Variant µs Logical GB/s QK+PV TFLOP/s
0 O-only v1 1424.667 1708.3 193.0
0 O-only v2 1037.106 2346.7 265.2
0 O-only gluon 2196.712 1107.9 125.2
0 O+LSE v1 1421.388 1712.6 193.5
0 O+LSE v2 1037.446 2346.4 265.1
0 O+LSE gluon 2184.492 1114.4 125.9
32768 O-only v1 1492.288 3071.4 368.7
32768 O-only v2 1153.906 3972.0 476.8
32768 O-only gluon 2937.836 1560.1 187.3
32768 O+LSE v1 1455.188 3150.0 378.1
32768 O+LSE v2 1135.966 4035.2 484.3
32768 O+LSE gluon 2916.156 1571.9 188.7

With a 32K prefix, cold/disjoint v2 reaches 3,972–4,035 logical GB/s,
versus 1,560–1,572 GB/s
for tuned Gluon. Because the logical-byte numerator is identical, the bandwidth
ratio expresses the measured latency gain; it does not independently establish
a reduction in physical traffic or identify the source of stalls.

H64 prefill across chunk sizes — disjoint KV, warm

All result columns report Faster (%) versus tuned Gluon.

Q First chunk, O-only First chunk, O+LSE Prefix32768, O-only Prefix32768, O+LSE
1 -12.61 -13.10 -11.91 -14.82
32 -15.98 -14.24 -1.05 -2.20
128 -11.25 -11.26 +29.96 +30.61
256 -7.46 -6.81 +69.58 +70.98
512 +3.33 +2.17 +163.21 +165.09
1024 +47.06 +43.43 +162.03 +163.49
2048 +112.04 +110.89 +153.97 +160.46

Large chunks provide the strongest gains. Small first-chunk shapes regress;
the all-Q first-chunk warm GM Faster (%) is +9.64 to +9.95
with disjoint KV and -0.75 to -0.51 with shared KV. The measured gain depends on the chunk
and KV geometry; these results do not define an automatic routing policy.

Operator and implementations

aiter.sparse_mla_bf16_fwd consumes BF16 queries and a BF16 latent used
for both K and V, with D512, no appended RoPE, and H16/H64: the absorbed attention
geometry used by GLM-5.3-Flash. It accepts per-query CSR global slot ids,
masks invalid slots and preserves the full selected set, including a
2051-entry tail.

One aiter.sparse_mla_bf16_fwd API exposes two reproducible implementations:

Version H64 implementation H64 reducer
v1 Full140, eight waves, 139,520 B LDS fully unrolled
v2 Hybrid: head tiles when Q×S≤128, otherwise four waves/65,792 B LDS fully unrolled

The caller can reuse contiguous O, optional natural-log LSE and normalized
FP32 split workspace. The torch-free ctypes entry uses the current stream;
mutation/fake registration makes the operation visible to torch.compile.
Metadata checks do not copy CSR values to the host during graph capture.
The module reuses AITER's pinned HipKittens dependency and CK-free headers,
with FP32 denormal preservation and an explicit pinned-VGPR build check.

The existing Gluon dispatch remains unchanged. The new API defaults to v2
and explicit S=1; these defaults are not a claim of the optimal split for
every workload. Engine selection and fallback belong to a subsequent vLLM
integration, using this API without copying the kernel implementation.

Validation

Submission base e0cb2ddbde2e7e71c1fe2c8d05f3f6f656360e14, implementation
ea21fa0a6c5d40916d2a7112812727f54217ebd7, MI355X/gfx950.
The fixed-K appendix tables were measured at base
80a3b0b09448bb5b8ade9a5946bec7549601331a, implementation
1121378ad62b1d1e1babe366c3fcb5ed2d6d14cb (build09).
Rebasing across seven upstream commits changed none of the 14 PR files or
the Gluon/JIT/header dependencies used here. A fresh JIT build (build10_rebase)
has identical preprocessed device source and assembly after normalizing only
the generated HIP CUID symbol. The full 514-test suite and 336 native trace
dispatch checks pass again on that build. The original fixed-K sweep was not repeated
after the rebase; the causal-prefill results above were measured on
build10_rebase. Trace durations are excluded from performance tables.
HipKittens d3cd9b31cb0ff611ff64b5701f57ccdeb7712f39 is the existing AITER pin.
PyTorch 2.12.0+git6bbd260, Triton 3.7.1+gitf0b55c07, HIP7.2 runtime,
ROCm7.2.3 compiler; one otherwise idle GPU.

Check Result
Actual AITER JIT + retained ISA 21 kernels pass; main compiler VGPR<40, FP32 denormals preserved, zero spills/private memory
Final operator suite 514/514 passed: FP64432, graph48, subnormal8, H64 dispatch boundary12, views/empty-Q/compile6, metadata8
Graph scope non-default stream, changed Q/index values, both output contracts, 30 replays/case
Compile scope backend=eager, fullgraph=True; no Inductor claim
Actual launch trace 336/336 expected native dispatches, no extra/missing launches, zero scratch
H64 causal-prefill timing 4,160 records / 249,600 event samples; zero accuracy exclusions
Supplemental fixed-K timing 4,224 records / 253,440 event samples; 4,208 records eligible after address adjudication
Independent runs reversed case/variant order, fresh allocations; full input hashes match; sampled FP64 before/after
GPU ownership quarter-second KFD sampling, zero observed foreign allocations in accepted phases
Model accuracy / serving sampled model operator checks and full GSM8K OFF/v2 scores below; other model scores pending; no serving speed claim

Black 26.3.0 and clang-format 18.1.8 pass on the changed files. Ruff 0.15.7
passes on all four new Python modules; the touched aiter/__init__.py has
the same 46 pre-existing F403 diagnostics as the submission base, with no
new diagnostics. The PR contains one signed-off commit with the HIP operator,
API, tests and benchmark. The Gluon address-control patch is separate diagnostic
evidence and is not included in this PR.

Reproduce the operator tests and the supplemental fixed-K sweep:

python -m pytest -q op_tests/test_sparse_mla_bf16.py
python -m op_tests.op_benchmarks.bench_sparse_mla_bf16 \
  --verify --output sparse_mla_o_only.jsonl
python -m op_tests.op_benchmarks.bench_sparse_mla_bf16 \
  --verify --reverse --output sparse_mla_o_only_repeat.jsonl
python -m op_tests.op_benchmarks.bench_sparse_mla_bf16 \
  --verify --return-lse --output sparse_mla_o_lse.jsonl
python -m op_tests.op_benchmarks.bench_sparse_mla_bf16 \
  --verify --return-lse --reverse --output sparse_mla_o_lse_repeat.jsonl

The sweep uses disjoint 32K-slot pools per query, H16/H64,
Q=1/32/128/256/512/1024/2048, K=2048/2051, and S=1/2/4/8/16/32.
Every point has 60 graph/event samples and 10 warmups. The 512 MiB cold flush
is outside the timing event. Native versions share input/output/workspace
addresses; Gluon shares inputs/final O and keeps its own partial layout.
Trace durations are not used as timing measurements.

Model accuracy and integration

Supplemental eager GLM-5.3-Flash TP4/H16 checks exercised both versions on
short, 32K and 128K prompts. All 11 sparse layers reached the native API in
prefill and decode, with no unsupported fallback. FP64 checks passed on
385 sampled query rows for v1 and 374 for v2; the existing attention path
also passed on those same Q/KV inputs. Maximum native normalized RMS error
was 0.00281, and maximum LSE absolute error was 5.73e-6. These are sampled
operator checks inside the model. Baseline greedy A/A varied before native
execution; token identity is not an acceptance requirement.

GSM8K — GLM-5.3-Flash

Evaluation GSM8K accuracy Correct / total
Published SGLang reference, BF16 KV + TileLang 97.50% —
Kernel OFF — existing vLLM attention 97.50% 1,286 / 1,319
Kernel ON — original v2, latest full rerun 97.65% 1,288 / 1,319

Local OFF/ON scores use lm-eval 5-shot strict-match on all 1,319 questions,
4×MI355X TP4/H16, BF16 KV, batch4, temperature1.0/top_p0.95, a 32,768-token
output limit and max reasoning effort, with the same weights, prompts and
generation settings (seed0). ON uses this PR's original v2/S1 kernel.
The published reference evaluated 1,319 questions on 4×GB300 with SGLang;
its hardware, backend and evaluator differ, so it is a reference point.

Scores vary between runs, including with the unchanged kernel. The table
shows the existing full OFF result and the latest full v2 rerun; it does not
establish a statistically significant accuracy improvement or equivalence.
The latest ON evaluation resumed from saved batches after model restarts,
counting each completed batch once. GPQA-Diamond, RULER and v1 full-model
scores remain pending.

The new API defaults to v2/S1; tuned results require the displayed explicit
split. A future vLLM PR should call this operator without copying kernel
source, retain fallback, validate physical CSR mapping and graph workspace
reuse, and measure the actual serving path. Broader model accuracy and
end-to-end serving speedups, especially for TP4/H16, remain unestablished.

All validation is associated with the source/binary hashes above. A later
rebase or dispatch change requires the relevant checks again. Full wheel
installation and Inductor validation remain outside the completed scope.

Related work

Reviewed at the following pinned revisions on 2026-09-30; PR statuses below
refer to that review.

  • #3459, merged: reuses its HK
    fragment/register helpers and BF16 LDS readers. That V4 FP8+RoPE decode
    interface differs from this BF16 D512 per-query CSR interface.
  • #4919 and
    #5551, merged: the same-checkout
    Gluon sparse-MLA implementation is the primary direct comparison. It has
    broader format/geometry coverage. Both its shipped policy and explicit
    split sweep are included in the measured comparison.
  • #5721, open at 847d0148:
    gfx942 support is complementary to this gfx950-only HIP implementation.
  • #6002, open at c69f539b:
    related BF16 D512 sparse-prefill work for gfx942, including attention sinks
    and low-head-count tuning. It targets a different backend/device path;
    its reported performance is not a denominator for this gfx950 operator.
  • #5301, open at 9b5cf23e:
    FP8 KV, separate RoPE, H1–16; its reported numbers are not a same-contract
    BF16 D512 comparison and are not used as a speedup denominator here.

Supplementary performance data

H16 causal-prefill reference, including all Q2048 warm regressions

H16 is supported as a reference path. V1 uses one wave / 33,024 B LDS and
the original reducer; v2 uses four waves / 68,352 B LDS and the unrolled
reducer. The causal sweep adds 4,256 timing records / 255,360 event samples
for H16, bringing the H64+H16 total to 8,416 records / 504,960 samples.

At Q2048/warm, v2 Faster (%) versus tuned Gluon ranges from
-39.78 to -11.21 across these prefill conditions. GLM-5.3-Flash TP4 uses H16; the H64 prefill headline
therefore does not demonstrate a serving speedup for the TP4 model.

H16/Q2048/warm reference, including regressions:

KV Prefix O-only Faster (%) O+LSE Faster (%)
disjoint 0 -24.68 -24.57
disjoint 32768 -11.21 -11.34
shared 0 -39.78 -39.58
shared 32768 -26.35 -26.18

Gluon uses S2 for the first chunk, except H16/shared/O-only uses S4;
all long-prefix entries use S1. V2 uses S1 throughout this table.

Decode / small-Q reference and fixed-K aggregate results

The fixed-K operator sweep includes the small-Q shapes relevant to decode.
It does not measure end-to-end decode latency or batched serving throughput.
Its all-Q geometric means are supplementary to the causal-prefill headline
and include every measured Q, including regressions.

Splits below are selected in the first run and fixed for the independent
reverse-order repeat. Comparison columns report Faster (%), computed as
100 × (GM(baseline time / v2 time) − 1), with regressions shown as negative
values. Each summary gives 14 Q×K shapes equal weight; it is not a
serving-workload average or an end-to-end throughput measurement.

H64 fixed-K reference — Faster (%)

Output Cache v2 vs tuned Gluon v2 vs v1 v2 vs shipped Gluon
O-only warm +63.43 +14.15 +68.38
O-only cold +60.12 +9.34 +67.93
O+LSE warm +63.20 +13.96 +68.49
O+LSE cold +59.50 +9.13 +67.84

For H64/Q2048/K2048/warm, O-only is 165.22% faster than tuned
Gluon: 1108.085 µs (v2) versus 2938.834 µs (Gluon), all S1; v1 is
1436.547 µs. O+LSE is 163.33% faster: 1094.447 µs versus
2881.978 µs; v1 is 1421.368 µs.
H64 small-Q regressions remain: O-only warm Q1 and Q32 have Faster (%)
of -13.34 and -7.70 respectively (GM over both K values).

H16 reference — Faster (%)

Output Cache v2 vs tuned Gluon v2 vs v1 v2 vs shipped Gluon
O-only warm -14.89 +70.14 -9.83
O-only cold -8.14 +52.12 -2.80
O+LSE warm -15.09 +70.70 -10.10
O+LSE cold -8.21 +52.84 -2.64

H16 Faster (%) is negative versus tuned Gluon across all measured warm Q values.
For O+LSE/Q2048/K2048/warm, v2 is 825.385 µs versus 733.724 µs for Gluon.
Full shape/cache/contract tables, with H64 first and H16 reference second,
follow below. These results do not support replacing Gluon across all shapes.

The trace observes one CTA per query for v2 H64/S1 and four for Gluon;
H16 uses one in both. This validates the dispatch/head-sharing distinction.
It does not establish physical HBM reread counts or their latency share.

Gluon address-boundary exclusions and corrected-baseline control

Baseline accuracy qualification

Original Gluon H64/Q2048/S16 and S32 fail the sampled FP64 gate for both K
values. Separate address-only controls identify the partial-buffer resource
boundary: the retained LLVM IR range is 2 GiB−2 bytes. An int64 global-pointer
control and a query-rebased buffer control both pass all nine probe cases;
their sampled BF16 outputs agree bitwise. At exactly 2 GiB, two final output
elements can be wrong; at 4 GiB, later queries are affected.

H64/Q1024/S32 reaches the same boundary but can pass the tolerance gate.
Its four raw timing points per phase are also excluded after this diagnosis.
The original raw records are retained. All best/first-selected results are
unchanged; matched-split coverage accounts for these additional exclusions.

Supplemental timing of the query-rebased Gluon control covers H64 Q1024/2048,
K2048/2051 and every split, both output contracts and independent repeats: 608
timing records, 36,480 samples, zero accuracy exclusions
. S1 remains its best
split in every shape/cache/contract. With first-selected splits fixed, v2 vs
this corrected Gluon has the following Faster (%) (GM over four Q×K shapes):

Output Cache v2 vs corrected Gluon — Faster (%)
O-only warm +164.83
O-only cold +161.12
O+LSE warm +166.06
O+LSE cold +160.72

The controls are separate lab modules, not changes to the measured upstream
baseline or a general Gluon fix included in this PR. Other formats/architectures
would need their own correctness validation before an upstream addressing fix.

Complete fixed-K independent-repeat tables: small-Q/decode reference, all other shapes, H16, cache and output contracts, bandwidth and roofline

AITER sparse MLA: independent performance results

Measured AITER base 80a3b0b09448bb5b8ade9a5946bec7549601331a, measured source 1121378ad62b1d1e1babe366c3fcb5ed2d6d14cb.

v1 is Full140+U at H64 and the legacy one-wave implementation at H16. v2 is the hybrid+U implementation. This measures the isolated attention main+reducer, not a model or serving engine.

Each query owns a disjoint 32K-slot pool. H16/H64, Q=1/32/128/256/512/1024/2048, K=2048/2051, S=1/2/4/8/16/32. Native input/output/workspace addresses are shared within each case; Gluon shares inputs/final output and retains its native partial layout.

The tables use the reverse-order repeat, with each variant's split fixed from its first run. Each point has 60 graph/event samples. Warm/cold are separate; a 512 MiB cold flush is outside the timing event. The shipped Gluon policy is reported separately from its tuned split. Native API S=1 is an explicit default, so these tuned results require the displayed split.

The audit checks completed coverage, source/binary/input hashes, accuracy and sampled KFD ownership. All regressions and accuracy exclusions remain in JSON/CSV. Same-split and retrospective best-in-run results are separate in the underlying summaries.

These fixed-K results are supplementary to the H64 causal-prefill headline. All comparison columns use Faster (%) = 100 × (baseline time / v2 time − 1); negative values indicate regressions. GM summaries apply this formula to the geometric mean of 14 equally weighted Q×K speed ratios, not a serving-workload average. The timing, bandwidth and split cells retain their original measured values.

O-only: 2,112 timing records / 126,720 event samples across both runs.

After the address controls, four raw tolerance-passing Q1024/S32 timings per phase are additionally excluded: 2,104 eligible timings per output contract. The raw records remain unchanged. All best/first-selected results are unchanged; matched-split coverage uses the corrected exclusions.

O+LSE: 2,112 timing records / 126,720 event samples across both runs.

After the address controls, four raw tolerance-passing Q1024/S32 timings per phase are additionally excluded: 2,104 eligible timings per output contract. The raw records remain unchanged. All best/first-selected results are unchanged; matched-split coverage uses the corrected exclusions.

H64 — fixed-K reference

O-only: Faster (%) (GM)

Cache v2 vs v1 v2 vs tuned Gluon v2 vs shipped Gluon
warm +14.15 (14/14) +63.43 (14/14) +68.38 (14/14)
cold +9.34 (14/14) +60.12 (14/14) +67.93 (14/14)

O-only: H64, K2048, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 22.680 / 98.6 (S32) 21.180 / 105.6 (S32) 17.880 / 125.1 (S32) 23.680 / 94.4 (S8) -15.58
32 44.680 / 1601.7 (S8) 49.180 / 1455.2 (S16) 43.600 / 1641.4 (S4) 43.180 / 1657.4 (S4) -11.35
128 112.561 / 2543.2 (S2) 101.601 / 2817.5 (S4) 129.140 / 2216.7 (S1) 125.440 / 2282.1 (S1) +27.11
256 191.681 / 2986.9 (S1) 170.760 / 3352.8 (S2) 286.881 / 1995.7 (S1) 288.741 / 1982.8 (S1) +68.00
512 364.522 / 3141.2 (S1) 284.901 / 4019.1 (S1) 737.503 / 1552.6 (S1) 735.144 / 1557.6 (S1) +158.86
1024 706.563 / 3241.2 (S1) 553.943 / 4134.2 (S1) 1499.308 / 1527.4 (S1) 1495.907 / 1530.9 (S1) +170.66
2048 1436.547 / 3188.3 (S1) 1108.085 / 4133.4 (S1) 2938.834 / 1558.5 (S1) 2931.314 / 1562.5 (S1) +165.22

O-only: H64, K2051, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 23.920 / 93.6 (S32) 23.900 / 93.7 (S32) 21.260 / 105.3 (S32) 26.520 / 84.4 (S8) -11.05
32 46.141 / 1553.2 (S8) 49.301 / 1453.6 (S16) 47.380 / 1512.5 (S4) 47.081 / 1522.2 (S4) -3.90
128 114.261 / 2508.8 (S2) 102.520 / 2796.1 (S4) 131.041 / 2187.5 (S1) 127.140 / 2254.6 (S1) +27.82
256 193.441 / 2963.8 (S1) 174.360 / 3288.1 (S2) 293.522 / 1953.2 (S1) 292.981 / 1956.8 (S1) +68.34
512 369.001 / 3107.4 (S1) 287.141 / 3993.2 (S1) 747.524 / 1533.9 (S1) 745.683 / 1537.7 (S1) +160.33
1024 721.844 / 3176.9 (S1) 560.342 / 4092.6 (S1) 1471.947 / 1558.0 (S1) 1469.147 / 1560.9 (S1) +162.69
2048 1459.687 / 3142.1 (S1) 1123.806 / 4081.2 (S1) 2916.294 / 1572.7 (S1) 2912.794 / 1574.6 (S1) +159.50

O-only: H64, K2048, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 24.680 / 90.6 (S32) 21.820 / 102.5 (S32) 18.440 / 121.3 (S32) 26.760 / 83.6 (S8) -15.49
32 52.661 / 1359.0 (S8) 56.600 / 1264.4 (S8) 57.040 / 1254.6 (S4) 56.560 / 1265.3 (S4) +0.78
128 124.500 / 2299.3 (S2) 129.121 / 2217.0 (S4) 152.561 / 1876.4 (S1) 151.740 / 1886.5 (S1) +18.15
256 211.221 / 2710.5 (S1) 208.001 / 2752.5 (S2) 323.362 / 1770.5 (S1) 323.602 / 1769.2 (S1) +55.46
512 382.121 / 2996.6 (S1) 325.422 / 3518.7 (S1) 796.704 / 1437.2 (S1) 795.404 / 1439.6 (S1) +144.82
1024 725.963 / 3154.6 (S1) 589.063 / 3887.7 (S1) 1559.848 / 1468.2 (S1) 1558.888 / 1469.1 (S1) +164.80
2048 1437.946 / 3185.2 (S1) 1123.165 / 4077.9 (S1) 2982.674 / 1535.6 (S1) 2981.354 / 1536.3 (S1) +165.56

O-only: H64, K2051, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 26.241 / 85.3 (S32) 24.960 / 89.7 (S32) 21.601 / 103.7 (S32) 29.820 / 75.1 (S8) -13.46
32 54.181 / 1322.7 (S8) 58.160 / 1232.2 (S8) 60.161 / 1191.2 (S4) 59.780 / 1198.8 (S4) +3.44
128 133.001 / 2155.3 (S2) 138.641 / 2067.6 (S4) 162.281 / 1766.4 (S1) 161.561 / 1774.3 (S1) +17.05
256 212.961 / 2692.1 (S1) 211.001 / 2717.1 (S2) 328.982 / 1742.7 (S1) 329.542 / 1739.7 (S1) +55.91
512 390.002 / 2940.1 (S1) 324.841 / 3529.8 (S1) 803.024 / 1427.9 (S1) 802.124 / 1429.5 (S1) +147.20
1024 730.743 / 3138.2 (S1) 584.722 / 3921.9 (S1) 1516.507 / 1512.2 (S1) 1515.388 / 1513.3 (S1) +159.36
2048 1471.187 / 3117.6 (S1) 1144.905 / 4006.0 (S1) 2956.154 / 1551.5 (S1) 2952.254 / 1553.6 (S1) +158.20

O+LSE: Faster (%) (GM)

Cache v2 vs v1 v2 vs tuned Gluon v2 vs shipped Gluon
warm +13.96 (14/14) +63.20 (14/14) +68.49 (14/14)
cold +9.13 (14/14) +59.50 (14/14) +67.84 (14/14)

O+LSE: H64, K2048, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 22.240 / 100.6 (S32) 20.640 / 108.4 (S32) 17.141 / 130.5 (S32) 22.920 / 97.6 (S8) -16.95
32 44.421 / 1611.3 (S8) 48.840 / 1465.5 (S16) 43.140 / 1659.1 (S4) 42.780 / 1673.1 (S4) -11.67
128 111.901 / 2558.5 (S2) 100.561 / 2847.0 (S4) 131.620 / 2175.2 (S1) 124.361 / 2302.1 (S1) +30.89
256 188.621 / 3035.7 (S1) 168.961 / 3388.9 (S2) 279.861 / 2046.0 (S1) 281.582 / 2033.5 (S1) +65.64
512 361.602 / 3167.0 (S1) 283.742 / 4036.0 (S1) 758.825 / 1509.1 (S1) 756.244 / 1514.3 (S1) +167.43
1024 711.105 / 3220.8 (S1) 555.984 / 4119.5 (S1) 1499.549 / 1527.4 (S1) 1499.329 / 1527.6 (S1) +169.71
2048 1421.368 / 3222.7 (S1) 1094.447 / 4185.4 (S1) 2881.978 / 1589.4 (S1) 2879.878 / 1590.6 (S1) +163.33

O+LSE: H64, K2051, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 23.800 / 94.1 (S32) 23.760 / 94.3 (S32) 20.100 / 111.4 (S32) 25.900 / 86.5 (S8) -15.40
32 45.040 / 1591.3 (S8) 49.240 / 1455.6 (S16) 46.881 / 1528.8 (S4) 46.541 / 1540.0 (S4) -4.79
128 112.700 / 2543.8 (S2) 101.581 / 2822.3 (S4) 129.641 / 2211.4 (S1) 125.721 / 2280.4 (S1) +27.62
256 193.021 / 2970.6 (S1) 173.721 / 3300.6 (S2) 290.441 / 1974.2 (S1) 292.361 / 1961.2 (S1) +67.19
512 368.623 / 3110.9 (S1) 286.562 / 4001.8 (S1) 745.385 / 1538.5 (S1) 744.485 / 1540.3 (S1) +160.11
1024 718.125 / 3193.8 (S1) 556.603 / 4120.6 (S1) 1469.689 / 1560.5 (S1) 1468.069 / 1562.3 (S1) +164.05
2048 1462.270 / 3136.9 (S1) 1133.428 / 4047.0 (S1) 3023.439 / 1517.2 (S1) 3020.739 / 1518.5 (S1) +166.75

O+LSE: H64, K2048, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 24.900 / 89.8 (S32) 22.120 / 101.1 (S32) 18.240 / 122.6 (S32) 26.961 / 83.0 (S8) -17.54
32 51.180 / 1398.5 (S8) 56.020 / 1277.6 (S8) 55.341 / 1293.3 (S4) 54.800 / 1306.1 (S4) -1.21
128 124.641 / 2297.0 (S2) 129.421 / 2212.1 (S4) 152.081 / 1882.5 (S1) 152.821 / 1873.4 (S1) +17.51
256 205.462 / 2786.8 (S1) 202.741 / 2824.2 (S2) 311.642 / 1837.3 (S1) 310.341 / 1845.0 (S1) +53.71
512 377.483 / 3033.7 (S1) 321.002 / 3567.5 (S1) 808.585 / 1416.3 (S1) 806.985 / 1419.1 (S1) +151.89
1024 729.805 / 3138.3 (S1) 591.763 / 3870.4 (S1) 1560.490 / 1467.7 (S1) 1557.530 / 1470.5 (S1) +163.70
2048 1455.389 / 3147.4 (S1) 1135.087 / 4035.6 (S1) 2924.319 / 1566.4 (S1) 2921.318 / 1568.0 (S1) +157.63

O+LSE: H64, K2051, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 26.360 / 85.0 (S32) 25.040 / 89.4 (S32) 21.360 / 104.9 (S32) 30.000 / 74.7 (S8) -14.70
32 54.421 / 1317.0 (S8) 58.861 / 1217.7 (S8) 60.321 / 1188.2 (S4) 59.841 / 1197.7 (S4) +2.48
128 127.600 / 2246.8 (S2) 132.220 / 2168.3 (S4) 154.541 / 1855.1 (S1) 155.001 / 1849.6 (S1) +16.88
256 214.061 / 2678.6 (S1) 210.501 / 2723.9 (S2) 328.682 / 1744.5 (S1) 329.642 / 1739.4 (S1) +56.14
512 388.483 / 2951.9 (S1) 323.542 / 3544.4 (S1) 803.545 / 1427.1 (S1) 803.445 / 1427.3 (S1) +148.36
1024 730.624 / 3139.1 (S1) 586.123 / 3913.0 (S1) 1515.410 / 1513.5 (S1) 1515.409 / 1513.5 (S1) +158.55
2048 1463.929 / 3133.4 (S1) 1149.747 / 3989.6 (S1) 3073.460 / 1492.5 (S1) 3071.719 / 1493.3 (S1) +167.32

H16 — reference results

O-only: Faster (%) (GM)

Cache v2 vs v1 v2 vs tuned Gluon v2 vs shipped Gluon
warm +70.14 (14/14) -14.89 (14/14) -9.83 (14/14)
cold +52.12 (14/14) -8.14 (14/14) -2.80 (14/14)

O-only: H16, K2048, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 41.920 / 51.0 (S32) 18.940 / 112.9 (S32) 15.960 / 134.0 (S32) 22.360 / 95.6 (S8) -15.73
32 54.460 / 1256.3 (S32) 31.180 / 2194.3 (S16) 25.820 / 2649.8 (S16) 28.720 / 2382.3 (S8) -17.19
128 109.480 / 2499.8 (S8) 69.140 / 3958.3 (S4) 55.241 / 4954.3 (S4) 55.500 / 4931.1 (S4) -20.10
256 202.201 / 2707.0 (S4) 129.401 / 4229.9 (S2) 105.840 / 5171.6 (S2) 109.581 / 4995.0 (S2) -18.21
512 374.641 / 2922.0 (S2) 222.881 / 4911.7 (S1) 197.761 / 5535.5 (S1) 195.701 / 5593.8 (S1) -11.27
1024 709.363 / 3086.5 (S1) 424.742 / 5154.7 (S1) 383.962 / 5702.2 (S1) 381.382 / 5740.8 (S1) -9.60
2048 1406.906 / 3112.4 (S1) 830.624 / 5271.8 (S1) 742.064 / 5900.9 (S1) 740.424 / 5914.0 (S1) -10.66

O-only: H16, K2051, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 45.440 / 47.1 (S32) 22.600 / 94.7 (S32) 18.720 / 114.4 (S32) 25.061 / 85.4 (S8) -17.17
32 59.520 / 1151.2 (S32) 35.040 / 1955.4 (S16) 30.040 / 2280.9 (S16) 31.580 / 2169.6 (S8) -14.27
128 113.000 / 2425.4 (S8) 73.040 / 3752.4 (S4) 57.681 / 4751.6 (S4) 57.560 / 4761.5 (S4) -21.03
256 207.041 / 2647.5 (S4) 135.281 / 4051.9 (S2) 110.920 / 4941.8 (S2) 113.320 / 4837.1 (S2) -18.01
512 377.942 / 2900.7 (S2) 226.981 / 4829.9 (S1) 202.081 / 5425.0 (S1) 200.021 / 5480.9 (S1) -10.97
1024 713.163 / 3074.5 (S1) 432.262 / 5072.4 (S1) 381.842 / 5742.1 (S1) 380.342 / 5764.8 (S1) -11.66
2048 1421.767 / 3084.3 (S1) 843.244 / 5200.4 (S1) 747.164 / 5869.1 (S1) 744.844 / 5887.4 (S1) -11.39

O-only: H16, K2048, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 50.560 / 42.3 (S32) 21.200 / 100.9 (S32) 17.960 / 119.0 (S32) 27.680 / 77.2 (S8) -15.28
32 67.281 / 1016.9 (S32) 45.900 / 1490.6 (S16) 42.600 / 1606.1 (S8) 42.701 / 1602.3 (S8) -7.19
128 133.941 / 2043.3 (S8) 114.060 / 2399.4 (S4) 110.000 / 2488.0 (S4) 109.740 / 2493.9 (S4) -3.56
256 229.561 / 2384.4 (S4) 186.141 / 2940.6 (S2) 173.401 / 3156.6 (S2) 173.221 / 3159.9 (S2) -6.84
512 403.482 / 2713.2 (S2) 274.601 / 3986.6 (S1) 258.242 / 4239.1 (S1) 256.721 / 4264.2 (S1) -5.96
1024 741.504 / 2952.7 (S1) 473.282 / 4626.1 (S1) 437.782 / 5001.2 (S1) 438.162 / 4996.9 (S1) -7.50
2048 1440.087 / 3040.7 (S1) 876.984 / 4993.1 (S1) 801.024 / 5466.6 (S1) 799.324 / 5478.2 (S1) -8.66

O-only: H16, K2051, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 54.181 / 39.5 (S32) 25.000 / 85.6 (S32) 21.040 / 101.8 (S32) 30.680 / 69.8 (S8) -15.84
32 74.120 / 924.4 (S32) 49.640 / 1380.3 (S16) 45.760 / 1497.3 (S8) 45.920 / 1492.1 (S8) -7.82
128 145.201 / 1887.5 (S8) 122.381 / 2239.5 (S4) 117.701 / 2328.6 (S4) 117.580 / 2330.9 (S4) -3.82
256 233.661 / 2345.9 (S4) 182.521 / 3003.2 (S2) 173.721 / 3155.3 (S2) 172.841 / 3171.4 (S2) -4.82
512 408.822 / 2681.6 (S2) 277.422 / 3951.7 (S1) 258.282 / 4244.6 (S1) 257.481 / 4257.8 (S1) -6.90
1024 739.244 / 2966.0 (S1) 472.502 / 4640.4 (S1) 431.602 / 5080.1 (S1) 429.882 / 5100.4 (S1) -8.66
2048 1444.107 / 3036.6 (S1) 881.264 / 4976.0 (S1) 792.284 / 5534.9 (S1) 791.724 / 5538.8 (S1) -10.10

O+LSE: Faster (%) (GM)

Cache v2 vs v1 v2 vs tuned Gluon v2 vs shipped Gluon
warm +70.70 (14/14) -15.09 (14/14) -10.10 (14/14)
cold +52.84 (14/14) -8.21 (14/14) -2.64 (14/14)

O+LSE: H16, K2048, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 41.680 / 51.3 (S32) 19.240 / 111.1 (S32) 15.860 / 134.8 (S32) 22.400 / 95.5 (S8) -17.57
32 55.760 / 1227.1 (S32) 31.060 / 2202.9 (S16) 25.900 / 2641.8 (S16) 28.740 / 2380.7 (S8) -16.61
128 109.301 / 2504.0 (S8) 69.580 / 3933.4 (S4) 54.920 / 4983.4 (S4) 54.721 / 5001.5 (S4) -21.07
256 200.401 / 2731.4 (S4) 125.261 / 4369.9 (S2) 104.601 / 5233.0 (S2) 106.440 / 5142.5 (S2) -16.49
512 372.142 / 2941.7 (S2) 220.901 / 4955.8 (S1) 193.622 / 5654.1 (S1) 192.021 / 5701.2 (S1) -12.35
1024 707.144 / 3096.3 (S1) 425.702 / 5143.3 (S1) 384.062 / 5700.9 (S1) 379.382 / 5771.2 (S1) -9.78
2048 1407.869 / 3110.4 (S1) 825.385 / 5305.4 (S1) 733.724 / 5968.2 (S1) 730.644 / 5993.3 (S1) -11.11

O+LSE: H16, K2051, warm

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 45.360 / 47.2 (S32) 22.600 / 94.7 (S32) 18.780 / 114.0 (S32) 25.320 / 84.6 (S8) -16.90
32 60.180 / 1138.6 (S32) 34.960 / 1960.0 (S16) 30.040 / 2280.9 (S16) 31.601 / 2168.3 (S8) -14.07
128 113.501 / 2414.8 (S8) 73.021 / 3753.5 (S4) 57.241 / 4788.2 (S4) 57.300 / 4783.3 (S4) -21.61
256 206.322 / 2656.8 (S4) 134.241 / 4083.4 (S2) 109.660 / 4998.7 (S2) 112.681 / 4864.8 (S2) -18.31
512 376.242 / 2913.9 (S2) 225.382 / 4864.3 (S1) 200.441 / 5469.6 (S1) 198.401 / 5525.8 (S1) -11.07
1024 710.065 / 3088.0 (S1) 430.863 / 5089.0 (S1) 380.782 / 5758.3 (S1) 379.742 / 5774.1 (S1) -11.62
2048 1418.088 / 3092.4 (S1) 842.285 / 5206.4 (S1) 744.824 / 5887.7 (S1) 742.705 / 5904.5 (S1) -11.57

O+LSE: H16, K2048, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 50.721 / 42.2 (S32) 21.560 / 99.2 (S32) 18.120 / 118.0 (S32) 28.240 / 75.7 (S8) -15.96
32 67.300 / 1016.7 (S32) 44.380 / 1541.7 (S16) 41.861 / 1634.5 (S8) 41.880 / 1633.8 (S8) -5.68
128 134.521 / 2034.5 (S8) 114.461 / 2391.1 (S4) 110.121 / 2485.3 (S4) 109.981 / 2488.5 (S4) -3.79
256 222.921 / 2455.5 (S4) 174.421 / 3138.2 (S2) 162.641 / 3365.5 (S2) 161.501 / 3389.3 (S2) -6.75
512 396.822 / 2758.8 (S2) 264.861 / 4133.3 (S1) 246.961 / 4432.9 (S1) 247.421 / 4424.6 (S1) -6.76
1024 739.145 / 2962.2 (S1) 473.423 / 4624.8 (S1) 438.983 / 4987.7 (S1) 438.843 / 4989.2 (S1) -7.27
2048 1433.408 / 3055.0 (S1) 864.205 / 5067.1 (S1) 781.365 / 5604.3 (S1) 781.845 / 5600.8 (S1) -9.59

O+LSE: H16, K2051, cold

Cells: median µs / logical GB/s (split S).

Q v1 v2 Gluon tuned Gluon shipped v2 vs tuned Gluon — Faster (%)
1 54.640 / 39.2 (S32) 25.200 / 85.0 (S32) 21.320 / 100.4 (S32) 31.440 / 68.1 (S8) -15.40
32 74.421 / 920.7 (S32) 49.440 / 1385.9 (S16) 46.140 / 1485.1 (S8) 46.240 / 1481.8 (S8) -6.67
128 139.041 / 1971.2 (S8) 118.921 / 2304.7 (S4) 112.740 / 2431.1 (S4) 112.041 / 2446.3 (S4) -5.20
256 234.261 / 2340.0 (S4) 182.701 / 3000.3 (S2) 173.441 / 3160.5 (S2) 173.081 / 3167.1 (S2) -5.07
512 408.822 / 2681.7 (S2) 277.801 / 3946.4 (S1) 257.981 / 4249.6 (S1) 259.002 / 4232.9 (S1) -7.13
1024 737.885 / 2971.5 (S1) 474.022 / 4625.6 (S1) 431.903 / 5076.7 (S1) 431.683 / 5079.3 (S1) -8.89
2048 1448.409 / 3027.7 (S1) 888.805 / 4933.9 (S1) 801.405 / 5472.0 (S1) 801.985 / 5468.1 (S1) -9.83

Roofline reference: Q2048/K2048 cold

AMD's MI355X datasheet gives per-GPU theoretical ceilings of 8 TB/s HBM and 2.5166 PFLOP/s dense BF16. These are advertised ceilings, not measured sustained rates. This kernel does not use hardware structured sparsity.

Logical bytes count Q, one gathered KV copy, CSR indices, O and optional LSE. Repeated reads and partial buffers are excluded. Logical bandwidth is a useful-byte proxy, not a physical HBM counter or utilization measurement. QK/PV TFLOP/s excludes softmax, address arithmetic and other scalar/vector work. The ideal model is memory-limited at these intensities; the actual bottleneck and extra traffic remain unmeasured. Warm cache hits can invalidate an HBM-only lower bound, hence this table uses cold points.

Contract H Variant µs Logical GB/s QK+PV TFLOP/s Ideal lower bound µs
O-only 64 v1 1437.946 3185.2 382.3 572.524
O-only 64 v2 1123.165 4077.9 489.5 572.524
O-only 64 gluon 2982.674 1535.6 184.3 572.524
O-only 64 gluon_default 2981.354 1536.3 184.4 572.524
O+LSE 64 v1 1455.389 3147.4 377.7 572.589
O+LSE 64 v2 1135.087 4035.6 484.3 572.589
O+LSE 64 gluon 2924.319 1566.4 188.0 572.589
O+LSE 64 gluon_default 2921.318 1568.0 188.2 572.589
O-only 16 v1 1440.087 3040.7 95.4 547.358
O-only 16 v2 876.984 4993.1 156.7 547.358
O-only 16 gluon 801.024 5466.6 171.6 547.358
O-only 16 gluon_default 799.324 5478.2 171.9 547.358
O+LSE 16 v1 1433.408 3055.0 95.9 547.374
O+LSE 16 v2 864.205 5067.1 159.0 547.374
O+LSE 16 gluon 781.365 5604.3 175.9 547.374
O+LSE 16 gluon_default 781.845 5600.8 175.8 547.374

Interpretation and routing boundary

The H16 and H64 results must remain separate. The measured H64 medium/large-Q gains support an opt-in HIP path; they do not support replacing Gluon for every head count or query shape. The existing Gluon dispatcher remains unchanged. A future vLLM adapter must explicitly select supported geometry and an appropriate split, retain fallback, and separately validate its physical CSR mapping and serving behavior.

At Q2048/S1, the retained trace observes one CTA per query for v2 H64 versus four for Gluon H64. At H16 both use one. This supports the source-level distinction in head sharing; it does not measure physical KV rereads or establish how much of the latency difference they cause.

The raw tolerance gates exclude H64/Q2048/S16 and S32. Address controls additionally invalidate H64/Q1024/S32 at the 2 GiB partial boundary, even when its tolerance gate passes. The original errors and controls are retained separately. Gluon's shipped and best valid paths at Q1024/Q2048 use S1, so the primary comparisons retain full shape coverage.

AI assistance

OpenAI Codex assisted with implementation, tests, analysis and documentation.
The submitter is responsible for reviewing the code and the stated evidence.

Expose Full140 v1 and hybrid v2 through one CSR attention API for BF16
D512 latent K/V and H16/H64. Reuse HipKittens, support caller-owned output
and split workspace, and preserve current-stream and graph semantics.

Include FP64 and graph tests, dispatch-boundary coverage, a pinned-VGPR
assembly check, and a reproducible same-checkout Gluon benchmark.

Signed-off-by: Sumin Hong <sumin.hong@moreh.io>
Co-authored-by: OpenAI Codex <codex@openai.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs (added automatically when gfx942 configs change); main branch always runs both MI35X and MI300X
ci:triton-355 Run the full Triton test suite on MI35X, not only the tests the change affects
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 6037 --add-label <label>

One backend per PR:
A PR changes one kernel backend: [Triton/Gluon] (Triton and Gluon count as one), [HIP], [ASM], [CK], [OPUS] or [FlyDSL]. If the title ends up with two backend tags, split the PR -- as stacked pull requests when one part cannot merge without the other.

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to stop the title rewrites; labels stay in sync either way.

@sumin-hong sumin-hong changed the title [Kernel] Add HIP BF16 sparse MLA for H64 prefill on gfx950 [Kernel] Add HIP BF16 sparse MLA for GLM-5.3-Flash H64 prefill on gfx950 Oct 1, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant