[HIP] IQ2R: GLM-5.3 packed MoE path for TP4/TP8 on gfx950 - #5728
ssharma4-amd wants to merge 7 commits into
Conversation
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
|
GLM-5.3 benchmark results — 2026-09-25 23:16 UTC ATOM TP8 — 1k input / 1k output
TP8 — 8k input / 1k output
TP4: pending for both workloads. MXFP4 1k/1k C16 failed; no throughput accepted. All numeric results have zero failed requests. Random length ratio 0.8; 10×C measured requests + 2×C warmups; FP8 KV; MTP off. MXFP4 1k/1k C1–C8 uses the mean of before/after runs. Other comparisons use a later baseline on the same node and still need fresh before/after validation. |
|
Single-GPU MoE optimization update One synthetic TP rank on MI355X. Microseconds per local MoE call; lower is better. These are not
TP8: E261. TP4: E262. Both use exact index/sign repacking and interleaved down/reduction for4096 tokens. Stored weight size is unchanged. Each row averages fresh clean bookends; changing-input/route IQ2R checks are exact and native kernels have zero scratch. The goal remains unmet. Dense candidates remain isolated; serving sweeps are paused. Details and all attempted optimizations are in |
|
Single-GPU MoE results — E280 small-token candidate Synthetic TP rank on MI355X. µs per complete local MoE call; lower is better. Tokens below are not serving concurrency. Same E280 configuration in all small-token rows.
The gate now uses102 VGPRs versus129 for the earlier packed gate, restoring two-workgroup residency without spills. Both TPs have fresh clean bookends and rocprof counters. E284 passes30 shape/routing cases and240 changing-input steps with exact final BF16 and intermediate FP8 values/scales. Dense selection remains E261 (TP8) / E262 (TP4):
The goal remains unmet. The new down-decoder scout E283 failed correctness and is excluded. Candidates remain isolated; production integration and official |
|
Single-GPU MoE update: TP8 M16 now beats its fresh MXFP4 control. Microseconds per complete local MoE call; lower is better. These are isolated shape-specific experiments, not serving concurrency or production results.
E289 combines the new down schedule, M16 frontend and batched reduction. E292 improves the M4 fused-down epilogue. All rows have clean bookends; TP8 M4 hot still loses. E280 passed 32 real-capture/TP cases. The E289 combination passed 30 broader routing/shape cases with exact final and intermediate outputs, plus 45 frontend tests. TP4 M8 remains unqualified after an unchanged-MXFP4 graph/eager tolerance failure. Dense gaps and production qualification remain; serving sweeps are still paused. All attempts through E292 are documented in |
|
TP8 isolated MoE: combined IQ2R versus fresh MXFP4. Microseconds per complete local MoE call, lower is better. Tokens are not serving concurrency.
All rows have less than 1% bookend drift, exact IQ2R checks and native dispatch The combination also passed 30 synthetic TP8/TP4 cases and 32 real-capture/TP No new TP4 MXFP4 numbers in this update; the prior TP4 table is retained. |
|
TP4 isolated MoE update: fresh MXFP4 versus IQ2R. Microseconds per complete local MoE call, lower is better. These token counts are not serving concurrency. M16 uses the combined small-token path; 1024/4096 use E308 N512 dense down.
Dense IQ2R improves 0.7–3.5% over E262, but still trails MXFP4. All table rows The dense candidate also passed 39 shape/routing cases and 936 exact checks, All attempts through E311 are documented in |
|
M4 isolated MoE update: MXFP4 versus IQ2R. Microseconds per complete MoE call; lower is better. Tokens here are not serving concurrency.
TP8 now reaches hot-route parity; its 0.5% difference is too small to call a Changes: wave sorting and paired FP8 conversion in the frontend; TP8 reuses Larger-token/dense gaps remain; overall serving parity is not achieved. |
|
New isolated MoE results: MXFP4 versus IQ2R, TP8. Microseconds per complete call; lower is better. Token counts are not serving concurrency.
Three cases now beat MXFP4; M128 hot remains 10.8% behind. Fresh clean bookends, Changes: independent output waves skip empty down subtiles and reuse inputs; The overall goal is still open. TP4 medium/dense gaps and broader qualification |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The fixed compact-gate policy beats MXFP4 in six of eight stable M64/M128 synthetic cases. TP8 M128 hot remains 1.4% slower; TP4 M64 hot remains 0.3% slower. These small remaining gaps are still open. TP4 r2 repeats the frozen selected candidate after the original hot rows exceeded the unchanged drift limit. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The selected isolated kernels beat MXFP4 on six of eight TP8 rows and three of eight TP4 rows shown. TP8 hot gaps are down to 0.6% at 128 tokens and 4.5% at 256; TP4 M64 hot remains 2.0% behind and dense cases remain 4.5–27.4% behind. The full performance goal is not achieved. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The selected isolated kernels beat MXFP4 on six of eight TP8 rows and all four compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 3.1% at 256. The combined dense TP4 kernel remains 3.3–23.5% behind. The full performance goal is not achieved. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The selected kernels beat MXFP4 on six of the eight reported TP8 rows and five of six compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 2.3% at 256. Newly qualified TP4 M256 spread is 19.9% faster, but hot is 29.7% slower. The four reported dense TP4 rows remain 3.3–23.5% slower. The full isolated and serving goal is not achieved. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. TP8 M256 token-major routing improves its control by 0.8–1.7%, but hot remains 4.3% behind fresh MXFP4. TP4 M256 hot remains 29.7% behind, and the reported dense TP4 rows remain 3.3–23.5% behind. N1024 down provides a small M4096 improvement. The full isolated and serving goal remains unmet. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. New gate scheduling reduces the reported TP8 M256 hot gap to 1.5% and TP4 M256 hot gap to 23.4%. Dense TP4 remains 3.3–23.5% behind. Latest compact and dense candidates pass 4,320 real-weight operator checks. Complete isolated and serving parity remain unmet. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
Complete isolated and serving parity remain unmet. The selected table still has six of eight TP8 rows and five of six compact TP4 rows faster than matched MXFP4. Direct codebook byte addresses improve both dense TP4 M4096 routes by 0.3–0.6% against their current control, leaving 22.5%/22.1% MXFP4 gaps. Extra activation lookahead is rejected; the compact TP8 address change is unselected. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
Selected isolated MoE results: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. These synthetic token counts are not benchmark_serving concurrency. The selected numbers are unchanged by the latest experiments.
Complete isolated and serving parity remain unmet. The selected comparison is unchanged: six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. Dense TP4 M4096 remains 22.5%/22.1% slower. E388–E390 do not improve the selected policy. E388–E390 are rejected: wider MFMA, paired weight loads and larger M reuse did not beat the selected kernels. E391 passes all 1,080 real-weight checks for the retained TP4 M4096 address-calculation change. The optimization inventory includes all attempts and preserved failures. Every selected row has matched clean MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. MXFP4 is A4W4 requantized from the synthetic IQ2R fixture; IQ2R uses FP8 activations. This is operator performance evidence, not original-checkpoint quality. Production integration, whole-model quality and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: updated isolated MoE results. Microseconds per complete call; lower is better. Token counts below are synthetic batch sizes, not benchmark_serving concurrency. M256 TP8/TP4 rows are newly measured; other rows retain the previous qualified selection.
Complete isolated and serving parity remain unmet. Six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. New ordered-down sign planes improve the selected M256 kernels. TP8 M256 hot remains 0.9% behind; TP4 M256 hot is 25.8% behind its fresh baseline. Dense TP4 M4096 remains about 22% behind. E393/E394 sign-plane down kernels are retained and pass 2,160 new real-weight checks. Concurrent filtered gates and direct dense gate epilogues are rejected. All attempts and failures are in the optimization inventory. Every selected row has fresh matched MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R; IQ2R uses FP8 activations. These results do not establish original-checkpoint quality. Production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: updated isolated MoE numbers. Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. TP8 M128 and TP4 M256 rows are newly measured.
The tested TP8 M128 hot gap is closed: IQ2R is 1.9% faster than fresh MXFP4, with exact real-weight operator qualification. Seven of eight selected TP8 rows and five of six compact TP4 rows now beat matched MXFP4. TP8 M256 hot remains 0.9% behind; TP4 M256 hot remains 25.8% behind its latest baseline, and dense TP4 M4096 remains about 22% behind. Complete isolated and serving parity remain unmet. The M128 gate now passes 1,080 real-weight checks and beats fresh MXFP4 on spread/hot/mixed. TP4 barrier removal passes another 1,080 real-weight checks and improves its control by 0.7–1.5%; its hot gap remains 25.8%. E402 is tracked separately and is excluded from this publication. All selected rows pass the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: current selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. The selected numbers are unchanged from the previous update.
The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet. The latest grid, LDS-layout and M32 reuse experiments produced no broad winner. Three completed experiments add 1,008 exact synthetic checks with stable bookends; the padded layout stopped at its alignment probe, and the two-word carry stopped on a nonfinite output. The inventory records each result and preserves its failure evidence. E408 diagnosis is ongoing and excluded. All selected rows pass the unchanged 3% drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.
The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet. Latest TP4 M256 experiment (E412), in us:
Overlap improves hot, but cold regressions and the remaining 20.5% hot gap keep it out of the selected table. E408–E412 add 1,917 completed checks; the inventory records all outcomes and failures. E413 TP8 transfer is running and excluded. Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.
All eight selected TP8 comparison rows now beat their matched MXFP4 baselines. E413 closes M256 hot with 58.53 versus 61.32 us (4.5% faster), while preserving wins on spread and mixed using one fixed policy. Five of six compact TP4 rows beat baseline; TP4 M256 hot and dense gaps remain. This is isolated operator parity within the selected TP8 coverage, not complete TP8 coverage or serving parity. E413 TP8 M256 also passes mixed routing: 118.41 us IQ2R versus 132.92 us MXFP4 (10.9% faster). The candidate passes 315 synthetic checks plus 1,440 exact real-weight checks across three layer/rank slices, with the same compiled binary. It gives up 6.1% of the previous spread speed and 4.2% of mixed speed to close hot parity using one fixed policy. The inventory preserves that tradeoff. TP4 hot/dense and uncovered TP8 cases remain open; E415 TP4 scheduling work continues. Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.
All eight selected TP8 rows still beat matched MXFP4; five of six selected compact TP4 rows beat baseline. E416 improves its matched dense TP4 control by 0.52–1.44% and passes real-weight qualification, but dense parity and TP4 M256 hot remain open. No production or serving parity is claimed. New TP4 results: E416 M4096 spread/hot is 730.14/602.89 us versus matched MXFP4 594.30/488.88 us, still 22.9%/23.3% behind. It improves the matched E386 control by 0.58%/1.44% and passes 2,160 exact real-weight checks with the same binary. E415 scalar M16 hot reaches 81.66 us versus 68.92 us MXFP4, but its cold-route tradeoff keeps E400 selected. E417/E419 M32 variants are rejected after profiling showed reduced residency. This update adds 1,554 synthetic checks; E420 continues the register-pressure experiment. Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
TP4 M256: latest isolated MXFP4 versus IQ2R candidate. Microseconds per complete MoE call; lower is better. Synthetic tokens are not benchmark_serving concurrency.
The hot gap is now 10.2%. This candidate passes 1,440 exact real-weight checks with the same binary used for timing. The M32 residency change and parallel epilogue pass 1,386 synthetic checks across E420–E423; all timings satisfy the unchanged 3% per-arm bookend-drift limit, with zero scratch/spills. Cold routes still favor scalar M16, so the selected 18-row coverage table remains unchanged. Synthetic MXFP4 is A4W4 requantized from IQ2R. Production integration, full-model quality and ATOM benchmark_serving acceptance remain pending. Experiments continue on a fresh Fleet node after verified runtime/fixture restoration. |
|
TP4 M256: latest isolated comparison. Microseconds per complete MoE call; lower is better.
Vector partial exchange improves the matched M32 control by 1.29% on hot routing, but the gap to MXFP4 is still 9.5%. Spread/mixed remain faster with M16; separately, grid3 improves M16 by 4.1%/3.4% on those patterns. E426–E428 pass 1,008 exact synthetic checks. The grid3 and vector candidates pass 1,440 real-weight checks using their respective timed binaries, with native grids and zero scratch/spills verified. All timing rows satisfy the unchanged 3% drift limit. The 18 selected coverage rows remain unchanged. Synthetic tokens are not serving concurrency; MXFP4 is A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending. Cross-collection DRAM counters are under separate investigation; these comparisons use clean latency. |
|
TP4 M256: latest completed isolated comparison. Microseconds per complete MoE call; lower is better.
E433 improves its matched M32 control by 3.46% on hot routing; the remaining MXFP4 gap is 6.1%. All 315 synthetic checks and 720 real-weight checks pass on the same timed binary, with native grids, zero scratch/spills and the unchanged 3% drift rule. The 18 selected coverage rows remain unchanged; cold routing still favors the separately measured M16 ingredient. Rejected: E432 component stores and E435 spilling M64 reuse. The known-byte counter probe supports current TCC units but leaves the historical E426 discrepancy unresolved; no old counts were rescaled. Synthetic tokens are not serving concurrency. MXFP4 uses A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending. |
9248908 to
ba4b9a8
Compare
GLM-5.3 IQ2R vs stock MXFP4: TP8 serving results, ISL 1024Status: at TP8 with ISL 1024 and OSL 1024, IQ2R beats stock ATOM/AITER MXFP4 at 8 of 9 concurrencies. C=4 misses the parity bar by 0.02 points, and a fix for it is included below. Every request completed, with zero failures. Method: unmodified
The two MXFP4 runs agree to within 0.5% at every point. Changes in this push
Still open
🤖 Generated with Claude Code |
4b450bc to
d04773b
Compare
Add a 2-bit IQ2R routed-expert path for GLM-5.3 (256 routed + fused
shared expert, top-9, hidden 6144) on MI355X at TP4 and TP8.
- csrc/kernels/iq2r: IQ2R device encoder and materializer, MXFP8 route
gather/quant, and the GLM-5.3 packed gate (quad) and down kernels.
- aiter/iq2r_glm53.py: packing to the glm53-packed-v1 layout, TP slicing,
workspace, and tuned dispatch (configs/iq2r_glm53_tuned.csv).
- aiter/iq2r_glm5_compile.py: offline FP8 -> IQ2R expert encoder.
- aiter/ops/iq2r_{format,encoder,reference}.py: format metadata, host
encoder and pure-Torch decoder.
- QuantType.iq2r_2bit for ATOM.
- op tests for the format, encoders, compiler and the packed MoE against
a dense Torch reference; docs/iq2r_glm53.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A tuned FlyDSL row can be valid while the local FlyDSL/LLD toolchain cannot link its code object. Catch DSLCompileError (raised before launch), remember the kernel name per process and use the native CK (a8w8 bpreshuffle) or torch (hgemm) path. Launch/runtime errors still propagate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aphs The tuner timed eager launches, so host overhead hid the best decode configs at small M. Time candidates from CUDA graph replay instead, let decode kernels cover up to 1024 tokens (MTP verify batches are 4x the concurrency), and regenerate the TP8/TP4 table. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add a "single" down variant (one 16-row atom per pass, three workgroups per CU) for decode-sized tasks on TP4 and TP8, let the nobarrier gate run at TP4 with 16 threads per row, and issue the codebook load ahead of the LDS store so it overlaps the first weight loads. Retune M=2..256 under CUDA graphs. MALL-cold MoE time vs MXFP4 improves most at TP4 (M=8 -2% -> 6%, M=16 6% -> 16%, M=256 0% -> 10%). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d04773b to
3fa37f5
Compare
Add the two steps after aiter.iq2r_glm5_compile that produce the packed checkpoint ATOM serves: - aiter.iq2r_overlay validates every compiled shard against the compiled config, links the compiled and FP8 shards into one directory, indexes the compiled tensors in place of the FP8 experts and writes the aiter-iq2r-overlay v2 quantization_config. - aiter.iq2r_glm53_pack_checkpoint relays each layer with iq2r_glm53_pack and writes the self-contained glm53-packed-v1 checkpoint. The compiler now accepts Redline shared-expert importance stored as [1, K], and docs/iq2r_glm53.md describes the full build. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI runs each op_test as `python3 FILE`, so every pytest-only test needs a __main__ entry point. Update the tuned-config assertions to the committed CSV, apply black, and silence ruff B023 in the tuner, whose closures run inside the loop iteration that defines them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sors' GPU iq2r_encode_out, iq2r_materialize_out and the two routing gather/quant launchers used the current device, so a call with tensors on another GPU faulted. They now take a HipDeviceGuard like the GLM-5.3 MoE launchers. The two-GPU test covers the encoder and the materializer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
184ba59 to
73068ac
Compare
Summary
Adds a GLM-5.3 IQ2R MoE (
aiter/iq2r_glm53.py) for MI355X. It covers 256routed experts plus the shared expert fused as expert 256, top-9,
hidden 6144, and intermediate 512 (TP4) or 256 (TP8).
iq2r_glm53_packturns the generic IQ2R records intoone packed stack per layer: quad-interleaved gate/up, with the codebook
signs folded in. The packed stack is the only expert copy on the GPU
(58.6 GB peak weights per GPU at TP4). The layout can be sliced to TP4 or
TP8 with contiguous slices.
iq2r_gemm_gfx950.cu,iq2r_moe_aux_gfx950.cu):decode,nobarrier) and down (route9,single,packed,ordered) kernels, usable up to 1024 tokens so MTP verifybatches can stay on them;
HipDeviceGuard, so tensors on a GPU other thanthe current device work.
aiter/configs/iq2r_glm53_tuned.csv),keyed on gfx, cu_num, token, model_dim, inter_dim, expert and topk.
AITER_CONFIG_IQ2R_GLM53overrides the CSV path.csrc/kernels/iq2r/iq2r_glm53_tune.py, withiq2r_glm53_untuned.csv. It times candidates by CUDA graph replay;eager timing adds host overhead that hid the best small-M launches.
and an importance calibration file into the packed checkpoint ATOM serves
(
docs/iq2r_glm53.mdhas the commands):aiter.iq2r_glm5_compileencodes the routed experts of layers 3–77and the fused shared expert into per-layer IQ2R shards (multi-GPU).
aiter.iq2r_overlayvalidates each compiled shard and writes a modeldirectory with the generic IQ2R
quantization_config(
aiter-iq2r-overlayv2).aiter.iq2r_glm53_pack_checkpointwrites the self-containedglm53-packed-v1checkpoint (about 235 GB; loads at TP4 and TP8).docs/iq2r_glm53.md.Test plan
them (
python3 op_tests/FILE, MI355X, one GPU visible):test_iq2r_glm53.py25: packed MoE vs a dense Torch reference forM = 1..3000 at TP4 and TP8, long-prefill chunking, packing commutes
with TP slicing, CSV lookup;
test_iq2r_hip.py4 + 1 skipped (the two-GPU test). With two GPUsvisible: 5 passed;
test_iq2r_encoder.py8,test_iq2r_format.py5,test_iq2r_reference.py5;test_iq2r_glm5_compile.py11,test_iq2r_overlay.py6,test_iq2r_glm53_pack_checkpoint.py2;test_tuned_gemm_flydsl_fallback.py2,test_gemm_a8w8_bpreshuffle_pad_k.py11.the packed MoE at TP4 and TP8 shapes for M = 1..2048, plus the encoder,
gives byte-identical output (40/40 tensors). Each output is also
identical across two runs.
checkpoint, greedy: coherent, correct answers on all prompts.
Accuracy
GSM8K (lm-eval 0.4.13, chat 5-shot, all 1319 questions), ATOM TP4,
FP8 KV cache, MI355X:
The IQ2R score was measured on an earlier revision of this branch. That
revision's GLM-5.3 kernels and tuned CSV match this one; see the bitwise
check above.
Results
ATOM
benchmark_serving, MI355X, OSL 1024, 10×C prompts, default launch(no IQ2R env vars). MXFP4 = stock ATOM with the experts quantized online to
MXFP4. The non-MTP tables were measured before commits 3 and 4 (the
M ≤ 1024 extension and retune); the MTP tables include them. Output tok/s:
TP4, ISL 1024 (IQ2R peak weights 58.6 GB/GPU, 198.7k KV blocks):
TP4, ISL 8192:
TP8, ISL 1024 (IQ2R peak weights 32.7 GB/GPU, 230.5k KV blocks):
TP8 at ISL 8192 has not been rerun on this build yet.
MTP results
GLM-5.3's MTP layer (checkpoint layer 78) drafts 3 tokens per step
(
--method mtp --num-speculative-tokens 3). The target then verifies1 + 3 tokens per sequence, so the routed MoE runs at M = 4 × C. The MTP gain at
concurrency C therefore follows the non-MTP gain at 4C. Both arms force the
same acceptance length (
--spec-decode-acceptance-length 2.8), so they do thesame work per step and only kernel speed differs.
Quantization parity matters for the draft layer: the MXFP4 arm quantizes the
layer-78 experts to MXFP4. The IQ2R arm must do the same, by excluding only
layers 3–77 from online quantization:
TP8, ISL 1024, OSL 1024, FP8 KV cache, same
benchmark_servingsettings asabove. Output tok/s:
TP4, ISL 1024, OSL 1024, BF16 KV cache, on a second MI355X machine. Output tok/s:
These MTP tables were measured on a development build with the same GLM-5.3
kernels and tuned CSV. A TP4 C8 check with this PR's kernels and
ROCm/ATOM#2335 (self-contained checkpoint, second pass) gives 1009.1 tok/s,
+11.3% over MXFP4; the development build gave 1025.0 to 1044.9 on that row.
At TP8 C128 and C256 the verify batch is 512 or 1024 tokens. That runs on the
prefill kernels, where IQ2R is level with MXFP4 or slightly slower (TP8 MoE at
M = 1024: 212 µs vs 204 µs), so the gain shrinks there.
Included fix
[FlyDSL] Fall back to CK/torch when a tuned kernel fails to compile
A tuned FlyDSL row can be valid while the local FlyDSL/LLD toolchain cannot
link its code object.
DSLCompileError(raised before launch) and remember the kernelname per process.
Tests:
op_tests/test_tuned_gemm_flydsl_fallback.pyandop_tests/test_gemm_a8w8_bpreshuffle_pad_k.py(2 new cases).🤖 Generated with Claude Code