Skip to content

[HIP] IQ2R: GLM-5.3 packed MoE path for TP4/TP8 on gfx950 - #5728

Draft
ssharma4-amd wants to merge 7 commits into
ROCm:mainfrom
ssharma4-amd:glm-5.3-iq2r
Draft

ssharma4-amd wants to merge 7 commits into
ROCm:mainfrom
ssharma4-amd:glm-5.3-iq2r

Conversation

@ssharma4-amd

@ssharma4-amd ssharma4-amd commented Sep 21, 2026 •

Copy link
Copy Markdown

Updated 2026-10-04: force-pushed. This PR is now self-contained on main (b68b0e5c7) and carries only GLM-5.3 IQ2R. Commits:

  1. 2da66c89c GLM-5.3 packed MoE path for TP4/TP8 (this PR)
  2. c947240d7 FlyDSL compile-failure fallback (small fix, needed for serving)
  3. 06d06563a decode kernels accept up to 1024 tokens; tuner times CUDA graph replay
  4. e961e89d9 single down kernel, TP4 nobarrier gate, decode retune
  5. ff693e7b5 checkpoint overlay and pack tools (compile → overlay → pack)
  6. 4452e5846 op_tests runnable as python3 FILE (as the CI GPU job runs them), black/ruff clean
  7. 73068acef encoder, materializer and routing launchers run on the tensors' GPU

Summary

Adds a GLM-5.3 IQ2R MoE (aiter/iq2r_glm53.py) for MI355X. It covers 256
routed experts plus the shared expert fused as expert 256, top-9,
hidden 6144, and intermediate 512 (TP4) or 256 (TP8).

  • Packed layout. iq2r_glm53_pack turns the generic IQ2R records into
    one packed stack per layer: quad-interleaved gate/up, with the codebook
    signs folded in. The packed stack is the only expert copy on the GPU
    (58.6 GB peak weights per GPU at TP4). The layout can be sliced to TP4 or
    TP8 with contiguous slices.
  • Kernels (iq2r_gemm_gfx950.cu, iq2r_moe_aux_gfx950.cu):
    • device IQ2R encoder and materializer;
    • decode gate (decode, nobarrier) and down (route9, single,
      packed, ordered) kernels, usable up to 1024 tokens so MTP verify
      batches can stay on them;
    • prefill gate/down for larger token counts;
    • MXFP8 route gather/quant, sort and reduce helpers.
    • Every launcher takes a HipDeviceGuard, so tensors on a GPU other than
      the current device work.
  • Dispatch through a tuned CSV (aiter/configs/iq2r_glm53_tuned.csv),
    keyed on gfx, cu_num, token, model_dim, inter_dim, expert and topk.
    • AITER_CONFIG_IQ2R_GLM53 overrides the CSV path.
    • No other environment variables.
    • Tuner: csrc/kernels/iq2r/iq2r_glm53_tune.py, with
      iq2r_glm53_untuned.csv. It times candidates by CUDA graph replay;
      eager timing adds host overhead that hid the best small-M launches.
  • Checkpoint tools. Three steps turn the block-FP8 GLM-5.3 checkpoint
    and an importance calibration file into the packed checkpoint ATOM serves
    (docs/iq2r_glm53.md has the commands):
    1. aiter.iq2r_glm5_compile encodes the routed experts of layers 3–77
      and the fused shared expert into per-layer IQ2R shards (multi-GPU).
    2. aiter.iq2r_overlay validates each compiled shard and writes a model
      directory with the generic IQ2R quantization_config
      (aiter-iq2r-overlay v2).
    3. aiter.iq2r_glm53_pack_checkpoint writes the self-contained
      glm53-packed-v1 checkpoint (about 235 GB; loads at TP4 and TP8).
  • Docs: docs/iq2r_glm53.md.

Test plan

  • Every op_test this PR adds passes when run the way the CI GPU job runs
    them (python3 op_tests/FILE, MI355X, one GPU visible):
    • test_iq2r_glm53.py 25: packed MoE vs a dense Torch reference for
      M = 1..3000 at TP4 and TP8, long-prefill chunking, packing commutes
      with TP slicing, CSV lookup;
    • test_iq2r_hip.py 4 + 1 skipped (the two-GPU test). With two GPUs
      visible: 5 passed;
    • test_iq2r_encoder.py 8, test_iq2r_format.py 5,
      test_iq2r_reference.py 5;
    • test_iq2r_glm5_compile.py 11, test_iq2r_overlay.py 6,
      test_iq2r_glm53_pack_checkpoint.py 2;
    • test_tuned_gemm_flydsl_fallback.py 2,
      test_gemm_a8w8_bpreshuffle_pad_k.py 11.
  • black (stable) and ruff 0.16.0 on the diff are clean.
  • Bitwise check against the build that scored the GSM8K result below:
    the packed MoE at TP4 and TP8 shapes for M = 1..2048, plus the encoder,
    gives byte-identical output (40/40 tensors). Each output is also
    identical across two runs.
  • TP4 serving smoke with IQ2R: load GLM-5.3 packed MoE checkpoints with TP slicing ATOM#2335, local WikiText-calibrated packed
    checkpoint, greedy: coherent, correct answers on all prompts.

Accuracy

GSM8K (lm-eval 0.4.13, chat 5-shot, all 1319 questions), ATOM TP4,
FP8 KV cache, MI355X:

model strict-match
FP8 checkpoint 97.57
MXFP4 (experts quantized online) 97.19
IQ2R, WikiText-2 32K-token calibration 93.40

The IQ2R score was measured on an earlier revision of this branch. That
revision's GLM-5.3 kernels and tuned CSV match this one; see the bitwise
check above.

Results

ATOM benchmark_serving, MI355X, OSL 1024, 10×C prompts, default launch
(no IQ2R env vars). MXFP4 = stock ATOM with the experts quantized online to
MXFP4. The non-MTP tables were measured before commits 3 and 4 (the
M ≤ 1024 extension and retune); the MTP tables include them. Output tok/s:

TP4, ISL 1024 (IQ2R peak weights 58.6 GB/GPU, 198.7k KV blocks):

C IQ2R TPOT ms MXFP4 IQ2R vs MXFP4
1 87.8 11.30 82.6 +6.3%
2 166.1 11.64 164.1 +1.3%
4 331.5 11.80 318.1 +4.2%
8 569.8 13.59 565.0 +0.9%
16 1006.1 15.43 943.9 +6.6%
32 1599.3 19.08 1441.5 +10.9%
64 2472.6 24.85 2219.6 +11.4%
128 3656.0 33.58 3321.6 +10.1%
256 4990.5 49.23 4732.0 +5.5%

TP4, ISL 8192:

C IQ2R TPOT ms MXFP4 IQ2R vs MXFP4
1 80.0 12.05 76.2 +5.1%
2 155.4 12.05 154.4 +0.7%
4 288.6 13.09 280.4 +2.9%
8 457.8 16.44 465.6 -1.7%
16 705.7 21.45 692.0 +2.0%
32 971.1 30.97 936.0 +3.7%
64 1248.2 48.50 1221.9 +2.2%
128 1514.2 80.03 1514.9 -0.0%
256 1718.6 141.16 1762.6 -2.5%

TP8, ISL 1024 (IQ2R peak weights 32.7 GB/GPU, 230.5k KV blocks):

C IQ2R TPOT ms MXFP4 IQ2R vs MXFP4
1 84.8 11.71 79.5 +6.6%
2 162.4 11.94 151.8 +7.0%
4 329.2 11.90 325.4 +1.1%
8 603.0 12.75 592.7 +1.7%
16 1109.2 13.98 1080.4 +2.7%
32 1826.0 16.57 1683.1 +8.5%
64 2894.4 21.15 2640.9 +9.6%
128 4532.0 26.98 4089.4 +10.8%
256 6330.3 38.79 5808.4 +9.0%

TP8 at ISL 8192 has not been rerun on this build yet.

MTP results

GLM-5.3's MTP layer (checkpoint layer 78) drafts 3 tokens per step
(--method mtp --num-speculative-tokens 3). The target then verifies
1 + 3 tokens per sequence, so the routed MoE runs at M = 4 × C. The MTP gain at
concurrency C therefore follows the non-MTP gain at 4C. Both arms force the
same acceptance length (--spec-decode-acceptance-length 2.8), so they do the
same work per step and only kernel speed differs.

Quantization parity matters for the draft layer: the MXFP4 arm quantizes the
layer-78 experts to MXFP4. The IQ2R arm must do the same, by excluding only
layers 3–77 from online quantization:

--online-quant-config '{"global_quant_config": "ptpc_fp8", "layer_quant_config": {"*expert*": "mxfp4"}, "exclude_layer": ["lm_head", "model.embed_tokens", "*.mlp.gate", "re:^model\\.layers\\.([3-9]|[1-6][0-9]|7[0-7])\\.mlp\\.experts"]}'

TP8, ISL 1024, OSL 1024, FP8 KV cache, same benchmark_serving settings as
above. Output tok/s:

C IQ2R MXFP4 IQ2R vs MXFP4 non-MTP at 4C
1 230.7 227.4 +1.5% +1.1%
2 423.3 409.7 +3.3% +1.7%
4 797.4 758.8 +5.1% +2.7%
8 1327.6 1238.3 +7.2% +8.5%
16 2139.4 1884.8 +13.5% +9.6%
32 3249.1 2947.2 +10.2% +10.8%
64 4685.8 4238.5 +10.6% +9.0%
128 6280.1 5833.1 +7.7% –
256 7831.0 7605.3 +3.0% –

TP4, ISL 1024, OSL 1024, BF16 KV cache, on a second MI355X machine. Output tok/s:

C IQ2R MXFP4 IQ2R vs MXFP4 non-MTP at 4C
1 199.8 192.7 +3.7% +4.2%
2 354.1 346.5 +2.2% +0.9%
4 644.3 573.5 +12.3% +6.6%
8 1044.9 906.5 +15.3% +10.9%
16 1601.9 1400.8 +14.4% +11.4%
32 2332.4 2080.2 +12.1% +10.1%

These MTP tables were measured on a development build with the same GLM-5.3
kernels and tuned CSV. A TP4 C8 check with this PR's kernels and
ROCm/ATOM#2335 (self-contained checkpoint, second pass) gives 1009.1 tok/s,
+11.3% over MXFP4; the development build gave 1025.0 to 1044.9 on that row.

At TP8 C128 and C256 the verify batch is 512 or 1024 tokens. That runs on the
prefill kernels, where IQ2R is level with MXFP4 or slightly slower (TP8 MoE at
M = 1024: 212 µs vs 204 µs), so the gain shrinks there.

Included fix

[FlyDSL] Fall back to CK/torch when a tuned kernel fails to compile

A tuned FlyDSL row can be valid while the local FlyDSL/LLD toolchain cannot
link its code object.

  • Catch DSLCompileError (raised before launch) and remember the kernel
    name per process.
  • Then use the native CK path (a8w8 bpreshuffle) or the torch path (hgemm).
  • Launch and runtime errors still propagate.

Tests: op_tests/test_tuned_gemm_flydsl_fallback.py and
op_tests/test_gemm_a8w8_bpreshuffle_pad_k.py (2 new cases).

🤖 Generated with Claude Code

@ssharma4-amd
ssharma4-amd requested a review from a team September 21, 2026 11:09
@github-actions github-actions Bot changed the title [HIP] [JIT] [Build] Add plain GLM-5.3 IQ2R 2-bit MoE support [HIP] [FlyDSL] [JIT] Add plain GLM-5.3 IQ2R 2-bit MoE support Sep 21, 2026
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5728 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@ssharma4-amd

ssharma4-amd commented Sep 25, 2026 •

Copy link
Copy Markdown
Author

GLM-5.3 benchmark results — 2026-09-25 23:16 UTC

ATOM benchmark_serving, one MI355X node. Output tokens/s; higher is better.

TP8 — 1k input / 1k output

Concurrency MXFP4 IQ2R IQ2R vs MXFP4
1 80.31 84.00 +4.59%
2 152.63 159.33 +4.39%
4 327.31 323.41 -1.19%
8 594.12 592.86 -0.21%
16 Failed 1,052.71 —
32 1,673.35 1,636.81 -2.18%
64 2,622.04 2,549.67 -2.76%
128 4,070.54 3,879.19 -4.70%
256 5,786.88 5,662.29 -2.15%

TP8 — 8k input / 1k output

Concurrency MXFP4 IQ2R IQ2R vs MXFP4
1 74.48 77.14 +3.57%
2 146.39 150.12 +2.55%
4 296.91 284.58 -4.15%
8 510.50 478.93 -6.19%
16 841.10 741.55 -11.84%
32 1,164.83 1,022.59 -12.21%
64 1,545.03 1,296.99 -16.05%
128 1,975.87 1,593.88 -19.33%
256 Stopped 1,840.95 —

TP4: pending for both workloads. MXFP4 1k/1k C16 failed; no throughput accepted.

All numeric results have zero failed requests. Random length ratio 0.8; 10×C measured requests + 2×C warmups; FP8 KV; MTP off.

MXFP4 1k/1k C1–C8 uses the mean of before/after runs. Other comparisons use a later baseline on the same node and still need fresh before/after validation.

@ssharma4-amd

ssharma4-amd commented Sep 26, 2026 •

Copy link
Copy Markdown
Author

Single-GPU MoE optimization update

One synthetic TP rank on MI355X. Microseconds per local MoE call; lower is better. These are not benchmark_serving concurrency results.

TP Tokens Routes MXFP4 µs IQ2R candidate µs IQ2R latency overhead
8 1024 spread 213.38 215.14 +0.8%
8 1024 hot 124.76 143.96 +15.4%
8 4096 spread 452.02 518.98 +14.8%
8 4096 hot 395.27 441.25 +11.6%
4 1024 spread 294.69 323.04 +9.6%
4 1024 hot 190.81 212.45 +11.3%
4 4096 spread 606.06 817.97 +35.0%
4 4096 hot 500.69 678.25 +35.5%

TP8: E261. TP4: E262. Both use exact index/sign repacking and interleaved down/reduction for4096 tokens. Stored weight size is unchanged. Each row averages fresh clean bookends; changing-input/route IQ2R checks are exact and native kernels have zero scratch.

The goal remains unmet. Dense candidates remain isolated; serving sweeps are paused. Details and all attempted optimizations are in docs/iq2r/SINGLE_GPU_ANALYSIS.md and OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

Single-GPU MoE results — E280 small-token candidate

Synthetic TP rank on MI355X. µs per complete local MoE call; lower is better. Tokens below are not serving concurrency. Same E280 configuration in all small-token rows.

TP Tokens Routes MXFP4 IQ2R IQ2R latency overhead
8 4 spread 25.21 24.67 -2.2%
8 4 hot 20.06 22.37 +11.5%
8 8 spread 36.52 35.66 -2.4%
8 8 hot 26.92 22.56 -16.2%
8 16 spread 48.72 54.72 +12.3%
8 16 hot 28.29 26.84 -5.1%
4 4 spread 36.03 32.83 -8.9%
4 4 hot 27.74 26.51 -4.4%
4 8 spread 55.12 54.23 -1.6%
4 8 hot 23.15 27.40 +18.3%
4 16 spread 93.41 81.75 -12.5%
4 16 hot 31.04 28.85 -7.1%

The gate now uses102 VGPRs versus129 for the earlier packed gate, restoring two-workgroup residency without spills. Both TPs have fresh clean bookends and rocprof counters. E284 passes30 shape/routing cases and240 changing-input steps with exact final BF16 and intermediate FP8 values/scales.

Dense selection remains E261 (TP8) / E262 (TP4):

TP Tokens Routes MXFP4 IQ2R IQ2R latency overhead
8 1024 spread 213.38 215.14 +0.8%
8 1024 hot 124.76 143.96 +15.4%
8 4096 spread 452.02 518.98 +14.8%
8 4096 hot 395.27 441.25 +11.6%
4 1024 spread 294.69 323.04 +9.6%
4 1024 hot 190.81 212.45 +11.3%
4 4096 spread 606.06 817.97 +35.0%
4 4096 hot 500.69 678.25 +35.5%

The goal remains unmet. The new down-decoder scout E283 failed correctness and is excluded. Candidates remain isolated; production integration and official benchmark_serving qualification are still pending. Details and every attempted optimization through E284 are in docs/iq2r/SINGLE_GPU_ANALYSIS.md and OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

Single-GPU MoE update: TP8 M16 now beats its fresh MXFP4 control.

Microseconds per complete local MoE call; lower is better. These are isolated shape-specific experiments, not serving concurrency or production results.

Experiment TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
E292 8 4 spread 25.30 24.47 -3.3%
E292 8 4 hot 19.97 22.23 +11.3%
E289 8 16 spread 49.86 48.88 -2.0%
E289 8 16 hot 28.59 25.51 -10.8%
E292 4 4 spread 36.20 32.64 -9.8%
E292 4 4 hot 27.86 26.32 -5.5%

E289 combines the new down schedule, M16 frontend and batched reduction. E292 improves the M4 fused-down epilogue. All rows have clean bookends; TP8 M4 hot still loses.

E280 passed 32 real-capture/TP cases. The E289 combination passed 30 broader routing/shape cases with exact final and intermediate outputs, plus 45 frontend tests. TP4 M8 remains unqualified after an unchanged-MXFP4 graph/eager tolerance failure. Dense gaps and production qualification remain; serving sweeps are still paused.

All attempts through E292 are documented in docs/iq2r/SINGLE_GPU_ANALYSIS.md and OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP8 isolated MoE: combined IQ2R versus fresh MXFP4.

Microseconds per complete local MoE call, lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
8 4 spread 25.34 24.41 -3.7%
8 4 hot 20.43 22.18 +8.5%
8 8 spread 36.46 35.59 -2.4%
8 8 hot 26.76 22.49 -16.0%
8 16 spread 48.95 46.64 -4.7%
8 16 hot 28.18 24.71 -12.3%

All rows have less than 1% bookend drift, exact IQ2R checks and native dispatch
with zero scratch. TP8 four-token hot routing still trails by 8.5%.

The combination also passed 30 synthetic TP8/TP4 cases and 32 real-capture/TP
cases with exact final and intermediate values. TP4 uses TP8 captures with
TP4 weight slices. These capture timings compare against original IQ2R.

No new TP4 MXFP4 numbers in this update; the prior TP4 table is retained.
TP4 M8's baseline correctness-bound failure remains unresolved. Dense gaps
and production/serving qualification remain; serving sweeps are paused.
Attempts through E303 are in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 isolated MoE update: fresh MXFP4 versus IQ2R.

Microseconds per complete local MoE call, lower is better. These token counts are not serving concurrency. M16 uses the combined small-token path; 1024/4096 use E308 N512 dense down.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
4 16 spread 91.83 79.70 -13.2%
4 16 hot 31.13 26.03 -16.4%
4 1024 spread 292.89 319.97 +9.2%
4 1024 hot 187.53 204.78 +9.2%
4 4096 spread 601.03 803.23 +33.6%
4 4096 hot 499.58 650.75 +30.3%

Dense IQ2R improves 0.7–3.5% over E262, but still trails MXFP4. All table rows
pass fixed correctness checks, native dispatch and timing-drift limits, zero spills.

The dense candidate also passed 39 shape/routing cases and 936 exact checks,
including route-aligned intermediate FP8 values/scales. Actual dense captures,
production integration and serving parity remain. TP4 M4's new timing is
provisional due drift; TP4 M8's older baseline-bound failure remains unresolved.
TP8's six qualified small-token comparisons remain in the preceding update.

All attempts through E311 are documented in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

M4 isolated MoE update: MXFP4 versus IQ2R.

Microseconds per complete MoE call; lower is better. Tokens here are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
8 4 spread 25.26 23.24 -8.0%
8 4 hot 19.99 19.89 -0.5%
4 4 spread 36.93 32.60 -11.7%
4 4 hot 28.47 25.78 -9.4%

TP8 now reaches hot-route parity; its 0.5% difference is too small to call a
robust lead. Both TP4 rows and TP8 spread are ahead. Fresh clean bookends pass
the drift limit, fixed correctness checks and native dispatch, zero spills.

Changes: wave sorting and paired FP8 conversion in the frontend; TP8 reuses
down weights across matching tokens, with a GPU-checked fallback. TP4 uses
wave-private codebook completion. Broader synthetic exact checks pass; E317
frontend also passes actual-capture qualification. New down paths still need
captured qualification and production integration.

Larger-token/dense gaps remain; overall serving parity is not achieved.
All attempts through E322 are in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE results: MXFP4 versus IQ2R, TP8.

Microseconds per complete call; lower is better. Token counts are not serving concurrency.

TP Tokens Routes MXFP4 µs Original IQ2R µs IQ2R candidate µs IQ2R overhead
8 64 spread 103.71 133.41 99.25 -4.3%
8 64 hot 41.47 43.87 38.43 -7.3%
8 128 spread 122.51 160.67 119.47 -2.5%
8 128 hot 44.62 60.63 49.44 +10.8%

Three cases now beat MXFP4; M128 hot remains 10.8% behind. Fresh clean bookends,
32 rotating banks, maximum drift 1.03%. Fixed correctness checks pass and native
dispatch is verified with zero scratch spills.

Changes: independent output waves skip empty down subtiles and reuse inputs;
one frontend launch overlaps sorting and input quantization. E324 also passes
real-weight/captured-input qualification of the earlier TP8/TP4 M4 changes.

The overall goal is still open. TP4 medium/dense gaps and broader qualification
remain; no new production integration and serving sweeps remain paused.
All attempts through E326 are in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 64 spread 103.76 86.67 -16.5%
8 64 hot 41.06 34.49 -16.0%
8 128 spread 122.56 104.61 -14.6%
8 128 hot 44.29 44.90 +1.4%
4 64 spread 193.15 159.39 -17.5%
4 64 hot 42.51 42.64 +0.3%
4 128 spread 213.66 170.10 -20.4%
4 128 hot 61.10 60.61 -0.8%

The fixed compact-gate policy beats MXFP4 in six of eight stable M64/M128 synthetic cases. TP8 M128 hot remains 1.4% slower; TP4 M64 hot remains 0.3% slower. These small remaining gaps are still open. TP4 r2 repeats the frozen selected candidate after the original hot rows exceeded the unchanged drift limit.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Fixed numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover static sign packing, compact gate tasks, task-prefix
generation and gate output reduction. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.47 102.79 -15.4%
8 128 hot 43.79 44.04 +0.6%
8 256 spread 135.48 116.80 -13.8%
8 256 hot 62.69 65.48 +4.5%
4 64 spread 192.76 160.17 -16.9%
4 64 hot 42.39 43.25 +2.0%
4 128 spread 213.04 170.07 -20.2%
4 128 hot 60.96 60.28 -1.1%
4 1024 spread 292.24 309.81 +6.0%
4 1024 hot 188.55 197.06 +4.5%
4 4096 spread 604.94 770.63 +27.4%
4 4096 hot 501.74 625.64 +24.7%

The selected isolated kernels beat MXFP4 on six of eight TP8 rows and three of eight TP4 rows shown. TP8 hot gaps are down to 0.6% at 128 tokens and 4.5% at 256; TP4 M64 hot remains 2.0% behind and dense cases remain 4.5–27.4% behind. The full performance goal is not achieved.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover exact M256 accumulation, codebook-read scheduling, compact gate
completion, shared-expert fusion and activation-delivery experiments. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 134.87 113.61 -15.8%
8 256 hot 62.76 64.71 +3.1%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.04 744.74 +23.5%
4 4096 hot 500.70 610.02 +21.8%

The selected isolated kernels beat MXFP4 on six of eight TP8 rows and all four compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 3.1% at 256. The combined dense TP4 kernel remains 3.3–23.5% behind. The full performance goal is not achieved.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover wider reduction, shared activation fragments, dense-kernel
combination, MFMA32 variants, real-weight checks and workgroup ordering. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.15 111.86 -17.2%
8 256 hot 62.64 64.10 +2.3%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.58 186.30 -19.9%
4 256 hot 71.31 92.46 +29.7%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.04 744.74 +23.5%
4 4096 hot 500.70 610.02 +21.8%

The selected kernels beat MXFP4 on six of the eight reported TP8 rows and five of six compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 2.3% at 256. Newly qualified TP4 M256 spread is 19.9% faster, but hot is 29.7% slower. The four reported dense TP4 rows remain 3.3–23.5% slower. The full isolated and serving goal is not achieved.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover register-record scheduling, vector gate output, ordered quad
down, persistent grid size and M32 gate pipelines. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 136.49 111.22 -18.5%
8 256 hot 63.64 66.38 +4.3%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.58 186.30 -19.9%
4 256 hot 71.31 92.46 +29.7%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 598.92 739.46 +23.5%
4 4096 hot 496.62 612.02 +23.2%

The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. TP8 M256 token-major routing improves its control by 0.8–1.7%, but hot remains 4.3% behind fresh MXFP4. TP4 M256 hot remains 29.7% behind, and the reported dense TP4 rows remain 3.3–23.5% behind. N1024 down provides a small M4096 improvement. The full isolated and serving goal remains unmet.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. New experiments cover codebook scheduling, active-row masking, task-fill
selection, gate tile width, route visits and dense down reuse. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.27 106.07 -21.6%
8 256 hot 63.12 64.07 +1.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.97 183.05 -21.4%
4 256 hot 72.54 89.55 +23.4%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 598.92 739.46 +23.5%
4 4096 hot 496.62 612.02 +23.2%

The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. New gate scheduling reduces the reported TP8 M256 hot gap to 1.5% and TP4 M256 hot gap to 23.4%. Dense TP4 remains 3.3–23.5% behind. Latest compact and dense candidates pass 4,320 real-weight operator checks. Complete isolated and serving parity remain unmet.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. New experiments cover cross-K codebook scheduling, codebook reuse, register
weight loading, removal of an obsolete barrier, and real-weight qualification. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.27 106.07 -21.6%
8 256 hot 63.12 64.07 +1.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.97 183.05 -21.4%
4 256 hot 72.54 89.55 +23.4%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

Complete isolated and serving parity remain unmet. The selected table still has six of eight TP8 rows and five of six compact TP4 rows faster than matched MXFP4. Direct codebook byte addresses improve both dense TP4 M4096 routes by 0.3–0.6% against their current control, leaving 22.5%/22.1% MXFP4 gaps. Extra activation lookahead is rejected; the compact TP8 address change is unselected.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. New experiments cover a three-slot activation ring, direct codebook byte addresses
and paired sign shifts. The ring is rejected. Byte addresses are retained only
at TP4 M4096; the compact TP8 transfer is unselected. New real-weight
qualification is still required for the retained address change. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

Selected isolated MoE results: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. These synthetic token counts are not benchmark_serving concurrency. The selected numbers are unchanged by the latest experiments.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.27 106.07 -21.6%
8 256 hot 63.12 64.07 +1.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.97 183.05 -21.4%
4 256 hot 72.54 89.55 +23.4%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

Complete isolated and serving parity remain unmet. The selected comparison is unchanged: six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. Dense TP4 M4096 remains 22.5%/22.1% slower. E388–E390 do not improve the selected policy.

E388–E390 are rejected: wider MFMA, paired weight loads and larger M reuse did not beat the selected kernels. E391 passes all 1,080 real-weight checks for the retained TP4 M4096 address-calculation change. The optimization inventory includes all attempts and preserved failures.

Every selected row has matched clean MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. MXFP4 is A4W4 requantized from the synthetic IQ2R fixture; IQ2R uses FP8 activations. This is operator performance evidence, not original-checkpoint quality.

Production integration, whole-model quality and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: updated isolated MoE results.

Microseconds per complete call; lower is better. Token counts below are synthetic batch sizes, not benchmark_serving concurrency. M256 TP8/TP4 rows are newly measured; other rows retain the previous qualified selection.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.20 181.33 -21.9%
4 256 hot 70.37 88.52 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

Complete isolated and serving parity remain unmet. Six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. New ordered-down sign planes improve the selected M256 kernels. TP8 M256 hot remains 0.9% behind; TP4 M256 hot is 25.8% behind its fresh baseline. Dense TP4 M4096 remains about 22% behind.

E393/E394 sign-plane down kernels are retained and pass 2,160 new real-weight checks. Concurrent filtered gates and direct dense gate epilogues are rejected. All attempts and failures are in the optimization inventory.

Every selected row has fresh matched MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R; IQ2R uses FP8 activations. These results do not establish original-checkpoint quality. Production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: updated isolated MoE numbers.

Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. TP8 M128 and TP4 M256 rows are newly measured.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

The tested TP8 M128 hot gap is closed: IQ2R is 1.9% faster than fresh MXFP4, with exact real-weight operator qualification. Seven of eight selected TP8 rows and five of six compact TP4 rows now beat matched MXFP4. TP8 M256 hot remains 0.9% behind; TP4 M256 hot remains 25.8% behind its latest baseline, and dense TP4 M4096 remains about 22% behind. Complete isolated and serving parity remain unmet.

The M128 gate now passes 1,080 real-weight checks and beats fresh MXFP4 on spread/hot/mixed. TP4 barrier removal passes another 1,080 real-weight checks and improves its control by 0.7–1.5%; its hot gap remains 25.8%. E402 is tracked separately and is excluded from this publication.

All selected rows pass the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: current selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. The selected numbers are unchanged from the previous update.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet.

The latest grid, LDS-layout and M32 reuse experiments produced no broad winner. Three completed experiments add 1,008 exact synthetic checks with stable bookends; the padded layout stopped at its alignment probe, and the two-word carry stopped on a nonfinite output. The inventory records each result and preserves its failure evidence. E408 diagnosis is ongoing and excluded.

All selected rows pass the unchanged 3% drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet.

Latest TP4 M256 experiment (E412), in us:

Routes MXFP4 Selected control Combined overlap
spread 224.23 179.86 188.55
hot 68.93 88.75 83.05
mixed 225.81 191.57 201.44

Overlap improves hot, but cold regressions and the remaining 20.5% hot gap keep it out of the selected table. E408–E412 add 1,917 completed checks; the inventory records all outcomes and failures. E413 TP8 transfer is running and excluded.

Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 131.03 111.48 -14.9%
8 256 hot 61.32 58.53 -4.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

All eight selected TP8 comparison rows now beat their matched MXFP4 baselines. E413 closes M256 hot with 58.53 versus 61.32 us (4.5% faster), while preserving wins on spread and mixed using one fixed policy. Five of six compact TP4 rows beat baseline; TP4 M256 hot and dense gaps remain. This is isolated operator parity within the selected TP8 coverage, not complete TP8 coverage or serving parity.

E413 TP8 M256 also passes mixed routing: 118.41 us IQ2R versus 132.92 us MXFP4 (10.9% faster). The candidate passes 315 synthetic checks plus 1,440 exact real-weight checks across three layer/rank slices, with the same compiled binary. It gives up 6.1% of the previous spread speed and 4.2% of mixed speed to close hot parity using one fixed policy. The inventory preserves that tradeoff. TP4 hot/dense and uncovered TP8 cases remain open; E415 TP4 scheduling work continues.

Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 131.03 111.48 -14.9%
8 256 hot 61.32 58.53 -4.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 594.30 730.14 +22.9%
4 4096 hot 488.88 602.89 +23.3%

All eight selected TP8 rows still beat matched MXFP4; five of six selected compact TP4 rows beat baseline. E416 improves its matched dense TP4 control by 0.52–1.44% and passes real-weight qualification, but dense parity and TP4 M256 hot remain open. No production or serving parity is claimed.

New TP4 results: E416 M4096 spread/hot is 730.14/602.89 us versus matched MXFP4 594.30/488.88 us, still 22.9%/23.3% behind. It improves the matched E386 control by 0.58%/1.44% and passes 2,160 exact real-weight checks with the same binary. E415 scalar M16 hot reaches 81.66 us versus 68.92 us MXFP4, but its cold-route tradeoff keeps E400 selected. E417/E419 M32 variants are rejected after profiling showed reduced residency. This update adds 1,554 synthetic checks; E420 continues the register-pressure experiment.

Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 M256: latest isolated MXFP4 versus IQ2R candidate.

Microseconds per complete MoE call; lower is better. Synthetic tokens are not benchmark_serving concurrency.

TP4 M256 routes MXFP4 us Scalar M16 us Parallel M32 candidate us Candidate vs MXFP4
spread 224.37 187.06 188.86 -15.8%
hot 69.00 82.58 76.06 +10.2%
mixed 225.56 200.12 206.31 -8.5%

The hot gap is now 10.2%. This candidate passes 1,440 exact real-weight checks with the same binary used for timing. The M32 residency change and parallel epilogue pass 1,386 synthetic checks across E420–E423; all timings satisfy the unchanged 3% per-arm bookend-drift limit, with zero scratch/spills.

Cold routes still favor scalar M16, so the selected 18-row coverage table remains unchanged. Synthetic MXFP4 is A4W4 requantized from IQ2R. Production integration, full-model quality and ATOM benchmark_serving acceptance remain pending. Experiments continue on a fresh Fleet node after verified runtime/fixture restoration.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 M256: latest isolated comparison. Microseconds per complete MoE call; lower is better.

TP4 M256 routes MXFP4 us M32 vector exchange us IQ2R time difference
spread 225.69 186.42 -17.4%
hot 68.98 75.57 +9.5%
mixed 226.96 200.18 -11.8%

Vector partial exchange improves the matched M32 control by 1.29% on hot routing, but the gap to MXFP4 is still 9.5%. Spread/mixed remain faster with M16; separately, grid3 improves M16 by 4.1%/3.4% on those patterns.

E426–E428 pass 1,008 exact synthetic checks. The grid3 and vector candidates pass 1,440 real-weight checks using their respective timed binaries, with native grids and zero scratch/spills verified. All timing rows satisfy the unchanged 3% drift limit. The 18 selected coverage rows remain unchanged.

Synthetic tokens are not serving concurrency; MXFP4 is A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending. Cross-collection DRAM counters are under separate investigation; these comparisons use clean latency.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 M256: latest completed isolated comparison. Microseconds per complete MoE call; lower is better.

TP4 M256 routes MXFP4 us E428 control us E433 specialized us IQ2R time difference
spread 224.63 185.96 183.58 -18.3%
hot 69.10 75.93 73.31 +6.1%
mixed 226.83 199.24 195.78 -13.7%

E433 improves its matched M32 control by 3.46% on hot routing; the remaining MXFP4 gap is 6.1%. All 315 synthetic checks and 720 real-weight checks pass on the same timed binary, with native grids, zero scratch/spills and the unchanged 3% drift rule. The 18 selected coverage rows remain unchanged; cold routing still favors the separately measured M16 ingredient.

Rejected: E432 component stores and E435 spilling M64 reuse. The known-byte counter probe supports current TCC units but leaves the historical E426 discrepancy unresolved; no old counts were rescaled.

Synthetic tokens are not serving concurrency. MXFP4 uses A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

GLM-5.3 IQ2R vs stock MXFP4: TP8 serving results, ISL 1024

Status: at TP8 with ISL 1024 and OSL 1024, IQ2R beats stock ATOM/AITER MXFP4 at 8 of 9 concurrencies. C=4 misses the parity bar by 0.02 points, and a fix for it is included below. Every request completed, with zero failures.

Method: unmodified atom.benchmarks.benchmark_serving, same node (MI355X, TP8), in the order MXFP4 → IQ2R → MXFP4.

  • Random dataset, range ratio 0.8, 10×C prompts, 2×C warmups, seed 0, --ignore-eos, request rate inf.
  • FP8 KV cache, no prefix caching, max_num_batched_tokens 4096, max_num_seqs 256, CUDA graphs [1…256].
  • The MXFP4 column is the average of the two MXFP4 runs.
  • Parity bar: IQ2R ≥ MXFP4 average − 0.5%.
C MXFP4 avg (tok/s) IQ2R (tok/s) Δ TPOT ms (MXFP4 → IQ2R) TTFT ms (MXFP4 → IQ2R)
1 79.9 83.9 +5.06% 12.42 → 11.83 104.9 → 95.5
2 152.0 159.4 +4.83% 12.75 → 12.16 105.0 → 100.9
4 326.2 324.5 −0.52% 11.99 → 12.06 104.9 → 95.9
8 594.2 605.5 +1.89% 12.99 → 12.76 155.4 → 146.4
16 1081.8 1107.5 +2.38% 14.34 → 13.99 146.6 → 153.0
32 1686.7 1818.8 +7.84% 18.02 → 16.65 190.9 → 172.6
64 2646.1 2932.6 +10.83% 23.26 → 20.96 270.2 → 234.7
128 4088.7 4496.3 +9.97% 30.06 → 27.34 408.8 → 343.6
256 5806.4 6273.7 +8.05% 42.44 → 39.26 626.9 → 581.4

The two MXFP4 runs agree to within 0.5% at every point.

Changes in this push

  • Packed decode stack: a load-time repack of the MoE weights into a decode-friendly layout, used at M2–256.
    • Enabled by IQ2R_GLM53_PACKED_STACK=1 and IQ2R_GLM53_PACKED_SMALL=1.
    • ATOM builds it in Iq2rMoEMethod, in the companion ATOM PR.
  • Prefill: a multi-block route sort, and the E243 gate plus the E261 down+reduce kernels. The gate runs as variant 33 on the packed stack.
  • M2/M4 decode: runs the packed gate followed by the route9 down on the original weights. This is the fix for the C=4 point: in a microbenchmark, the M4 MoE takes 22.35 µs per layer against MXFP4's 24.6 µs. It has not yet been confirmed end to end.
  • E436 commit: these commits were rebased onto the E436 prefetch commit already on this branch. On the production dispatch path, the MoE output is bitwise identical to the pre-rebase build at M=1–4096, and the kernel timings are within noise.

Still open

  • C=4 end to end: the serving rerun that checks the M2/M4 fix is in progress.
  • ISL 8192 at TP8: stock MXFP4 hits a GPU memory fault in graph mode on the first 8k request. The fault is in _gluon_deepgemm_fp8_paged_mqa_logits_preshuffle, the DSA indexer kernel. The same request succeeds in eager mode, and the block tables and context lengths check out. Root cause is still being investigated; IQ2R runs fine at 8k.
  • TP4: next, at ISL 1024 and 8192.
  • MTP: after TP4, compare IQ2R and MXFP4 on throughput, TPOT and acceptance rate.

🤖 Generated with Claude Code

@ssharma4-amd ssharma4-amd changed the title [HIP] [FlyDSL] [JIT] Add plain GLM-5.3 IQ2R 2-bit MoE support [HIP] IQ2R: GLM-5.3 packed MoE path for TP4/TP8 on gfx950 Sep 30, 2026
ssharma4-amd and others added 4 commits September 30, 2026 19:46
Add a 2-bit IQ2R routed-expert path for GLM-5.3 (256 routed + fused
shared expert, top-9, hidden 6144) on MI355X at TP4 and TP8.

- csrc/kernels/iq2r: IQ2R device encoder and materializer, MXFP8 route
  gather/quant, and the GLM-5.3 packed gate (quad) and down kernels.
- aiter/iq2r_glm53.py: packing to the glm53-packed-v1 layout, TP slicing,
  workspace, and tuned dispatch (configs/iq2r_glm53_tuned.csv).
- aiter/iq2r_glm5_compile.py: offline FP8 -> IQ2R expert encoder.
- aiter/ops/iq2r_{format,encoder,reference}.py: format metadata, host
  encoder and pure-Torch decoder.
- QuantType.iq2r_2bit for ATOM.
- op tests for the format, encoders, compiler and the packed MoE against
  a dense Torch reference; docs/iq2r_glm53.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A tuned FlyDSL row can be valid while the local FlyDSL/LLD toolchain
cannot link its code object. Catch DSLCompileError (raised before
launch), remember the kernel name per process and use the native CK
(a8w8 bpreshuffle) or torch (hgemm) path. Launch/runtime errors still
propagate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…aphs

The tuner timed eager launches, so host overhead hid the best decode
configs at small M. Time candidates from CUDA graph replay instead, let
decode kernels cover up to 1024 tokens (MTP verify batches are 4x the
concurrency), and regenerate the TP8/TP4 table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add a "single" down variant (one 16-row atom per pass, three workgroups per
CU) for decode-sized tasks on TP4 and TP8, let the nobarrier gate run at TP4
with 16 threads per row, and issue the codebook load ahead of the LDS store so
it overlaps the first weight loads. Retune M=2..256 under CUDA graphs.

MALL-cold MoE time vs MXFP4 improves most at TP4 (M=8 -2% -> 6%, M=16
6% -> 16%, M=256 0% -> 10%).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ssharma4-amd and others added 3 commits October 4, 2026 03:57
Add the two steps after aiter.iq2r_glm5_compile that produce the packed
checkpoint ATOM serves:

- aiter.iq2r_overlay validates every compiled shard against the compiled
  config, links the compiled and FP8 shards into one directory, indexes the
  compiled tensors in place of the FP8 experts and writes the
  aiter-iq2r-overlay v2 quantization_config.
- aiter.iq2r_glm53_pack_checkpoint relays each layer with iq2r_glm53_pack
  and writes the self-contained glm53-packed-v1 checkpoint.

The compiler now accepts Redline shared-expert importance stored as
[1, K], and docs/iq2r_glm53.md describes the full build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI runs each op_test as `python3 FILE`, so every pytest-only test needs a
__main__ entry point. Update the tuned-config assertions to the committed
CSV, apply black, and silence ruff B023 in the tuner, whose closures run
inside the loop iteration that defines them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sors' GPU

iq2r_encode_out, iq2r_materialize_out and the two routing gather/quant
launchers used the current device, so a call with tensors on another GPU
faulted. They now take a HipDeviceGuard like the GLM-5.3 MoE launchers.
The two-GPU test covers the encoder and the materializer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant