Skip to content

IQ2R: load GLM-5.3 packed MoE checkpoints with TP slicing - #2335

Draft
ssharma4-amd wants to merge 4 commits into
ROCm:mainfrom
ssharma4-amd:glm-5.3-iq2r
Draft

ssharma4-amd wants to merge 4 commits into
ROCm:mainfrom
ssharma4-amd:glm-5.3-iq2r

Conversation

@ssharma4-amd

@ssharma4-amd ssharma4-amd commented Sep 21, 2026 •

Copy link
Copy Markdown

Updated 2026-10-04: force-pushed. This PR is now self-contained on main (922b35196) and carries only GLM-5.3 IQ2R. Commits:

  1. 06078d99c load GLM-5.3 packed MoE checkpoints with TP slicing (this PR)
  2. a9b8e36db compile-cache hash includes max_model_len / max_num_seqs (small fix)
  3. 6bae854d5 indexer CP compile-key test stub gets the two new fields
  4. 337460e0e ruff fixes in the new tests

Kernels, dispatch and the checkpoint tools are in ROCm/aiter#5728.

Summary

  • quant_spec: Iq2rParser reads the aiter-iq2r-overlay v2 contract.
    Modules listed in iq2r_modules get the IQ2R spec. Every other layer keeps
    the base checkpoint's quantization (base_quantization_config).
  • Iq2rMoEMethod recognizes quantization_config.iq2r_layout = "glm53-packed-v1".
    • It allocates the packed per-rank buffers.
    • It slices the packed gate/up and down records for the TP rank at load
      time; the same checkpoint serves TP4 and TP8.
    • It calls aiter.iq2r_glm53.iq2r_glm53_moe_out (SiLU activation).
    • Unknown layouts are rejected.
  • deepseek_v2.py: weight mappings for the packed tensors.
  • config.py: iq2r accepts --online-quant-config, so the non-expert
    layers can be PTPC FP8.

Launch:

python -m atom.entrypoints.openai_server --model <GLM-5.3-IQ2R-packed> \
  -tp 4 --kv-cache-dtype fp8 --online-quant-config \
  '{"global_quant_config": "ptpc_fp8", "layer_quant_config": {"*shared_experts*": "mxfp4"}, "exclude_layer": ["lm_head", "model.embed_tokens", "*.mlp.gate", "*.mlp.experts"]}'

With MTP, the draft layer (layer 78) keeps checkpoint experts, which online
quantization turns into MXFP4. Only layers 3–77 are excluded:

python -m atom.entrypoints.openai_server --model <GLM-5.3-IQ2R-packed> \
  -tp 4 --kv-cache-dtype fp8 --method mtp --num-speculative-tokens 3 \
  --online-quant-config '{"global_quant_config": "ptpc_fp8", "layer_quant_config": {"*expert*": "mxfp4"}, "exclude_layer": ["lm_head", "model.embed_tokens", "*.mlp.gate", "re:^model\\.layers\\.([3-9]|[1-6][0-9]|7[0-7])\\.mlp\\.experts"]}'

Test plan

  • With [HIP] IQ2R: GLM-5.3 packed MoE path for TP4/TP8 on gfx950 aiter#5728 installed (MI355X):
    • tests/test_iq2r_moe_method.py 11 passed: overlay v2 parsing, contract
      and layout rejection, exact checkpoint shapes, TP4 create_weights,
      rank-shard slicing at (4,3) and (8,5), and the packed apply. Without an
      IQ2R-capable aiter the module skips.
    • tests/test_config_compile_hash.py 3 passed. These tests are new and
      fail without the fix.
    • tests/test_indexer_cp_gate.py 23 passed.
  • black and ruff are clean.
  • TP4 serving smoke with aiter#5728, local WikiText-calibrated packed
    checkpoint, greedy: coherent, correct answers on all prompts.
  • GSM8K (lm-eval chat 5-shot, all 1319 questions, TP4, FP8 KV): IQ2R 93.40
    vs MXFP4 97.19 and FP8 97.57. Details are in [HIP] IQ2R: GLM-5.3 packed MoE path for TP4/TP8 on gfx950 aiter#5728.

Depends on ROCm/aiter#5728.

Results (MXFP4 vs IQ2R, TP4 and TP8)

ATOM benchmark_serving, MI355X, OSL 1024, 10×C prompts, default launch
(no IQ2R env vars). MXFP4 = stock ATOM with the experts quantized online to
MXFP4. The non-MTP tables were measured before the aiter M ≤ 1024 extension
and retune; the MTP tables include it. Output tok/s:

TP4, ISL 1024 (IQ2R peak weights 58.6 GB/GPU, 198.7k KV blocks):

C IQ2R TPOT ms MXFP4 IQ2R vs MXFP4
1 87.8 11.30 82.6 +6.3%
2 166.1 11.64 164.1 +1.3%
4 331.5 11.80 318.1 +4.2%
8 569.8 13.59 565.0 +0.9%
16 1006.1 15.43 943.9 +6.6%
32 1599.3 19.08 1441.5 +10.9%
64 2472.6 24.85 2219.6 +11.4%
128 3656.0 33.58 3321.6 +10.1%
256 4990.5 49.23 4732.0 +5.5%

TP4, ISL 8192:

C IQ2R TPOT ms MXFP4 IQ2R vs MXFP4
1 80.0 12.05 76.2 +5.1%
2 155.4 12.05 154.4 +0.7%
4 288.6 13.09 280.4 +2.9%
8 457.8 16.44 465.6 -1.7%
16 705.7 21.45 692.0 +2.0%
32 971.1 30.97 936.0 +3.7%
64 1248.2 48.50 1221.9 +2.2%
128 1514.2 80.03 1514.9 -0.0%
256 1718.6 141.16 1762.6 -2.5%

TP8, ISL 1024 (IQ2R peak weights 32.7 GB/GPU, 230.5k KV blocks):

C IQ2R TPOT ms MXFP4 IQ2R vs MXFP4
1 84.8 11.71 79.5 +6.6%
2 162.4 11.94 151.8 +7.0%
4 329.2 11.90 325.4 +1.1%
8 603.0 12.75 592.7 +1.7%
16 1109.2 13.98 1080.4 +2.7%
32 1826.0 16.57 1683.1 +8.5%
64 2894.4 21.15 2640.9 +9.6%
128 4532.0 26.98 4089.4 +10.8%
256 6330.3 38.79 5808.4 +9.0%

TP8 at ISL 8192 has not been rerun on this build yet.

MTP results

GLM-5.3's MTP layer (checkpoint layer 78) drafts 3 tokens per step
(--method mtp --num-speculative-tokens 3). The target then verifies
1 + 3 tokens per sequence, so the routed MoE runs at M = 4 × C. The MTP gain at
concurrency C therefore follows the non-MTP gain at 4C. Both arms force the
same acceptance length (--spec-decode-acceptance-length 2.8), so they do the
same work per step and only kernel speed differs.

Quantization parity matters for the draft layer: the MXFP4 arm quantizes the
layer-78 experts to MXFP4. The IQ2R arm must do the same, by excluding only
layers 3–77 from online quantization:

--online-quant-config '{"global_quant_config": "ptpc_fp8", "layer_quant_config": {"*expert*": "mxfp4"}, "exclude_layer": ["lm_head", "model.embed_tokens", "*.mlp.gate", "re:^model\\.layers\\.([3-9]|[1-6][0-9]|7[0-7])\\.mlp\\.experts"]}'

TP8, ISL 1024, OSL 1024, FP8 KV cache, same benchmark_serving settings as
above. Output tok/s:

C IQ2R MXFP4 IQ2R vs MXFP4 non-MTP at 4C
1 230.7 227.4 +1.5% +1.1%
2 423.3 409.7 +3.3% +1.7%
4 797.4 758.8 +5.1% +2.7%
8 1327.6 1238.3 +7.2% +8.5%
16 2139.4 1884.8 +13.5% +9.6%
32 3249.1 2947.2 +10.2% +10.8%
64 4685.8 4238.5 +10.6% +9.0%
128 6280.1 5833.1 +7.7% –
256 7831.0 7605.3 +3.0% –

TP4, ISL 1024, OSL 1024, BF16 KV cache, on a second MI355X machine. Output tok/s:

C IQ2R MXFP4 IQ2R vs MXFP4 non-MTP at 4C
1 199.8 192.7 +3.7% +4.2%
2 354.1 346.5 +2.2% +0.9%
4 644.3 573.5 +12.3% +6.6%
8 1044.9 906.5 +15.3% +10.9%
16 1601.9 1400.8 +14.4% +11.4%
32 2332.4 2080.2 +12.1% +10.1%

These MTP tables were measured on a development build with the same GLM-5.3
kernels and tuned CSV; its ATOM IQ2R method served only the packed layout. A
TP4 C8 check with this PR's loader and aiter#5728 (self-contained checkpoint,
second pass) gives 1009.1 tok/s, +11.3% over MXFP4; the development build gave
1025.0 to 1044.9 on that row.

At TP8 C128 and C256 the verify batch is 512 or 1024 tokens. That runs on the
prefill kernels, where IQ2R is level with MXFP4 or slightly slower (TP8 MoE at
M = 1024: 212 µs vs 204 µs), so the gain shrinks there.

Included fix

config: include max_model_len and max_num_seqs in the compile cache hash

The sparse-attention indexer passes max_model_len and
max_num_seqs * max_model_len to its custom op as Python ints, so Dynamo
bakes them into the compiled graph. The failure:

  • A server restarted with a larger --max-model-len reused an artifact
    built for a smaller one.
  • It then hit an illegal memory access on the first decode past the old
    width.

Reproduced on GLM-5.3 TP4: an ISL 1024 run followed by ISL 8192.

The fix adds both fields to Config.compute_hash. With it, the same
sequence completes; see the TP4 ISL 8192 sweep above.

Tests: tests/test_config_compile_hash.py (fails without the fix). The
indexer CP test stub in tests/test_indexer_cp_gate.py now has the two new
fields.

🤖 Generated with Claude Code

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 2335 --add-label <label>

@zufayu
zufayu requested a review from JiaoliangYu September 22, 2026 01:51
@jamesETsmith

jamesETsmith commented Sep 22, 2026 •

Copy link
Copy Markdown

@ssharma4-amd thanks for doing this, IQ2R would be great to have in ATOM, a couple of questions here:

  • can you share accuracy results here with and without IQ2R for GSM8k?
  • Can you provide details on how to generate the per-expert diagonal second-moments?
  • I'm not an expert in IQ2R, but don't you need a dataset to weight the quantization when you generate your importance matrix?
  • What did you use for testing here and can you share any details about how you generated it?

Comment thread recipes/GLM-5.3-Flash.md Outdated
This tree can replace the routed FP8 experts in transformer layers 3–44 with
AITER's native-basis IQ2R format. Attention, dense MLPs, the shared experts,
and checkpoint layer 45 (MTP) retain the base model's block-FP8 configuration.
The current runtime is TP1/EP1 only.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this correct? You show results in the PR description from TP4 runs.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi James, this PR is still a WIP. I need to review certain changes since we have been working both with GPT OSS 120B and GLM 5.3

@ssharma4-amd
ssharma4-amd marked this pull request as draft September 22, 2026 16:39
@zufayu
zufayu removed the request for review from JiaoliangYu September 23, 2026 00:29
@ssharma4-amd

ssharma4-amd commented Sep 25, 2026 •

Copy link
Copy Markdown
Author

GLM-5.3 benchmark results — 2026-09-25 23:16 UTC

ATOM benchmark_serving, one MI355X node. Output tokens/s; higher is better.

TP8 — 1k input / 1k output

Concurrency MXFP4 IQ2R IQ2R vs MXFP4
1 80.31 84.00 +4.59%
2 152.63 159.33 +4.39%
4 327.31 323.41 -1.19%
8 594.12 592.86 -0.21%
16 Failed 1,052.71 —
32 1,673.35 1,636.81 -2.18%
64 2,622.04 2,549.67 -2.76%
128 4,070.54 3,879.19 -4.70%
256 5,786.88 5,662.29 -2.15%

TP8 — 8k input / 1k output

Concurrency MXFP4 IQ2R IQ2R vs MXFP4
1 74.48 77.14 +3.57%
2 146.39 150.12 +2.55%
4 296.91 284.58 -4.15%
8 510.50 478.93 -6.19%
16 841.10 741.55 -11.84%
32 1,164.83 1,022.59 -12.21%
64 1,545.03 1,296.99 -16.05%
128 1,975.87 1,593.88 -19.33%
256 Stopped 1,840.95 —

TP4: pending for both workloads. MXFP4 1k/1k C16 failed; no throughput accepted.

All numeric results have zero failed requests. Random length ratio 0.8; 10×C measured requests + 2×C warmups; FP8 KV; MTP off.

MXFP4 1k/1k C1–C8 uses the mean of before/after runs. Other comparisons use a later baseline on the same node and still need fresh before/after validation.

@ssharma4-amd

ssharma4-amd commented Sep 26, 2026 •

Copy link
Copy Markdown
Author

Single-GPU MoE optimization update

One synthetic TP rank on MI355X. Microseconds per local MoE call; lower is better. These are not benchmark_serving concurrency results.

TP Tokens Routes MXFP4 µs IQ2R candidate µs IQ2R latency overhead
8 1024 spread 213.38 215.14 +0.8%
8 1024 hot 124.76 143.96 +15.4%
8 4096 spread 452.02 518.98 +14.8%
8 4096 hot 395.27 441.25 +11.6%
4 1024 spread 294.69 323.04 +9.6%
4 1024 hot 190.81 212.45 +11.3%
4 4096 spread 606.06 817.97 +35.0%
4 4096 hot 500.69 678.25 +35.5%

TP8: E261. TP4: E262. Both use exact index/sign repacking and interleaved down/reduction for4096 tokens. Stored weight size is unchanged. Each row averages fresh clean bookends; changing-input/route IQ2R checks are exact and native kernels have zero scratch.

The goal remains unmet. Dense candidates remain isolated; serving sweeps are paused. Details and all attempted optimizations are in docs/iq2r/SINGLE_GPU_ANALYSIS.md and OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

Single-GPU MoE results — E280 small-token candidate

Synthetic TP rank on MI355X. µs per complete local MoE call; lower is better. Tokens below are not serving concurrency. Same E280 configuration in all small-token rows.

TP Tokens Routes MXFP4 IQ2R IQ2R latency overhead
8 4 spread 25.21 24.67 -2.2%
8 4 hot 20.06 22.37 +11.5%
8 8 spread 36.52 35.66 -2.4%
8 8 hot 26.92 22.56 -16.2%
8 16 spread 48.72 54.72 +12.3%
8 16 hot 28.29 26.84 -5.1%
4 4 spread 36.03 32.83 -8.9%
4 4 hot 27.74 26.51 -4.4%
4 8 spread 55.12 54.23 -1.6%
4 8 hot 23.15 27.40 +18.3%
4 16 spread 93.41 81.75 -12.5%
4 16 hot 31.04 28.85 -7.1%

The gate now uses102 VGPRs versus129 for the earlier packed gate, restoring two-workgroup residency without spills. Both TPs have fresh clean bookends and rocprof counters. E284 passes30 shape/routing cases and240 changing-input steps with exact final BF16 and intermediate FP8 values/scales.

Dense selection remains E261 (TP8) / E262 (TP4):

TP Tokens Routes MXFP4 IQ2R IQ2R latency overhead
8 1024 spread 213.38 215.14 +0.8%
8 1024 hot 124.76 143.96 +15.4%
8 4096 spread 452.02 518.98 +14.8%
8 4096 hot 395.27 441.25 +11.6%
4 1024 spread 294.69 323.04 +9.6%
4 1024 hot 190.81 212.45 +11.3%
4 4096 spread 606.06 817.97 +35.0%
4 4096 hot 500.69 678.25 +35.5%

The goal remains unmet. The new down-decoder scout E283 failed correctness and is excluded. Candidates remain isolated; production integration and official benchmark_serving qualification are still pending. Details and every attempted optimization through E284 are in docs/iq2r/SINGLE_GPU_ANALYSIS.md and OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

Single-GPU MoE update: TP8 M16 now beats its fresh MXFP4 control.

Microseconds per complete local MoE call; lower is better. These are isolated shape-specific experiments, not serving concurrency or production results.

Experiment TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
E292 8 4 spread 25.30 24.47 -3.3%
E292 8 4 hot 19.97 22.23 +11.3%
E289 8 16 spread 49.86 48.88 -2.0%
E289 8 16 hot 28.59 25.51 -10.8%
E292 4 4 spread 36.20 32.64 -9.8%
E292 4 4 hot 27.86 26.32 -5.5%

E289 combines the new down schedule, M16 frontend and batched reduction. E292 improves the M4 fused-down epilogue. All rows have clean bookends; TP8 M4 hot still loses.

E280 passed 32 real-capture/TP cases. The E289 combination passed 30 broader routing/shape cases with exact final and intermediate outputs, plus 45 frontend tests. TP4 M8 remains unqualified after an unchanged-MXFP4 graph/eager tolerance failure. Dense gaps and production qualification remain; serving sweeps are still paused.

All attempts through E292 are documented in docs/iq2r/SINGLE_GPU_ANALYSIS.md and OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP8 isolated MoE: combined IQ2R versus fresh MXFP4.

Microseconds per complete local MoE call, lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
8 4 spread 25.34 24.41 -3.7%
8 4 hot 20.43 22.18 +8.5%
8 8 spread 36.46 35.59 -2.4%
8 8 hot 26.76 22.49 -16.0%
8 16 spread 48.95 46.64 -4.7%
8 16 hot 28.18 24.71 -12.3%

All rows have less than 1% bookend drift, exact IQ2R checks and native dispatch
with zero scratch. TP8 four-token hot routing still trails by 8.5%.

The combination also passed 30 synthetic TP8/TP4 cases and 32 real-capture/TP
cases with exact final and intermediate values. TP4 uses TP8 captures with
TP4 weight slices. These capture timings compare against original IQ2R.

No new TP4 MXFP4 numbers in this update; the prior TP4 table is retained.
TP4 M8's baseline correctness-bound failure remains unresolved. Dense gaps
and production/serving qualification remain; serving sweeps are paused.
Attempts through E303 are in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 isolated MoE update: fresh MXFP4 versus IQ2R.

Microseconds per complete local MoE call, lower is better. These token counts are not serving concurrency. M16 uses the combined small-token path; 1024/4096 use E308 N512 dense down.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
4 16 spread 91.83 79.70 -13.2%
4 16 hot 31.13 26.03 -16.4%
4 1024 spread 292.89 319.97 +9.2%
4 1024 hot 187.53 204.78 +9.2%
4 4096 spread 601.03 803.23 +33.6%
4 4096 hot 499.58 650.75 +30.3%

Dense IQ2R improves 0.7–3.5% over E262, but still trails MXFP4. All table rows
pass fixed correctness checks, native dispatch and timing-drift limits, zero spills.

The dense candidate also passed 39 shape/routing cases and 936 exact checks,
including route-aligned intermediate FP8 values/scales. Actual dense captures,
production integration and serving parity remain. TP4 M4's new timing is
provisional due drift; TP4 M8's older baseline-bound failure remains unresolved.
TP8's six qualified small-token comparisons remain in the preceding update.

All attempts through E311 are documented in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

M4 isolated MoE update: MXFP4 versus IQ2R.

Microseconds per complete MoE call; lower is better. Tokens here are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R latency overhead
8 4 spread 25.26 23.24 -8.0%
8 4 hot 19.99 19.89 -0.5%
4 4 spread 36.93 32.60 -11.7%
4 4 hot 28.47 25.78 -9.4%

TP8 now reaches hot-route parity; its 0.5% difference is too small to call a
robust lead. Both TP4 rows and TP8 spread are ahead. Fresh clean bookends pass
the drift limit, fixed correctness checks and native dispatch, zero spills.

Changes: wave sorting and paired FP8 conversion in the frontend; TP8 reuses
down weights across matching tokens, with a GPU-checked fallback. TP4 uses
wave-private codebook completion. Broader synthetic exact checks pass; E317
frontend also passes actual-capture qualification. New down paths still need
captured qualification and production integration.

Larger-token/dense gaps remain; overall serving parity is not achieved.
All attempts through E322 are in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE results: MXFP4 versus IQ2R, TP8.

Microseconds per complete call; lower is better. Token counts are not serving concurrency.

TP Tokens Routes MXFP4 µs Original IQ2R µs IQ2R candidate µs IQ2R overhead
8 64 spread 103.71 133.41 99.25 -4.3%
8 64 hot 41.47 43.87 38.43 -7.3%
8 128 spread 122.51 160.67 119.47 -2.5%
8 128 hot 44.62 60.63 49.44 +10.8%

Three cases now beat MXFP4; M128 hot remains 10.8% behind. Fresh clean bookends,
32 rotating banks, maximum drift 1.03%. Fixed correctness checks pass and native
dispatch is verified with zero scratch spills.

Changes: independent output waves skip empty down subtiles and reuse inputs;
one frontend launch overlaps sorting and input quantization. E324 also passes
real-weight/captured-input qualification of the earlier TP8/TP4 M4 changes.

The overall goal is still open. TP4 medium/dense gaps and broader qualification
remain; no new production integration and serving sweeps remain paused.
All attempts through E326 are in docs/iq2r/OPTIMIZATION_INVENTORY.md.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 64 spread 103.76 86.67 -16.5%
8 64 hot 41.06 34.49 -16.0%
8 128 spread 122.56 104.61 -14.6%
8 128 hot 44.29 44.90 +1.4%
4 64 spread 193.15 159.39 -17.5%
4 64 hot 42.51 42.64 +0.3%
4 128 spread 213.66 170.10 -20.4%
4 128 hot 61.10 60.61 -0.8%

The fixed compact-gate policy beats MXFP4 in six of eight stable M64/M128 synthetic cases. TP8 M128 hot remains 1.4% slower; TP4 M64 hot remains 0.3% slower. These small remaining gaps are still open. TP4 r2 repeats the frozen selected candidate after the original hot rows exceeded the unchanged drift limit.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Fixed numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover static sign packing, compact gate tasks, task-prefix
generation and gate output reduction. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.47 102.79 -15.4%
8 128 hot 43.79 44.04 +0.6%
8 256 spread 135.48 116.80 -13.8%
8 256 hot 62.69 65.48 +4.5%
4 64 spread 192.76 160.17 -16.9%
4 64 hot 42.39 43.25 +2.0%
4 128 spread 213.04 170.07 -20.2%
4 128 hot 60.96 60.28 -1.1%
4 1024 spread 292.24 309.81 +6.0%
4 1024 hot 188.55 197.06 +4.5%
4 4096 spread 604.94 770.63 +27.4%
4 4096 hot 501.74 625.64 +24.7%

The selected isolated kernels beat MXFP4 on six of eight TP8 rows and three of eight TP4 rows shown. TP8 hot gaps are down to 0.6% at 128 tokens and 4.5% at 256; TP4 M64 hot remains 2.0% behind and dense cases remain 4.5–27.4% behind. The full performance goal is not achieved.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover exact M256 accumulation, codebook-read scheduling, compact gate
completion, shared-expert fusion and activation-delivery experiments. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 134.87 113.61 -15.8%
8 256 hot 62.76 64.71 +3.1%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.04 744.74 +23.5%
4 4096 hot 500.70 610.02 +21.8%

The selected isolated kernels beat MXFP4 on six of eight TP8 rows and all four compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 3.1% at 256. The combined dense TP4 kernel remains 3.3–23.5% behind. The full performance goal is not achieved.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover wider reduction, shared activation fragments, dense-kernel
combination, MFMA32 variants, real-weight checks and workgroup ordering. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.15 111.86 -17.2%
8 256 hot 62.64 64.10 +2.3%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.58 186.30 -19.9%
4 256 hot 71.31 92.46 +29.7%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.04 744.74 +23.5%
4 4096 hot 500.70 610.02 +21.8%

The selected kernels beat MXFP4 on six of the eight reported TP8 rows and five of six compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 2.3% at 256. Newly qualified TP4 M256 spread is 19.9% faster, but hot is 29.7% slower. The four reported dense TP4 rows remain 3.3–23.5% slower. The full isolated and serving goal is not achieved.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. Changes cover register-record scheduling, vector gate output, ordered quad
down, persistent grid size and M32 gate pipelines. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 136.49 111.22 -18.5%
8 256 hot 63.64 66.38 +4.3%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.58 186.30 -19.9%
4 256 hot 71.31 92.46 +29.7%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 598.92 739.46 +23.5%
4 4096 hot 496.62 612.02 +23.2%

The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. TP8 M256 token-major routing improves its control by 0.8–1.7%, but hot remains 4.3% behind fresh MXFP4. TP4 M256 hot remains 29.7% behind, and the reported dense TP4 rows remain 3.3–23.5% behind. N1024 down provides a small M4096 improvement. The full isolated and serving goal remains unmet.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. New experiments cover codebook scheduling, active-row masking, task-fill
selection, gate tile width, route visits and dense down reuse. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.27 106.07 -21.6%
8 256 hot 63.12 64.07 +1.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.97 183.05 -21.4%
4 256 hot 72.54 89.55 +23.4%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 598.92 739.46 +23.5%
4 4096 hot 496.62 612.02 +23.2%

The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. New gate scheduling reduces the reported TP8 M256 hot gap to 1.5% and TP4 M256 hot gap to 23.4%. Dense TP4 remains 3.3–23.5% behind. Latest compact and dense candidates pass 4,320 real-weight operator checks. Complete isolated and serving parity remain unmet.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. New experiments cover cross-K codebook scheduling, codebook reuse, register
weight loading, removal of an obsolete barrier, and real-weight qualification. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

New isolated MoE comparison: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. Tokens are not serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.27 106.07 -21.6%
8 256 hot 63.12 64.07 +1.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.97 183.05 -21.4%
4 256 hot 72.54 89.55 +23.4%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

Complete isolated and serving parity remain unmet. The selected table still has six of eight TP8 rows and five of six compact TP4 rows faster than matched MXFP4. Direct codebook byte addresses improve both dense TP4 M4096 routes by 0.3–0.6% against their current control, leaving 22.5%/22.1% MXFP4 gaps. Extra activation lookahead is rejected; the compact TP8 address change is unselected.

Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit.
Selected-arm numerical checks pass, native kernels are verified, and scratch use is
zero. New experiments cover a three-slot activation ring, direct codebook byte addresses
and paired sign shifts. The ring is rejected. Byte addresses are retained only
at TP4 M4096; the compact TP8 transfer is unselected. New real-weight
qualification is still required for the retained address change. All attempts are listed in
docs/iq2r/OPTIMIZATION_INVENTORY.md.

The full serving goal remains open. These are isolated experiments; production
integration and ATOM benchmark_serving acceptance are still pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

Selected isolated MoE results: MXFP4 versus IQ2R.

Microseconds per complete call; lower is better. These synthetic token counts are not benchmark_serving concurrency. The selected numbers are unchanged by the latest experiments.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.27 106.07 -21.6%
8 256 hot 63.12 64.07 +1.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.97 183.05 -21.4%
4 256 hot 72.54 89.55 +23.4%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

Complete isolated and serving parity remain unmet. The selected comparison is unchanged: six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. Dense TP4 M4096 remains 22.5%/22.1% slower. E388–E390 do not improve the selected policy.

E388–E390 are rejected: wider MFMA, paired weight loads and larger M reuse did not beat the selected kernels. E391 passes all 1,080 real-weight checks for the retained TP4 M4096 address-calculation change. The optimization inventory includes all attempts and preserved failures.

Every selected row has matched clean MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. MXFP4 is A4W4 requantized from the synthetic IQ2R fixture; IQ2R uses FP8 activations. This is operator performance evidence, not original-checkpoint quality.

Production integration, whole-model quality and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: updated isolated MoE results.

Microseconds per complete call; lower is better. Token counts below are synthetic batch sizes, not benchmark_serving concurrency. M256 TP8/TP4 rows are newly measured; other rows retain the previous qualified selection.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 121.73 103.33 -15.1%
8 128 hot 44.16 44.32 +0.4%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 232.20 181.33 -21.9%
4 256 hot 70.37 88.52 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

Complete isolated and serving parity remain unmet. Six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. New ordered-down sign planes improve the selected M256 kernels. TP8 M256 hot remains 0.9% behind; TP4 M256 hot is 25.8% behind its fresh baseline. Dense TP4 M4096 remains about 22% behind.

E393/E394 sign-plane down kernels are retained and pass 2,160 new real-weight checks. Concurrent filtered gates and direct dense gate epilogues are rejected. All attempts and failures are in the optimization inventory.

Every selected row has fresh matched MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R; IQ2R uses FP8 activations. These results do not establish original-checkpoint quality. Production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: updated isolated MoE numbers.

Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. TP8 M128 and TP4 M256 rows are newly measured.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

The tested TP8 M128 hot gap is closed: IQ2R is 1.9% faster than fresh MXFP4, with exact real-weight operator qualification. Seven of eight selected TP8 rows and five of six compact TP4 rows now beat matched MXFP4. TP8 M256 hot remains 0.9% behind; TP4 M256 hot remains 25.8% behind its latest baseline, and dense TP4 M4096 remains about 22% behind. Complete isolated and serving parity remain unmet.

The M128 gate now passes 1,080 real-weight checks and beats fresh MXFP4 on spread/hot/mixed. TP4 barrier removal passes another 1,080 real-weight checks and improves its control by 0.7–1.5%; its hot gap remains 25.8%. E402 is tracked separately and is excluded from this publication.

All selected rows pass the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: current selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. The selected numbers are unchanged from the previous update.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet.

The latest grid, LDS-layout and M32 reuse experiments produced no broad winner. Three completed experiments add 1,008 exact synthetic checks with stable bookends; the padded layout stopped at its alignment probe, and the two-word carry stopped on a nonfinite output. The inventory records each result and preserves its failure evidence. E408 diagnosis is ongoing and excluded.

All selected rows pass the unchanged 3% drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 135.83 104.72 -22.9%
8 256 hot 62.80 63.36 +0.9%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet.

Latest TP4 M256 experiment (E412), in us:

Routes MXFP4 Selected control Combined overlap
spread 224.23 179.86 188.55
hot 68.93 88.75 83.05
mixed 225.81 191.57 201.44

Overlap improves hot, but cold regressions and the remaining 20.5% hot gap keep it out of the selected table. E408–E412 add 1,917 completed checks; the inventory records all outcomes and failures. E413 TP8 transfer is running and excluded.

Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 131.03 111.48 -14.9%
8 256 hot 61.32 58.53 -4.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 603.20 738.76 +22.5%
4 4096 hot 501.69 612.41 +22.1%

All eight selected TP8 comparison rows now beat their matched MXFP4 baselines. E413 closes M256 hot with 58.53 versus 61.32 us (4.5% faster), while preserving wins on spread and mixed using one fixed policy. Five of six compact TP4 rows beat baseline; TP4 M256 hot and dense gaps remain. This is isolated operator parity within the selected TP8 coverage, not complete TP8 coverage or serving parity.

E413 TP8 M256 also passes mixed routing: 118.41 us IQ2R versus 132.92 us MXFP4 (10.9% faster). The candidate passes 315 synthetic checks plus 1,440 exact real-weight checks across three layer/rank slices, with the same compiled binary. It gives up 6.1% of the previous spread speed and 4.2% of mixed speed to close hot parity using one fixed policy. The inventory preserves that tradeoff. TP4 hot/dense and uncovered TP8 cases remain open; E415 TP4 scheduling work continues.

Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

MXFP4 versus IQ2R: selected isolated MoE numbers.

Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.

TP Tokens Routes MXFP4 µs IQ2R µs IQ2R time difference
8 32 spread 78.78 66.57 -15.5%
8 32 hot 28.56 28.09 -1.7%
8 64 spread 103.36 86.56 -16.3%
8 64 hot 40.97 34.38 -16.1%
8 128 spread 119.15 101.60 -14.7%
8 128 hot 43.35 42.53 -1.9%
8 256 spread 131.03 111.48 -14.9%
8 256 hot 61.32 58.53 -4.5%
4 64 spread 191.78 159.06 -17.1%
4 64 hot 42.53 42.02 -1.2%
4 128 spread 212.21 169.64 -20.1%
4 128 hot 60.88 59.32 -2.6%
4 256 spread 223.90 179.49 -19.8%
4 256 hot 68.93 86.72 +25.8%
4 1024 spread 290.48 300.02 +3.3%
4 1024 hot 187.38 196.84 +5.0%
4 4096 spread 594.30 730.14 +22.9%
4 4096 hot 488.88 602.89 +23.3%

All eight selected TP8 rows still beat matched MXFP4; five of six selected compact TP4 rows beat baseline. E416 improves its matched dense TP4 control by 0.52–1.44% and passes real-weight qualification, but dense parity and TP4 M256 hot remain open. No production or serving parity is claimed.

New TP4 results: E416 M4096 spread/hot is 730.14/602.89 us versus matched MXFP4 594.30/488.88 us, still 22.9%/23.3% behind. It improves the matched E386 control by 0.58%/1.44% and passes 2,160 exact real-weight checks with the same binary. E415 scalar M16 hot reaches 81.66 us versus 68.92 us MXFP4, but its cold-route tradeoff keeps E400 selected. E417/E419 M32 variants are rejected after profiling showed reduced residency. This update adds 1,554 synthetic checks; E420 continues the register-pressure experiment.

Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 M256: latest isolated MXFP4 versus IQ2R candidate.

Microseconds per complete MoE call; lower is better. Synthetic tokens are not benchmark_serving concurrency.

TP4 M256 routes MXFP4 us Scalar M16 us Parallel M32 candidate us Candidate vs MXFP4
spread 224.37 187.06 188.86 -15.8%
hot 69.00 82.58 76.06 +10.2%
mixed 225.56 200.12 206.31 -8.5%

The hot gap is now 10.2%. This candidate passes 1,440 exact real-weight checks with the same binary used for timing. The M32 residency change and parallel epilogue pass 1,386 synthetic checks across E420–E423; all timings satisfy the unchanged 3% per-arm bookend-drift limit, with zero scratch/spills.

Cold routes still favor scalar M16, so the selected 18-row coverage table remains unchanged. Synthetic MXFP4 is A4W4 requantized from IQ2R. Production integration, full-model quality and ATOM benchmark_serving acceptance remain pending. Experiments continue on a fresh Fleet node after verified runtime/fixture restoration.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 M256: latest isolated comparison. Microseconds per complete MoE call; lower is better.

TP4 M256 routes MXFP4 us M32 vector exchange us IQ2R time difference
spread 225.69 186.42 -17.4%
hot 68.98 75.57 +9.5%
mixed 226.96 200.18 -11.8%

Vector partial exchange improves the matched M32 control by 1.29% on hot routing, but the gap to MXFP4 is still 9.5%. Spread/mixed remain faster with M16; separately, grid3 improves M16 by 4.1%/3.4% on those patterns.

E426–E428 pass 1,008 exact synthetic checks. The grid3 and vector candidates pass 1,440 real-weight checks using their respective timed binaries, with native grids and zero scratch/spills verified. All timing rows satisfy the unchanged 3% drift limit. The 18 selected coverage rows remain unchanged.

Synthetic tokens are not serving concurrency; MXFP4 is A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending. Cross-collection DRAM counters are under separate investigation; these comparisons use clean latency.

@ssharma4-amd

Copy link
Copy Markdown
Author

TP4 M256: latest completed isolated comparison. Microseconds per complete MoE call; lower is better.

TP4 M256 routes MXFP4 us E428 control us E433 specialized us IQ2R time difference
spread 224.63 185.96 183.58 -18.3%
hot 69.10 75.93 73.31 +6.1%
mixed 226.83 199.24 195.78 -13.7%

E433 improves its matched M32 control by 3.46% on hot routing; the remaining MXFP4 gap is 6.1%. All 315 synthetic checks and 720 real-weight checks pass on the same timed binary, with native grids, zero scratch/spills and the unchanged 3% drift rule. The 18 selected coverage rows remain unchanged; cold routing still favors the separately measured M16 ingredient.

Rejected: E432 component stores and E435 spilling M64 reuse. The known-byte counter probe supports current TCC units but leaves the historical E426 discrepancy unresolved; no old counts were rescaled.

Synthetic tokens are not serving concurrency. MXFP4 uses A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending.

@ssharma4-amd

Copy link
Copy Markdown
Author

GLM-5.3 IQ2R vs stock MXFP4: TP8 serving results, ISL 1024

Status: at TP8 with ISL 1024 and OSL 1024, IQ2R beats stock ATOM/AITER MXFP4 at 8 of 9 concurrencies. C=4 misses the parity bar by 0.02 points, and a fix for it is included below. Every request completed, with zero failures.

Method: unmodified atom.benchmarks.benchmark_serving, same node (MI355X, TP8), in the order MXFP4 → IQ2R → MXFP4.

  • Random dataset, range ratio 0.8, 10×C prompts, 2×C warmups, seed 0, --ignore-eos, request rate inf.
  • FP8 KV cache, no prefix caching, max_num_batched_tokens 4096, max_num_seqs 256, CUDA graphs [1…256].
  • The MXFP4 column is the average of the two MXFP4 runs.
  • Parity bar: IQ2R ≥ MXFP4 average − 0.5%.
C MXFP4 avg (tok/s) IQ2R (tok/s) Δ TPOT ms (MXFP4 → IQ2R) TTFT ms (MXFP4 → IQ2R)
1 79.9 83.9 +5.06% 12.42 → 11.83 104.9 → 95.5
2 152.0 159.4 +4.83% 12.75 → 12.16 105.0 → 100.9
4 326.2 324.5 −0.52% 11.99 → 12.06 104.9 → 95.9
8 594.2 605.5 +1.89% 12.99 → 12.76 155.4 → 146.4
16 1081.8 1107.5 +2.38% 14.34 → 13.99 146.6 → 153.0
32 1686.7 1818.8 +7.84% 18.02 → 16.65 190.9 → 172.6
64 2646.1 2932.6 +10.83% 23.26 → 20.96 270.2 → 234.7
128 4088.7 4496.3 +9.97% 30.06 → 27.34 408.8 → 343.6
256 5806.4 6273.7 +8.05% 42.44 → 39.26 626.9 → 581.4

The two MXFP4 runs agree to within 0.5% at every point.

Changes in this push

  • Packed decode stack: a load-time repack of the MoE weights into a decode-friendly layout, used at M2–256.
    • Enabled by IQ2R_GLM53_PACKED_STACK=1 and IQ2R_GLM53_PACKED_SMALL=1.
    • ATOM builds it in Iq2rMoEMethod, in the companion ATOM PR.
  • Prefill: a multi-block route sort, and the E243 gate plus the E261 down+reduce kernels. The gate runs as variant 33 on the packed stack.
  • M2/M4 decode: runs the packed gate followed by the route9 down on the original weights. This is the fix for the C=4 point: in a microbenchmark, the M4 MoE takes 22.35 µs per layer against MXFP4's 24.6 µs. It has not yet been confirmed end to end.
  • E436 commit: these commits were rebased onto the E436 prefetch commit already on this branch. On the production dispatch path, the MoE output is bitwise identical to the pre-rebase build at M=1–4096, and the kernel timings are within noise.

Still open

  • C=4 end to end: the serving rerun that checks the M2/M4 fix is in progress.
  • ISL 8192 at TP8: stock MXFP4 hits a GPU memory fault in graph mode on the first 8k request. The fault is in _gluon_deepgemm_fp8_paged_mqa_logits_preshuffle, the DSA indexer kernel. The same request succeeds in eager mode, and the block tables and context lengths check out. Root cause is still being investigated; IQ2R runs fine at 8k.
  • TP4: next, at ISL 1024 and 8192.
  • MTP: after TP4, compare IQ2R and MXFP4 on throughput, TPOT and acceptance rate.

🤖 Generated with Claude Code

@ssharma4-amd ssharma4-amd changed the title Add plain GLM-5.3 IQ2R 2-bit MoE integration IQ2R: load GLM-5.3 packed MoE checkpoints with TP slicing Sep 30, 2026
ssharma4-amd and others added 2 commits September 30, 2026 18:56
Add an IQ2R quantization path for GLM-5.3 checkpoints whose routed experts
and fused shared expert are stored in AITER's glm53-packed-v1 layout.

- quant_spec: Iq2rParser reads the aiter-iq2r-overlay v2 contract. The
  modules in iq2r_modules get the IQ2R spec; every other layer keeps the
  base checkpoint's quantization (base_quantization_config).
- config: IQ2R checkpoints are eligible for online quantization of their
  non-IQ2R layers.
- moe: Iq2rMoEMethod allocates the packed expert tensors, slices each full
  checkpoint layer to this rank's intermediate shard at load time (TP4 and
  TP8 share one checkpoint), and runs aiter.iq2r_glm53.iq2r_glm53_moe_out
  after the usual expert selection. Other layouts are rejected.
- deepseek_v2: map the stacked IQ2R tensor names of GlmMoeDsaForCausalLM
  onto the FusedMoE parameters.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sparse-attention indexer passes max_model_len and
max_num_seqs * max_model_len to its custom op as Python ints, so Dynamo
bakes them into the compiled graph. A server restarted with a larger
--max-model-len reused an artifact built for a smaller one and faulted
on the first decode past the old width.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ssharma4-amd and others added 2 commits October 4, 2026 03:48
…_seqs

Config.compute_hash now keys on both fields, and this stub is built to fail
loudly when the factor list grows a field it does not model.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants