IQ2R: load GLM-5.3 packed MoE checkpoints with TP slicing - #2335
ssharma4-amd wants to merge 4 commits into
Conversation
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
|
@ssharma4-amd thanks for doing this, IQ2R would be great to have in ATOM, a couple of questions here:
|
| This tree can replace the routed FP8 experts in transformer layers 3–44 with | ||
| AITER's native-basis IQ2R format. Attention, dense MLPs, the shared experts, | ||
| and checkpoint layer 45 (MTP) retain the base model's block-FP8 configuration. | ||
| The current runtime is TP1/EP1 only. |
There was a problem hiding this comment.
Is this correct? You show results in the PR description from TP4 runs.
There was a problem hiding this comment.
Hi James, this PR is still a WIP. I need to review certain changes since we have been working both with GPT OSS 120B and GLM 5.3
|
GLM-5.3 benchmark results — 2026-09-25 23:16 UTC ATOM TP8 — 1k input / 1k output
TP8 — 8k input / 1k output
TP4: pending for both workloads. MXFP4 1k/1k C16 failed; no throughput accepted. All numeric results have zero failed requests. Random length ratio 0.8; 10×C measured requests + 2×C warmups; FP8 KV; MTP off. MXFP4 1k/1k C1–C8 uses the mean of before/after runs. Other comparisons use a later baseline on the same node and still need fresh before/after validation. |
|
Single-GPU MoE optimization update One synthetic TP rank on MI355X. Microseconds per local MoE call; lower is better. These are not
TP8: E261. TP4: E262. Both use exact index/sign repacking and interleaved down/reduction for4096 tokens. Stored weight size is unchanged. Each row averages fresh clean bookends; changing-input/route IQ2R checks are exact and native kernels have zero scratch. The goal remains unmet. Dense candidates remain isolated; serving sweeps are paused. Details and all attempted optimizations are in |
|
Single-GPU MoE results — E280 small-token candidate Synthetic TP rank on MI355X. µs per complete local MoE call; lower is better. Tokens below are not serving concurrency. Same E280 configuration in all small-token rows.
The gate now uses102 VGPRs versus129 for the earlier packed gate, restoring two-workgroup residency without spills. Both TPs have fresh clean bookends and rocprof counters. E284 passes30 shape/routing cases and240 changing-input steps with exact final BF16 and intermediate FP8 values/scales. Dense selection remains E261 (TP8) / E262 (TP4):
The goal remains unmet. The new down-decoder scout E283 failed correctness and is excluded. Candidates remain isolated; production integration and official |
|
Single-GPU MoE update: TP8 M16 now beats its fresh MXFP4 control. Microseconds per complete local MoE call; lower is better. These are isolated shape-specific experiments, not serving concurrency or production results.
E289 combines the new down schedule, M16 frontend and batched reduction. E292 improves the M4 fused-down epilogue. All rows have clean bookends; TP8 M4 hot still loses. E280 passed 32 real-capture/TP cases. The E289 combination passed 30 broader routing/shape cases with exact final and intermediate outputs, plus 45 frontend tests. TP4 M8 remains unqualified after an unchanged-MXFP4 graph/eager tolerance failure. Dense gaps and production qualification remain; serving sweeps are still paused. All attempts through E292 are documented in |
|
TP8 isolated MoE: combined IQ2R versus fresh MXFP4. Microseconds per complete local MoE call, lower is better. Tokens are not serving concurrency.
All rows have less than 1% bookend drift, exact IQ2R checks and native dispatch The combination also passed 30 synthetic TP8/TP4 cases and 32 real-capture/TP No new TP4 MXFP4 numbers in this update; the prior TP4 table is retained. |
|
TP4 isolated MoE update: fresh MXFP4 versus IQ2R. Microseconds per complete local MoE call, lower is better. These token counts are not serving concurrency. M16 uses the combined small-token path; 1024/4096 use E308 N512 dense down.
Dense IQ2R improves 0.7–3.5% over E262, but still trails MXFP4. All table rows The dense candidate also passed 39 shape/routing cases and 936 exact checks, All attempts through E311 are documented in |
|
M4 isolated MoE update: MXFP4 versus IQ2R. Microseconds per complete MoE call; lower is better. Tokens here are not serving concurrency.
TP8 now reaches hot-route parity; its 0.5% difference is too small to call a Changes: wave sorting and paired FP8 conversion in the frontend; TP8 reuses Larger-token/dense gaps remain; overall serving parity is not achieved. |
|
New isolated MoE results: MXFP4 versus IQ2R, TP8. Microseconds per complete call; lower is better. Token counts are not serving concurrency.
Three cases now beat MXFP4; M128 hot remains 10.8% behind. Fresh clean bookends, Changes: independent output waves skip empty down subtiles and reuse inputs; The overall goal is still open. TP4 medium/dense gaps and broader qualification |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The fixed compact-gate policy beats MXFP4 in six of eight stable M64/M128 synthetic cases. TP8 M128 hot remains 1.4% slower; TP4 M64 hot remains 0.3% slower. These small remaining gaps are still open. TP4 r2 repeats the frozen selected candidate after the original hot rows exceeded the unchanged drift limit. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The selected isolated kernels beat MXFP4 on six of eight TP8 rows and three of eight TP4 rows shown. TP8 hot gaps are down to 0.6% at 128 tokens and 4.5% at 256; TP4 M64 hot remains 2.0% behind and dense cases remain 4.5–27.4% behind. The full performance goal is not achieved. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The selected isolated kernels beat MXFP4 on six of eight TP8 rows and all four compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 3.1% at 256. The combined dense TP4 kernel remains 3.3–23.5% behind. The full performance goal is not achieved. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The selected kernels beat MXFP4 on six of the eight reported TP8 rows and five of six compact TP4 rows. TP8 hot gaps are 0.4% at 128 tokens and 2.3% at 256. Newly qualified TP4 M256 spread is 19.9% faster, but hot is 29.7% slower. The four reported dense TP4 rows remain 3.3–23.5% slower. The full isolated and serving goal is not achieved. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. TP8 M256 token-major routing improves its control by 0.8–1.7%, but hot remains 4.3% behind fresh MXFP4. TP4 M256 hot remains 29.7% behind, and the reported dense TP4 rows remain 3.3–23.5% behind. N1024 down provides a small M4096 improvement. The full isolated and serving goal remains unmet. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
The reported selected kernels beat MXFP4 on six of eight TP8 rows and five of six compact TP4 rows. New gate scheduling reduces the reported TP8 M256 hot gap to 1.5% and TP4 M256 hot gap to 23.4%. Dense TP4 remains 3.3–23.5% behind. Latest compact and dense candidates pass 4,320 real-weight operator checks. Complete isolated and serving parity remain unmet. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
New isolated MoE comparison: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. Tokens are not serving concurrency.
Complete isolated and serving parity remain unmet. The selected table still has six of eight TP8 rows and five of six compact TP4 rows faster than matched MXFP4. Direct codebook byte addresses improve both dense TP4 M4096 routes by 0.3–0.6% against their current control, leaving 22.5%/22.1% MXFP4 gaps. Extra activation lookahead is rejected; the compact TP8 address change is unselected. Fresh same-GPU MXFP4 bookends; every row passes the existing drift limit. The full serving goal remains open. These are isolated experiments; production |
|
Selected isolated MoE results: MXFP4 versus IQ2R. Microseconds per complete call; lower is better. These synthetic token counts are not benchmark_serving concurrency. The selected numbers are unchanged by the latest experiments.
Complete isolated and serving parity remain unmet. The selected comparison is unchanged: six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. Dense TP4 M4096 remains 22.5%/22.1% slower. E388–E390 do not improve the selected policy. E388–E390 are rejected: wider MFMA, paired weight loads and larger M reuse did not beat the selected kernels. E391 passes all 1,080 real-weight checks for the retained TP4 M4096 address-calculation change. The optimization inventory includes all attempts and preserved failures. Every selected row has matched clean MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. MXFP4 is A4W4 requantized from the synthetic IQ2R fixture; IQ2R uses FP8 activations. This is operator performance evidence, not original-checkpoint quality. Production integration, whole-model quality and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: updated isolated MoE results. Microseconds per complete call; lower is better. Token counts below are synthetic batch sizes, not benchmark_serving concurrency. M256 TP8/TP4 rows are newly measured; other rows retain the previous qualified selection.
Complete isolated and serving parity remain unmet. Six of eight TP8 rows and five of six compact TP4 rows beat matched MXFP4. New ordered-down sign planes improve the selected M256 kernels. TP8 M256 hot remains 0.9% behind; TP4 M256 hot is 25.8% behind its fresh baseline. Dense TP4 M4096 remains about 22% behind. E393/E394 sign-plane down kernels are retained and pass 2,160 new real-weight checks. Concurrent filtered gates and direct dense gate epilogues are rejected. All attempts and failures are in the optimization inventory. Every selected row has fresh matched MXFP4 bookends within the unchanged 3% drift limit, exact IQ2R checks, verified native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R; IQ2R uses FP8 activations. These results do not establish original-checkpoint quality. Production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: updated isolated MoE numbers. Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. TP8 M128 and TP4 M256 rows are newly measured.
The tested TP8 M128 hot gap is closed: IQ2R is 1.9% faster than fresh MXFP4, with exact real-weight operator qualification. Seven of eight selected TP8 rows and five of six compact TP4 rows now beat matched MXFP4. TP8 M256 hot remains 0.9% behind; TP4 M256 hot remains 25.8% behind its latest baseline, and dense TP4 M4096 remains about 22% behind. Complete isolated and serving parity remain unmet. The M128 gate now passes 1,080 real-weight checks and beats fresh MXFP4 on spread/hot/mixed. TP4 barrier removal passes another 1,080 real-weight checks and improves its control by 0.7–1.5%; its hot gap remains 25.8%. E402 is tracked separately and is excluded from this publication. All selected rows pass the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: current selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. These are synthetic token batches, not benchmark_serving concurrency. The selected numbers are unchanged from the previous update.
The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet. The latest grid, LDS-layout and M32 reuse experiments produced no broad winner. Three completed experiments add 1,008 exact synthetic checks with stable bookends; the padded layout stopped at its alignment probe, and the two-word carry stopped on a nonfinite output. The inventory records each result and preserves its failure evidence. E408 diagnosis is ongoing and excluded. All selected rows pass the unchanged 3% drift rule, exact IQ2R checks, native dispatch and zero scratch. MXFP4 is A4W4 requantized from IQ2R for this fixture. Original-checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.
The selected comparison is unchanged: seven of eight selected TP8 rows and five of six compact TP4 rows beat matched MXFP4. TP8 M256 hot remains 0.9% behind, TP4 M256 hot 25.8% behind, and dense TP4 M4096 about 22% behind. Complete isolated and serving parity remain unmet. Latest TP4 M256 experiment (E412), in us:
Overlap improves hot, but cold regressions and the remaining 20.5% hot gap keep it out of the selected table. E408–E412 add 1,917 completed checks; the inventory records all outcomes and failures. E413 TP8 transfer is running and excluded. Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.
All eight selected TP8 comparison rows now beat their matched MXFP4 baselines. E413 closes M256 hot with 58.53 versus 61.32 us (4.5% faster), while preserving wins on spread and mixed using one fixed policy. Five of six compact TP4 rows beat baseline; TP4 M256 hot and dense gaps remain. This is isolated operator parity within the selected TP8 coverage, not complete TP8 coverage or serving parity. E413 TP8 M256 also passes mixed routing: 118.41 us IQ2R versus 132.92 us MXFP4 (10.9% faster). The candidate passes 315 synthetic checks plus 1,440 exact real-weight checks across three layer/rank slices, with the same compiled binary. It gives up 6.1% of the previous spread speed and 4.2% of mixed speed to close hot parity using one fixed policy. The inventory preserves that tradeoff. TP4 hot/dense and uncovered TP8 cases remain open; E415 TP4 scheduling work continues. Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
MXFP4 versus IQ2R: selected isolated MoE numbers. Microseconds per complete MoE call; lower is better. Synthetic token batches, not benchmark_serving concurrency.
All eight selected TP8 rows still beat matched MXFP4; five of six selected compact TP4 rows beat baseline. E416 improves its matched dense TP4 control by 0.52–1.44% and passes real-weight qualification, but dense parity and TP4 M256 hot remain open. No production or serving parity is claimed. New TP4 results: E416 M4096 spread/hot is 730.14/602.89 us versus matched MXFP4 594.30/488.88 us, still 22.9%/23.3% behind. It improves the matched E386 control by 0.58%/1.44% and passes 2,160 exact real-weight checks with the same binary. E415 scalar M16 hot reaches 81.66 us versus 68.92 us MXFP4, but its cold-route tradeoff keeps E400 selected. E417/E419 M32 variants are rejected after profiling showed reduced residency. This update adds 1,554 synthetic checks; E420 continues the register-pressure experiment. Every selected row passes the unchanged 3% bookend-drift rule, exact IQ2R checks, native dispatch and zero scratch. Synthetic MXFP4 is A4W4 requantized from IQ2R. Real checkpoint quality, production integration and final ATOM benchmark_serving acceptance remain pending. |
|
TP4 M256: latest isolated MXFP4 versus IQ2R candidate. Microseconds per complete MoE call; lower is better. Synthetic tokens are not benchmark_serving concurrency.
The hot gap is now 10.2%. This candidate passes 1,440 exact real-weight checks with the same binary used for timing. The M32 residency change and parallel epilogue pass 1,386 synthetic checks across E420–E423; all timings satisfy the unchanged 3% per-arm bookend-drift limit, with zero scratch/spills. Cold routes still favor scalar M16, so the selected 18-row coverage table remains unchanged. Synthetic MXFP4 is A4W4 requantized from IQ2R. Production integration, full-model quality and ATOM benchmark_serving acceptance remain pending. Experiments continue on a fresh Fleet node after verified runtime/fixture restoration. |
|
TP4 M256: latest isolated comparison. Microseconds per complete MoE call; lower is better.
Vector partial exchange improves the matched M32 control by 1.29% on hot routing, but the gap to MXFP4 is still 9.5%. Spread/mixed remain faster with M16; separately, grid3 improves M16 by 4.1%/3.4% on those patterns. E426–E428 pass 1,008 exact synthetic checks. The grid3 and vector candidates pass 1,440 real-weight checks using their respective timed binaries, with native grids and zero scratch/spills verified. All timing rows satisfy the unchanged 3% drift limit. The 18 selected coverage rows remain unchanged. Synthetic tokens are not serving concurrency; MXFP4 is A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending. Cross-collection DRAM counters are under separate investigation; these comparisons use clean latency. |
|
TP4 M256: latest completed isolated comparison. Microseconds per complete MoE call; lower is better.
E433 improves its matched M32 control by 3.46% on hot routing; the remaining MXFP4 gap is 6.1%. All 315 synthetic checks and 720 real-weight checks pass on the same timed binary, with native grids, zero scratch/spills and the unchanged 3% drift rule. The 18 selected coverage rows remain unchanged; cold routing still favors the separately measured M16 ingredient. Rejected: E432 component stores and E435 spilling M64 reuse. The known-byte counter probe supports current TCC units but leaves the historical E426 discrepancy unresolved; no old counts were rescaled. Synthetic tokens are not serving concurrency. MXFP4 uses A4W4 requantized from IQ2R. Production integration, model quality and ATOM benchmark_serving acceptance remain pending. |
GLM-5.3 IQ2R vs stock MXFP4: TP8 serving results, ISL 1024Status: at TP8 with ISL 1024 and OSL 1024, IQ2R beats stock ATOM/AITER MXFP4 at 8 of 9 concurrencies. C=4 misses the parity bar by 0.02 points, and a fix for it is included below. Every request completed, with zero failures. Method: unmodified
The two MXFP4 runs agree to within 0.5% at every point. Changes in this push
Still open
🤖 Generated with Claude Code |
ffa498c to
0b3a15f
Compare
Add an IQ2R quantization path for GLM-5.3 checkpoints whose routed experts and fused shared expert are stored in AITER's glm53-packed-v1 layout. - quant_spec: Iq2rParser reads the aiter-iq2r-overlay v2 contract. The modules in iq2r_modules get the IQ2R spec; every other layer keeps the base checkpoint's quantization (base_quantization_config). - config: IQ2R checkpoints are eligible for online quantization of their non-IQ2R layers. - moe: Iq2rMoEMethod allocates the packed expert tensors, slices each full checkpoint layer to this rank's intermediate shard at load time (TP4 and TP8 share one checkpoint), and runs aiter.iq2r_glm53.iq2r_glm53_moe_out after the usual expert selection. Other layouts are rejected. - deepseek_v2: map the stacked IQ2R tensor names of GlmMoeDsaForCausalLM onto the FusedMoE parameters. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sparse-attention indexer passes max_model_len and max_num_seqs * max_model_len to its custom op as Python ints, so Dynamo bakes them into the compiled graph. A server restarted with a larger --max-model-len reused an artifact built for a smaller one and faulted on the first decode past the old width. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0b3a15f to
3445cc1
Compare
…_seqs Config.compute_hash now keys on both fields, and this stub is built to fail loudly when the factor list grows a field it does not model. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
3445cc1 to
337460e
Compare
Summary
quant_spec:Iq2rParserreads theaiter-iq2r-overlayv2 contract.Modules listed in
iq2r_modulesget the IQ2R spec. Every other layer keepsthe base checkpoint's quantization (
base_quantization_config).Iq2rMoEMethodrecognizesquantization_config.iq2r_layout = "glm53-packed-v1".time; the same checkpoint serves TP4 and TP8.
aiter.iq2r_glm53.iq2r_glm53_moe_out(SiLU activation).deepseek_v2.py: weight mappings for the packed tensors.config.py:iq2raccepts--online-quant-config, so the non-expertlayers can be PTPC FP8.
Launch:
With MTP, the draft layer (layer 78) keeps checkpoint experts, which online
quantization turns into MXFP4. Only layers 3–77 are excluded:
Test plan
tests/test_iq2r_moe_method.py11 passed: overlay v2 parsing, contractand layout rejection, exact checkpoint shapes, TP4
create_weights,rank-shard slicing at (4,3) and (8,5), and the packed apply. Without an
IQ2R-capable aiter the module skips.
tests/test_config_compile_hash.py3 passed. These tests are new andfail without the fix.
tests/test_indexer_cp_gate.py23 passed.checkpoint, greedy: coherent, correct answers on all prompts.
vs MXFP4 97.19 and FP8 97.57. Details are in [HIP] IQ2R: GLM-5.3 packed MoE path for TP4/TP8 on gfx950 aiter#5728.
Depends on ROCm/aiter#5728.
Results (MXFP4 vs IQ2R, TP4 and TP8)
ATOM
benchmark_serving, MI355X, OSL 1024, 10×C prompts, default launch(no IQ2R env vars). MXFP4 = stock ATOM with the experts quantized online to
MXFP4. The non-MTP tables were measured before the aiter M ≤ 1024 extension
and retune; the MTP tables include it. Output tok/s:
TP4, ISL 1024 (IQ2R peak weights 58.6 GB/GPU, 198.7k KV blocks):
TP4, ISL 8192:
TP8, ISL 1024 (IQ2R peak weights 32.7 GB/GPU, 230.5k KV blocks):
TP8 at ISL 8192 has not been rerun on this build yet.
MTP results
GLM-5.3's MTP layer (checkpoint layer 78) drafts 3 tokens per step
(
--method mtp --num-speculative-tokens 3). The target then verifies1 + 3 tokens per sequence, so the routed MoE runs at M = 4 × C. The MTP gain at
concurrency C therefore follows the non-MTP gain at 4C. Both arms force the
same acceptance length (
--spec-decode-acceptance-length 2.8), so they do thesame work per step and only kernel speed differs.
Quantization parity matters for the draft layer: the MXFP4 arm quantizes the
layer-78 experts to MXFP4. The IQ2R arm must do the same, by excluding only
layers 3–77 from online quantization:
TP8, ISL 1024, OSL 1024, FP8 KV cache, same
benchmark_servingsettings asabove. Output tok/s:
TP4, ISL 1024, OSL 1024, BF16 KV cache, on a second MI355X machine. Output tok/s:
These MTP tables were measured on a development build with the same GLM-5.3
kernels and tuned CSV; its ATOM IQ2R method served only the packed layout. A
TP4 C8 check with this PR's loader and aiter#5728 (self-contained checkpoint,
second pass) gives 1009.1 tok/s, +11.3% over MXFP4; the development build gave
1025.0 to 1044.9 on that row.
At TP8 C128 and C256 the verify batch is 512 or 1024 tokens. That runs on the
prefill kernels, where IQ2R is level with MXFP4 or slightly slower (TP8 MoE at
M = 1024: 212 µs vs 204 µs), so the gain shrinks there.
Included fix
config: include max_model_len and max_num_seqs in the compile cache hash
The sparse-attention indexer passes
max_model_lenandmax_num_seqs * max_model_lento its custom op as Python ints, so Dynamobakes them into the compiled graph. The failure:
--max-model-lenreused an artifactbuilt for a smaller one.
width.
Reproduced on GLM-5.3 TP4: an ISL 1024 run followed by ISL 8192.
The fix adds both fields to
Config.compute_hash. With it, the samesequence completes; see the TP4 ISL 8192 sweep above.
Tests:
tests/test_config_compile_hash.py(fails without the fix). Theindexer CP test stub in
tests/test_indexer_cp_gate.pynow has the two newfields.
🤖 Generated with Claude Code