Skip to content

perf(cpu-ep): packed-panel 6x16 microkernel for gemm_bt MoE prefill - #1364

Merged
justinchuby merged 1 commit into
mainfrom
squad/leon-p23-moe-packed-gemm
Aug 19, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/leon-p23-moe-packed-gemm

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 19, 2026 •

Copy link
Copy Markdown
Owner

What

A packed-panel 6x16 microkernel for gemm_bt, the kernel that computes
C = A·Btᵀ directly on the [out][in]-major MoE expert weights (#1241). It is
entered only from the single-threaded driver on prefill-shaped row extents;
decode, small-token MoE and every multi-threaded shape keep the existing
tile_4x2 path unchanged.

Why

MoE expert GEMMs were the largest remaining single-thread loss on the
production-native matrix — §31.5/§31.8 of the CPU-EP ledger put MoE prefill at
1.22–1.56x ORT at t=512 and attributed it to "microkernel efficiency
against MLAS's packed panels and wider register tiles". That attribution is
right, and it is fixable without handing anything to ORT's CPU EP.

tile_4x2 keeps k-partials in the SIMD lanes. Two consequences: every
output cell ends in a horizontal reduction, and the b side of the inner loop
is strided by k — six loads per eight FMAs.

How

tile_4x2 (unchanged, still the default) micro_6x16 (new)
lanes hold k-partials c columns
horizontal reductions one per output cell none
per k step 6 loads / 8 FMAs = 1.33 2 contiguous loads + 6 broadcasts / 12 FMAs = 1.5
b access strided by k two sequential cache lines
  • pack_panel transposes one column panel of bt into micro-panel-major order
    (packed[s*k*PNR + p*PNR + t]). This is the only transpose in the path, it is
    bounded by the panel rather than the whole expert bank, and every row band in
    the sweep reuses it.
  • block_packed drives column panel → row band → micro-panel. Both remainders
    (rows % 6, n % 16) fall through to the existing full-k dot, which is
    the same code the unpacked path already uses for its edges.
  • Single-threaded gemm_bt hands the whole row extent to one block when the
    packed kernel is selected, so the pack is amortised once rather than once per
    64-row MAX_MC block.

The gate is structural, not tuned

PACK_MIN_ROWS = PMR * PACK_MIN_BANDS = 72, deliberately above MAX_MC (64),
pinned by a const assertion:

const _: () = assert!(
    PACK_MIN_ROWS > super::MAX_MC,
    "the packed kernel must stay out of the multi-threaded driver's row \
     blocks, which are capped at MAX_MC and would each re-pack the panel"
);

The multi-threaded driver slices m into 2*threads blocks capped at MAX_MC,
and each block re-packs the same panel. Keeping the bar above that cap makes
"the panel is packed once per sweep" a property of the code rather than of the
shape, and means multi-threaded behaviour is byte-identical to main for
every shape
. The assertion stops a future MAX_MC bump from silently undoing
that.

Why I will not make a multi-threaded claim

I tried to. It does not survive its own control. A 2-thread A/B of two binaries
that I traced to be on the identical code path (every row block rows=56..64,
packed=false) still came out ~40% apart on the median. Multi-thread
measurement on this shared 32-core box is not usable at this effect size, in
either direction, so this PR takes the position that costs nothing: MT is left
exactly as it was. Sharing one pack across threads (or making
gemm_bt_col_parallel the packed path, since each of its tasks already owns all
m rows) is the natural follow-up, and it needs a quiet machine.

Evidence

Both binaries built --no-default-features --features bench-native,cuda-13000,
differing only in matmul.rs. The arm is proved by the linker, not the flag:
nm -C <bin> | grep -ci mlas = 0 on both. Arms interleaved innermost in one
driver process, arm order alternating per rep, thread-matched
--native-threads 1 --ort-intra-threads 1. Metric is native_min / ort_min
(lower is better; <1.000 beats ORT).

Single thread, t=512 — the shapes that reach the new kernel

Final gate, shipped binary (6 reps):

model before after Δ min Δ median parity
moe_phi35moe_h2048_i6400_e4_t512 1.356 1.009 −25.61% −16.50% PASS
moe_qwen3moe_h2048_i768_e16_t512 1.297 1.101 −15.12% −15.92% PASS
moe_mixtral_h1024_i3584_e8_t512 1.226 1.173 −4.27% −2.85% PASS

Four independent passes, Δ on the min ratio:

model pass 1 pass 2 pass 3 pass 4 (shipped)
qwen3moe t512 −16.7% −23.4% −19.6% −15.1%
phi35moe t512 −16.9% −24.9% −27.6% −25.6%
mixtral t512 −1.3% −6.6% −3.0% −4.3%

Phi-3.5-MoE prefill crosses 1.000 in the shipped pass and Qwen3-MoE lands at
1.10
, from a 1.30–1.54 loss. Mixtral moves least, and I believe that is real
rather than measurement: its expert bank is ~352 MB of fp32, so that shape is
closer to DRAM-bound than microkernel-bound and a better inner loop has less to
win.

t=1 / t=32 — unchanged by construction, and verified

Traced with a temporary instrumented build: every gemm_bt_block at t=32 has
rows in 14–19, far below the gate, so packed=false and the executed code
is the unpacked path. Their run-to-run spread (±10% on min, ±3% on median, sign
flipping between passes) is therefore a direct read of this box's noise
floor
. At t=512 the traced extents are 74–365 rows, all packed.

Correctness

  • cargo test -p onnx-runtime-ep-cpu --lib → 1434 passed, 0 failed, 17 ignored.
  • New test gemm_bt_packed_kernel_matches_naive_single_threaded: 10 shapes
    against naive_bt on a 1-thread pool, covering both remainders, k tails, the
    shortest admitted k, and multi-column-panel sweeps. It carries a
    non-vacuity assertion — if a policy change stops these shapes from reaching
    the packed kernel, the test fails loudly rather than silently measuring the old
    path.
  • Falsified twice: swapping the two stored halves of the micro-tile, and
    corrupting the pack layout to dst[t*k+p]. Each mutation fails only the new
    test; the other five gemm_bt tests stay green.
  • Clippy clean, cargo fmt --check clean, cargo check --features mlas compiles.
  • Parity PASS in every benchmark cell above.

Review

Reviewed by Opus (claude-opus-4.8), verdict APPROVE-WITH-NITS, no blocking
findings. It independently re-ran the packed test with 7 additional adversarial
shapes (the nc-floor branch at k>8192, single-micro-panel wide sweeps,
whole-m widening at m=1000, maximal row/column remainders) under
AddressSanitizer
— all clean — and verified the panel-index bounds, the
jend-j0 ≡ 0 (mod PNR) contract, remainder exhaustiveness/disjointness, and that
no consumer of gemm_bt depends on its accumulation order. Its one substantive
nit — that the original rows >= 48 gate did not keep large-m multi-threaded
blocks out of the packed path, contrary to the comment — is what produced the
structural > MAX_MC gate and the const assertion above.

Scope

  • No ORT CPU fallback, no deferral — assigned == executed is unchanged.
  • Multi-threaded blocking policy untouched; MT is on the identical code path.
  • tile_4x2, block, dot, gemm_bt_block_scalar and the mlas arm untouched.

The MoE expert GEMMs were the largest remaining single-thread loss against
ORT/MLAS on the production-native matrix (1.21-1.54x at t=512). The cause
was microkernel shape, not blocking: `tile_4x2` keeps k-partials in the
SIMD lanes, so every output cell ends in a horizontal reduction and the
inner loop issues 6 loads per 8 FMAs, with the `b` side strided by `k`.

Add a packed-panel path selected only for prefill-shaped row extents:

* `pack_panel` transposes one column panel of `bt` into micro-panel-major
  order once, bounded by the panel rather than the whole expert bank, and
  every row band in the sweep reuses it.
* `micro_6x16` holds twelve accumulators of *c columns*, so there is no
  horizontal reduction at all and each k step is two contiguous vector
  loads plus six broadcasts against twelve FMAs (1.5 FMA/load).
* `block_packed` drives column panel -> row band -> micro-panel; both
  remainders fall to the existing full-k `dot`, unchanged.
* Single-threaded `gemm_bt` now hands the whole row extent to one block
  when the packed kernel is selected, so the pack is amortised once
  instead of once per 64-row `MAX_MC` block.

The gate is twelve full micro-tile bands (`rows >= 72`), which is above
`MAX_MC`. That is structural, not tuned: the multi-threaded driver slices
`m` into `2*threads` blocks capped at `MAX_MC`, and each such block would
re-pack the same panel, so keeping the bar above the cap makes "one pack
per sweep" a property of the code rather than of the shape. A `const`
assertion pins the relationship so a future `MAX_MC` bump cannot silently
let the multi-threaded driver in.

Multi-threaded behaviour is therefore identical to before this commit for
every shape, which is the only honest position available: on the shared
host this was developed on, a 2-thread A/B of two binaries *traced to be
on the identical code path* (every row block `rows=56..64`, all unpacked)
still differed by ~40% on the median. No multi-thread claim is made in
either direction. Sharing one pack across threads is the follow-up.

Production-native single-thread A/B (both binaries `--no-default-features
--features bench-native,cuda-13000`, zero MLAS symbols by `nm`, arms
interleaved innermost with alternating order, `--native-threads 1
--ort-intra-threads 1`, metric native_min/ort_min, four passes):

  moe_qwen3moe_h2048_i768_e16_t512   1.297 -> 1.101   -15.1%
  moe_phi35moe_h2048_i6400_e4_t512   1.356 -> 1.009   -25.6%
  moe_mixtral_h1024_i3584_e8_t512    1.226 -> 1.173    -4.3%

reproduced across four independent passes (qwen -16.7/-23.4/-19.6/-15.1%,
phi35 -16.9/-24.9/-27.6/-25.6%, mixtral -1.3/-6.6/-3.0/-4.3%), parity PASS
in every cell. The t=1 and t=32 rows are provably on the identical code
path (traced row extents 14-19, below the gate) and their spread is the
host noise floor.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the squad/leon-p23-moe-packed-gemm branch from 0777291 to c37b0ea Compare August 19, 2026 02:33
@justinchuby
justinchuby merged commit 307c727 into main Aug 19, 2026
9 checks passed
@justinchuby
justinchuby deleted the squad/leon-p23-moe-packed-gemm branch August 19, 2026 02:38
justinchuby added a commit that referenced this pull request Aug 19, 2026
… that failed (#1373)

Documentation only — records Phase 17 in the CPU-EP ledger.

## What lands here

* **§37.1–37.2 — #1245**, the pure-Horner `exp8`. Full 7-model A/B
across two
independent passes, plus the accuracy contract: 1 ULP over
**1,118,743,631**
`f32` values by exhaustive sweep, and the story of the sweep that first
reported `max_ulp = 8388646` because it included the deliberate
flush-to-zero
  threshold.
* **§37.3 — #1364**, the packed-panel `6x16` `gemm_bt` microkernel. Four
independent passes. Phi-3.5-MoE prefill reaches **1.009** (parity with
ORT)
  from 1.356; Qwen3-MoE **1.101** from 1.297.
* **§37.5** — the unmerged two-pass online softmax, kept as a negative
result
(2.4× worse) with the reason it loses: §29 already removed the copy that
would
  have made the extra pass expensive.

## §37.4 is the part that matters

Choosing the packed kernel's row gate produced two negative results, and
the
second one is a methodology finding, not a perf finding.

A four-band gate looked fine single-threaded and lost ~29% on the
8-thread
Phi-3.5-MoE median. Raising it to eight bands removed that, but then the
2-thread
cells showed +6% to +11% — which reads as a real multi-threaded
regression.

It is not usable data. After raising the gate above `MAX_MC`, the two
binaries
are **traced to be on the identical code path at two threads** (every
row block
`rows=56..64`, `packed=false`, verified with an instrumented build) —
and they
still measured **~40% apart on the median**, in the favourable direction
for the
arm that had just been slower.

That is a worse null arm than §36.3's: ~40% at **two** threads on a code
path
proved identical, against §36.3's ±20–30% at thirty-two. Low thread
counts are
not the quiet regime on a shared box; they are the regime where one
noisy
neighbour is a larger fraction of what is running.

The consequence is stated plainly in the section: this document now has
two
independent measurements of its own instrument disagreeing with itself
by
20–40%, and every cross-arm claim in it that rests on a single pass
without a
null control should be read as provisional.

## Scope

No code. No behaviour change. Both PRs being documented are already
merged
(#1245 as `512885a7b`, #1364 as `307c7277d`).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
…h it (#1380)

## The problem

Two phases in a row reported a delta and then had to reconstruct, *after
the
fact*, whether the instrument could resolve it — §36.3 from cells the
change
could not reach, §37.4 from a code path traced to be identical. That
reconstruction is an argument, and it arrives too late to change what
was
measured.

## The change

`ab.py --null-control` adds a third arm that is **the first arm's own
binary,
with the first arm's `--arm-env`, under a second name**, interleaved and
order-alternated exactly like the real arms. It cannot measure the
change. Its
delta is the host's noise floor for that cell, *in that invocation*.

The driver also gained a deltas table (it previously printed only
per-arm
medians, leaving the subtraction to the reader):

```
=== deltas vs 'before' (median ratio; negative = arm is faster) ===
The 'null' column is the same binary as 'before'. Its delta is this host's noise
floor for the cell, so a real delta no larger than it is not a result.
moe_phi35moe_h2048_i6400_e4_t512 t=1    after:  -25.48%  > noise (0.19%)  null:   -0.19%
moe_mixtral_h1024_i3584_e8_t512  t=1    after:   -2.38%  > noise (0.63%)  null:   -0.63%
```

Anything inside the control prints `WITHIN NOISE` instead of a number
that looks
publishable.

## What it measured (§38)

Same-binary control arm, `|Δ|` of the median native/ort ratio against
the
identical binary:

| cell | 1 thread | 2 threads | 8 threads |
| --- | --- | --- | --- |
| `moe_mixtral_h1024_i3584_e8_t512` | 0.63% | 0.22% | **19.54%** |
| `moe_qwen3moe_h2048_i768_e16_t512` | 4.75% | 3.97% | **12.68%** |
| `moe_phi35moe_h2048_i6400_e4_t512` | 0.19% | 4.28% | 3.03% |
| `sm_bert_b8_s128` | 0.21% | — | — |
| `sm_decode_h32_kv8192` | 1.32% | — | — |
| `sm_prefill_h32_s512` | 0.22% | — | — |

**Single-thread cells are tight** — under 5%, mostly under 1.5% — so the
single-thread ratios this ledger publishes are resolvable at the effect
sizes it
claims. **Eight-thread cells are not**: a 12–20% floor retires the 5–15%
movements several earlier phases tabulated at `t=8`, exactly as §36.3
warned.

Both phase-17 merges are re-measured against their own control, in the
same
invocation — the first self-controlled cells in the document:

| change | cell (1 thread) | Δ median ratio | floor | verdict |
| --- | --- | --- | --- | --- |
| #1364 packed panel | `moe_phi35moe…t512` | **−25.48%** | 0.19% | 134×
the floor |
| #1364 packed panel | `moe_qwen3moe…t512` | **−21.56%** | 4.75% | 4.5×
the floor |
| #1364 packed panel | `moe_mixtral…t512` | −2.38% | 0.63% | 3.8× the
floor |
| #1245 Horner `exp8` | `sm_bert_b8_s128` | **−10.28%** | 0.21% | 49×
the floor |
| #1245 Horner `exp8` | `sm_decode_h32_kv8192` | **−11.73%** | 1.32% |
8.9× the floor |
| #1245 Horner `exp8` | `sm_prefill_h32_s512` | **−10.33%** | 0.22% |
47× the floor |

## It corrects my own §37.4

§37.4 reported ~40% between two distinct binaries traced to identical
code paths
at two threads, and read that as the two-thread noise floor. **That
reading was
wrong.** The same-binary control puts the two-thread floor at
0.22–4.28%, an
order of magnitude tighter, and the absolute ratios differ between the
sessions
(mixtral `t=2` reads 1.35 here, 2.23 there, same fixture, same thread
count).

The 40% was a *session*: a load episode long enough to outlast the arm
alternation. That is a more useful finding than a floor, because it is
not
something interleaving can fix — a whole invocation can be poisoned, and
only a
control arm inside that invocation exposes it. §38.3 states the narrower
rule
that follows: run the control in every invocation whose result will be
published, and discard the **invocation**, not the cell, when the
control moves
more than the effect.

## Scope

* Additive: without `--null-control` the driver behaves as before,
except that
  multi-arm runs now also print a deltas table.
* `ruff check` clean. (`ruff format` does not pass on this file before
or after
this change, so it is left alone rather than reformatted in a
perf-tooling
  PR.)
* No Rust, no behaviour change in the EP.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
Main already carries a section 39 (#1364's harness follow-up, merged while
this branch was open), so this appended one collided with it. Renumber to
40/Phase 20 rather than merging two sections under the same number.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown

✅ Benchmarks — No Regression

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
✅ matmul/large_generic_f32_threads=1/32x1024x1024 8.61 ms 9.43 ms +9.6%
✅ gather/large_f32_threads=1-internal/131072 23.15 µs 25.29 µs +9.3%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.33 ms 5.73 ms +7.4%
✅ gather/small_f32_threads=1-internal/4096 625.4 ns 656.3 ns +4.9%
✅ matmul/small_generic_f16_threads=8/1x256x256 26.92 µs 28.21 µs +4.8%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 360.49 µs 377.03 µs +4.6%
✅ matmul/small_generic_f16_threads=1/1x256x256 28.14 µs 29.27 µs +4.0%
✅ gather/large_f16_threads=1-internal/131072 11.83 µs 12.22 µs +3.2%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 494.22 µs 506.54 µs +2.5%
✅ gather/large_bf16_threads=1-internal/131072 10.39 µs 10.65 µs +2.5%
✅ sampling_latency/top_p_per_token 350.49 µs 358.77 µs +2.4%
✅ kv_cache/alloc_dealloc_pages 36.43 µs 36.95 µs +1.4%
✅ matmul/small_generic_bf16_threads=1/1x256x256 28.10 µs 28.49 µs +1.4%
✅ tokenization/decode_tokens_per_second 5.73 ms 5.80 ms +1.2%
✅ matmul/medium_generic_f16_threads=8/32x512x512 27.80 µs 28.04 µs +0.8%
✅ add/small_bf16_threads=1-internal/1024 409.7 ns 413.2 ns +0.8%
✅ sampling_latency/min_p_per_token 191.81 µs 193.17 µs +0.7%
✅ gather/medium_f16_threads=1-internal/32768 2.23 µs 2.24 µs +0.4%
✅ matmul/small_generic_f32_threads=1/1x256x256 34.15 µs 34.28 µs +0.4%
✅ gather/small_bf16_threads=1-internal/4096 441.0 ns 441.5 ns +0.1%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 487.93 µs 488.09 µs +0.0%
✅ reduce_mean/large_f32_threads=1-internal/262144 911.09 µs 911.13 µs +0.0%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.68 ms 3.67 ms -0.1%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.15 ms 2.15 ms -0.3%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 73.45 µs 73.24 µs -0.3%
✅ gather/medium_bf16_threads=1-internal/32768 2.21 µs 2.20 µs -0.5%
✅ sampling_latency/greedy_per_token 3.01 µs 2.99 µs -0.5%
✅ grammar_masking/llguidance_compute_mask/32 69.47 µs 69.04 µs -0.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 1.98 ms 1.97 ms -0.7%
✅ logit_processing/seven_processor_chain_per_step 298.52 µs 296.37 µs -0.7%
✅ add/medium_bf16_threads=1-internal/262144 97.77 µs 96.73 µs -1.1%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 79.16 µs 78.20 µs -1.2%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.31 ms 3.27 ms -1.2%
✅ add/large_f16_threads=1-internal/4194304 1.56 ms 1.54 ms -1.3%
✅ qwen3_sampling_processors/top_k_partial_selection 135.40 µs 133.02 µs -1.8%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.08 µs 13.82 µs -1.8%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.91 ms 1.88 ms -1.9%
✅ reduce_mean/medium_f32_threads=1-internal/65536 230.99 µs 226.03 µs -2.1%
✅ gather/small_f16_threads=1-internal/4096 448.6 ns 438.1 ns -2.4%
✅ sampling_latency/top_k_per_token 53.75 µs 52.40 µs -2.5%
✅ add/large_f32_threads=1-internal/4194304 603.63 µs 588.14 µs -2.6%
✅ tokenization/encode_tokens_per_second 358.63 µs 348.07 µs -2.9%
✅ matmul/small_generic_f32_threads=8/1x256x256 33.38 µs 32.22 µs -3.5%
✅ matmul/small_generic_bf16_threads=8/1x256x256 30.39 µs 29.15 µs -4.1%
✅ add/small_f32_threads=1-internal/1024 186.2 ns 178.5 ns -4.1%
✅ matmul/medium_generic_f16_threads=1/32x512x512 28.75 µs 27.45 µs -4.6%
✅ add/large_bf16_threads=1-internal/4194304 1.63 ms 1.53 ms -6.2%
✅ add/small_f16_threads=1-internal/1024 457.8 ns 427.0 ns -6.7%
✅ add/medium_f32_threads=1-internal/262144 24.49 µs 22.83 µs -6.8%
✅ qwen3_sampling_processors/top_k_top_p_fast 657.75 µs 609.61 µs -7.3%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 651.14 µs 559.17 µs -14.1%
✅ gather/medium_f32_threads=1-internal/32768 3.97 µs 3.40 µs -14.4%
✅ add/medium_f16_threads=1-internal/262144 113.32 µs 96.94 µs -14.5%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.18 ms 897.91 µs -23.7%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 1.79 ms 1.34 ms -25.0%
🟢 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 491.07 µs 361.86 µs -26.3%
🟢 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 72.73 µs 44.95 µs -38.2%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 66.97 µs 38.84 µs -42.0%
🟢 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 103.11 µs 50.67 µs -50.9%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.25 3.17 5.10 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 19, 2026
The packed 6x16 microkernel landed in #1364 single-threaded only, behind a
structural gate: `PACK_MIN_ROWS` (72) sits above `MAX_MC` (64), so no row
block the multi-threaded driver produces can ever be wide enough to pack.
That was right at the time — the row-block driver packs *inside* each
block, so `t` threads would pack the same panel `t` times — but it left
multi-threaded MoE GEMM on the unpacked path.

Add a second driver for the multi-threaded case that makes the column
panel the unit of work: each panel is packed exactly once, by whichever
lane owns it, and that lane sweeps the whole row extent against it. It is
the single-threaded packed algorithm with its outer loop handed to the
pool. Panel buffers are per-lane (`for_each_init`), sized once at `nc*k`
and reused across the panels that lane draws, so no panel is ever packed
twice and concurrent sessions cannot share one. Lanes write disjoint
column ranges of `c` through a `Send`/`Sync` wrapper, the same pattern and
the same disjointness argument as `x86_sgemm::CPtr`.

The obvious alternative — pack a panel serially and parallelise the row
bands under it, so one pack is literally shared across row blocks — was
built and measured first. It loses: it pays a fork-join per panel and
makes the pack a partly serial term, and on the
`{phi35,qwen3,mixtral} x {96,148,256,365} rows x {2..32} threads` grid it
came in at geomean 0.61-1.04 against the unpacked driver, worse the wider
the pool, at every panel width from 128KB to 4MB. Parallelising the pack
over the `k` dimension as well as over micro-panels recovered part of it
(0.61 -> 0.79 at 16 threads) but never enough. That result is recorded in
the driver's doc comment so it is not rediscovered.

Selection needs two conditions, both measured on that grid:

* a band for every thread, since the unpacked driver slices `m` finer and
  keeps more threads busy on a short `m`;
* `k >= 2048`, where one micro-panel reaches a quarter of the panel
  budget. Below it the per-panel costs stop being amortised and the
  unpacked kernel's strided `b` reads are already cache-resident:
  `mixtral` `k=1024` measured 0.90-1.02 and `qwen3` `k=768` 0.59-1.09,
  while every `k >= 2048` shape won.

Kernel grid, gated, vs the unpacked driver (geomean, higher is better):
t=2 1.178, t=4 1.208, t=8 1.268, t=16 1.237, t=32 1.014; by model
phi35 1.377, qwen3 1.077, mixtral 1.101; worst single cell 0.80.

End to end on the production MoE fixtures at t=512, pure-native both arms
(`nm -C | grep -ci mlas` = 0), interleaved arms with a same-binary null
control, median native/ort ratio: phi35moe -21.6% / -11.3% / -25.2% /
-35.0% / -12.9% at 2/4/8/16/32 threads and qwen3moe -27.6% / -19.3% /
-13.3% / -6.6% / +0.8%, every one of those outside its cell's noise floor
except qwen at 32. Mixtral's ratio is noise-dominated on this host (the
null arm moved up to 66%); on native time against the null it is -2.4% /
+1.7% / -5.3% / -6.3% / -3.6%, i.e. parity to a small win, which is what
the kernel grid predicts for a model with only one of its two GEMMs above
the `k` gate.

The new test drives the driver through a real pool at 2/3/4/8 threads and
asserts, per pool, that the shapes still reach it — plus that at least one
leaves a short final column panel, a hole an earlier version of the test
had. Mutations that skip the short final panel, the row remainder, the
last micro-panel of a panel, or the last band of a sweep each fail it and
nothing else.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
The packed 6x16 microkernel landed in #1364 single-threaded only, behind a
structural gate: `PACK_MIN_ROWS` (72) sits above `MAX_MC` (64), so no row
block the multi-threaded driver produces can ever be wide enough to pack.
That was right at the time — the row-block driver packs *inside* each
block, so `t` threads would pack the same panel `t` times — but it left
multi-threaded MoE GEMM on the unpacked path.

Add a second driver for the multi-threaded case that makes the column
panel the unit of work: each panel is packed exactly once, by whichever
lane owns it, and that lane sweeps the whole row extent against it. It is
the single-threaded packed algorithm with its outer loop handed to the
pool. Panel buffers are per-lane (`for_each_init`), sized once at `nc*k`
and reused across the panels that lane draws, so no panel is ever packed
twice and concurrent sessions cannot share one. Lanes write disjoint
column ranges of `c` through a `Send`/`Sync` wrapper, the same pattern and
the same disjointness argument as `x86_sgemm::CPtr`.

The obvious alternative — pack a panel serially and parallelise the row
bands under it, so one pack is literally shared across row blocks — was
built and measured first. It loses: it pays a fork-join per panel and
makes the pack a partly serial term, and on the
`{phi35,qwen3,mixtral} x {96,148,256,365} rows x {2..32} threads` grid it
came in at geomean 0.61-1.04 against the unpacked driver, worse the wider
the pool, at every panel width from 128KB to 4MB. Parallelising the pack
over the `k` dimension as well as over micro-panels recovered part of it
(0.61 -> 0.79 at 16 threads) but never enough. That result is recorded in
the driver's doc comment so it is not rediscovered.

Selection needs two conditions, both measured on that grid:

* a band for every thread, since the unpacked driver slices `m` finer and
  keeps more threads busy on a short `m`;
* `k >= 2048`, where one micro-panel reaches a quarter of the panel
  budget. Below it the per-panel costs stop being amortised and the
  unpacked kernel's strided `b` reads are already cache-resident:
  `mixtral` `k=1024` measured 0.90-1.02 and `qwen3` `k=768` 0.59-1.09,
  while every `k >= 2048` shape won.

Kernel grid, gated, vs the unpacked driver (geomean, higher is better):
t=2 1.178, t=4 1.208, t=8 1.268, t=16 1.237, t=32 1.014; by model
phi35 1.377, qwen3 1.077, mixtral 1.101; worst single cell 0.80.

End to end on the production MoE fixtures at t=512, pure-native both arms
(`nm -C | grep -ci mlas` = 0), interleaved arms with a same-binary null
control, median native/ort ratio: phi35moe -21.6% / -11.3% / -25.2% /
-35.0% / -12.9% at 2/4/8/16/32 threads and qwen3moe -27.6% / -19.3% /
-13.3% / -6.6% / +0.8%, every one of those outside its cell's noise floor
except qwen at 32. Mixtral's ratio is noise-dominated on this host (the
null arm moved up to 66%); on native time against the null it is -2.4% /
+1.7% / -5.3% / -6.3% / -3.6%, i.e. parity to a small win, which is what
the kernel grid predicts for a model with only one of its two GEMMs above
the `k` gate.

The new test drives the driver through a real pool at 2/3/4/8 threads and
asserts, per pool, that the shapes still reach it — plus that at least one
leaves a short final column panel, a hole an earlier version of the test
had. Mutations that skip the short final panel, the row remainder, the
last micro-panel of a panel, or the last band of a sweep each fail it and
nothing else.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
## What

Lets multi-threaded `gemm_bt` reach the packed 6x16 microkernel that
#1364 landed single-threaded only.

#1364's gate was structural: `PACK_MIN_ROWS` (72) sits above `MAX_MC`
(64), so no row block the multi-threaded driver produces is ever wide
enough to pack. That was the right call — the row-block driver packs
*inside* each block, so `t` threads would pack the same panel `t` times
— but it left multi-threaded MoE GEMM on the unpacked path.

This adds `gemm_bt_panel_parallel`, selected before the row-block driver
when `threads > 1`. The **column panel** is the unit of work: each panel
is packed exactly once, by whichever lane owns it, and that lane sweeps
the whole row extent against it. It is the single-threaded packed
algorithm with its outer loop handed to the pool.

Ownership of the panel buffer is per-lane and explicit:
- **Identity** — created by `for_each_init`, so it belongs to one worker
lane of one invocation. Nothing is cached or static, so concurrent
sessions cannot collide.
- **Capacity** — allocated once at exactly `nc*k` and reused across the
panels that lane draws, never regrown. A short final panel uses a prefix
and the sweep is told how many micro-panels are live.
- **No duplicate pack** — panel indices are disjoint across lanes, so
total pack work is exactly `n16*k` element moves however many threads
run.

Lanes write disjoint column ranges of `c` through a `Send`/`Sync`
wrapper — same pattern and same disjointness argument as the existing
`x86_sgemm::CPtr`.

## The design that lost

The literal "one immutable pack shared across row blocks" shape — pack a
panel serially, parallelise the row bands under it — was **built and
measured first**. It loses, and the negative result is recorded in the
driver's doc comment so it isn't rediscovered:

| panel width | t=2 | t=4 | t=8 | t=16 | t=32 |
|---|---|---|---|---|---|
| 128KB | 0.83 | 0.56 | 0.35 | 0.20 | 0.40 |
| 512KB | 1.04 | 0.84 | 0.76 | 0.61 | 0.77 |
| 4MB | 1.00 | 0.74 | 0.64 | 0.49 | 0.72 |

(geomean vs the unpacked driver, higher is better; 512KB/4MB rows are
after also parallelising the pack over `k`, which recovered a lot and
still wasn't enough). It pays a fork-join per panel and makes the pack a
partly serial term — at `k=6400` a panel is a single micro-panel, so
splitting the pack by micro-panel alone leaves it on one thread while
the other 31 wait.

Distributing whole panels keeps the pack fully parallel and the sweep
barrier-free, and it still packs each panel exactly once.

## Gate

Two conditions, both measured:

1. **A band for every thread** (`m/6 >= threads`). The unpacked driver
slices `m` into `MR`-row blocks and keeps more threads busy on a short
`m` than a `PMR`-row band sweep can; packing there trades a real loss
(idle cores) for a speculative win.
2. **`k >= 2048`** — where one micro-panel reaches a quarter of the
panel budget, `PANEL_BYTES / (PNR * 4 * 4)`. Below it the per-panel
costs stop being amortised and the unpacked kernel's strided `b` reads
are already cache-resident. Measured: `mixtral` `k=1024` 0.90–1.02,
`qwen3` `k=768` 0.59–1.09, every `k >= 2048` shape a win. Admitting `k
>= 1024` scored marginally better on grid geomean (1.194 vs 1.177) but
turned mixtral's first GEMM into a consistent small loss.

## Measurements

**Kernel grid** — `gemm_bt` in-process, best-of-7,
`{phi35,qwen3,mixtral} x {96,148,256,365} rows x {2,4,8,16,32} threads`,
gated, vs the unpacked driver (higher is better):

| | t=2 | t=4 | t=8 | t=16 | t=32 |
|---|---|---|---|---|---|
| geomean | **1.178** | **1.208** | **1.268** | **1.237** | **1.014** |

By model: phi35 **1.377**, qwen3 **1.077**, mixtral **1.101**. Worst
single cell 0.80.

**End to end — re-measured from a production default build.** The
current task asked
for the matrix to be rebuilt without reusing any MLAS-linked historical
claim, so
every number below comes from one fresh run and the earlier end-to-end
table has
been replaced rather than amended.

Build: `cargo build --release -p onnx-genai-bench --features
bench-native` — default
features on, `mlas` never named. Proved MLAS-free four ways, on **both**
arms:
`nm -C` 0, `nm -D` 0, `ldd` 0, `strings -a | grep
'mlas-sys|MlasGemm|mlas_sys'` 0.

Method: `bench_generic` runs ours and ORT **inside a single
invocation**, so the
`ours/ORT` ratio cannot be split by host drift; `ab.py` reverses arm
order every
other trial; `--null-control` re-runs the *baseline binary under a
second name*, so
every cell carries an A/A floor measured in the same invocation. 3
models x
1/2/4/8/16/32 threads x 5 trials x 3 arms = 270 invocations.

Both metrics are given because they disagree, and the disagreement is
the point.

| model | t | Δ ours/ORT | A/A floor | Δ our native | A/A floor |
verdict |
| --- | --- | --- | --- | --- | --- | --- |
| mixtral | 1 | +0.00% | 0.00% | -0.22% | 0.73% | parity |
| mixtral | 2 | -5.80% | 1.14% | -5.14% | 2.03% | **faster** |
| mixtral | 4 | -2.45% | 1.34% | -3.96% | 0.26% | **faster** |
| mixtral | 8 | -5.66% | 11.95% | -4.05% | 0.74% | **faster** |
| mixtral | 16 | +13.12% | 12.52% | -4.50% | 4.21% | **faster** |
| mixtral | 32 | +9.38% | 1.58% | -2.46% | 2.20% | **faster** |
| phi35moe | 1 | +3.12% | 0.00% | +5.36% | 0.22% | parity (re-run, see
below) |
| phi35moe | 2 | -20.38% | 1.05% | -17.71% | 0.06% | **faster** |
| phi35moe | 4 | -14.74% | 0.08% | -28.33% | 15.02% | **faster** |
| phi35moe | 8 | -40.10% | 11.66% | -35.87% | 13.91% | **faster** |
| phi35moe | 16 | -38.64% | 3.73% | -38.42% | 2.50% | **faster** |
| phi35moe | 32 | -11.93% | 0.43% | -12.03% | 0.39% | **faster** |
| qwen3moe | 1 | -0.99% | 0.27% | -1.09% | 1.06% | **faster** |
| qwen3moe | 2 | -18.26% | 0.27% | -19.88% | 1.37% | **faster** |
| qwen3moe | 4 | -22.96% | 10.74% | -41.19% | 22.38% | **faster** |
| qwen3moe | 8 | -21.84% | 6.73% | -47.98% | 36.76% | **faster** |
| qwen3moe | 16 | -9.55% | 5.21% | -15.95% | 2.12% | **faster** |
| qwen3moe | 32 | +4.17% | 7.16% | -9.52% | 9.84% | parity |

**15 faster, 3 parity, 0 slower** on native time — including all six
mixtral cells
and all six qwen3moe cells, which is the parity-or-better-beyond-phi bar
this was
asked to clear.

Two things in that table need stating plainly rather than being smoothed
over:

* **The two mixtral cells where the *ratio* regresses (+13.12% at t=16,
+9.38% at
t=32) are ORT-arm artifacts, not slowdowns.** Decomposing the same
invocations:
at t=32 our native time is **−2.46%** while the ORT arm drifted
**+5.64%** (its
own null moved −0.80%); at t=16 our native is **−4.50%** against a
−0.45% ORT
move and a 12.52% ratio floor. The ratio is the publishable metric when
the ORT
arm is stable, and at 16-32 threads on this host it is not. Reported as
measured.
* **phi35moe t=1 first measured +5.36% native against a 0.22% floor**,
which would
have been a real regression. It is not. The t=1 code path is provably
unchanged —
`gemm_bt_panel_parallel` is selected only under `threads > 1`, and the
supporting
refactor is behaviour-preserving (`panel_cols(k, n16, PANEL_BYTES)`
computes the
  same value the inlined arithmetic did; `PMR`, `PNR`, `PANEL_BYTES`,
`PACK_MIN_BANDS`, `PACK_MIN_ROWS` are all byte-identical to `main`, only
  visibility changed). Re-running that one cell at 9 trials confirms it:
native p50 main 1589.0 ms, **this PR 1665.9 ms, null (main's own binary)
1667.0 ms**
— the control moved with us, so per ledger §39.3 the invocation is
discarded, not
  the cell. Ratio delta +0.64% against a 0.40% floor.

## Tests

`gemm_bt_packed_mt_matches_naive_across_thread_counts` drives the driver
through a real rayon pool at 2/3/4/8 threads over 8 shapes covering both
remainders, one-panel and many-panel extents, panel counts above and
below the thread count, and `k` tails. It asserts per pool that the
shapes still reach this driver, and asserts that at least one shape
leaves a **short final column panel** — a hole an earlier version of
this test had (a mutation that dropped the short final panel survived
it).

Falsified:

| mutation | result |
|---|---|
| short final panel skipped (`n16 / nc`) | fails, only this test |
| row remainder dropped | fails, only this test |
| last micro-panel of each panel skipped | fails, only this test |
| last band of every sweep skipped | fails this + the single-threaded
packed test |
| panel width ignores thread count (perf-only) | stays green, correctly
|
| no-op control | stays green |

## Validation

- `cargo test -p onnx-runtime-ep-cpu --lib` — 1448 passed, 0 failed, 17
ignored
- `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` — 0 warnings
- `cargo fmt --check` — clean
- `cargo check -p onnx-runtime-ep-cpu --features mlas` — clean (the
`mlas` build excludes this whole path)

No ORT CPU fallback, no runtime or default MLAS: the bench binaries link
zero MLAS symbols.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant