Repository navigation
perf(cpu-ep): packed-panel 6x16 microkernel for gemm_bt MoE prefill - #1364
Merged
Merged
Conversation
The MoE expert GEMMs were the largest remaining single-thread loss against ORT/MLAS on the production-native matrix (1.21-1.54x at t=512). The cause was microkernel shape, not blocking: `tile_4x2` keeps k-partials in the SIMD lanes, so every output cell ends in a horizontal reduction and the inner loop issues 6 loads per 8 FMAs, with the `b` side strided by `k`. Add a packed-panel path selected only for prefill-shaped row extents: * `pack_panel` transposes one column panel of `bt` into micro-panel-major order once, bounded by the panel rather than the whole expert bank, and every row band in the sweep reuses it. * `micro_6x16` holds twelve accumulators of *c columns*, so there is no horizontal reduction at all and each k step is two contiguous vector loads plus six broadcasts against twelve FMAs (1.5 FMA/load). * `block_packed` drives column panel -> row band -> micro-panel; both remainders fall to the existing full-k `dot`, unchanged. * Single-threaded `gemm_bt` now hands the whole row extent to one block when the packed kernel is selected, so the pack is amortised once instead of once per 64-row `MAX_MC` block. The gate is twelve full micro-tile bands (`rows >= 72`), which is above `MAX_MC`. That is structural, not tuned: the multi-threaded driver slices `m` into `2*threads` blocks capped at `MAX_MC`, and each such block would re-pack the same panel, so keeping the bar above the cap makes "one pack per sweep" a property of the code rather than of the shape. A `const` assertion pins the relationship so a future `MAX_MC` bump cannot silently let the multi-threaded driver in. Multi-threaded behaviour is therefore identical to before this commit for every shape, which is the only honest position available: on the shared host this was developed on, a 2-thread A/B of two binaries *traced to be on the identical code path* (every row block `rows=56..64`, all unpacked) still differed by ~40% on the median. No multi-thread claim is made in either direction. Sharing one pack across threads is the follow-up. Production-native single-thread A/B (both binaries `--no-default-features --features bench-native,cuda-13000`, zero MLAS symbols by `nm`, arms interleaved innermost with alternating order, `--native-threads 1 --ort-intra-threads 1`, metric native_min/ort_min, four passes): moe_qwen3moe_h2048_i768_e16_t512 1.297 -> 1.101 -15.1% moe_phi35moe_h2048_i6400_e4_t512 1.356 -> 1.009 -25.6% moe_mixtral_h1024_i3584_e8_t512 1.226 -> 1.173 -4.3% reproduced across four independent passes (qwen -16.7/-23.4/-19.6/-15.1%, phi35 -16.9/-24.9/-27.6/-25.6%, mixtral -1.3/-6.6/-3.0/-4.3%), parity PASS in every cell. The t=1 and t=32 rows are provably on the identical code path (traced row extents 14-19, below the gate) and their spread is the host noise floor. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
force-pushed
the
squad/leon-p23-moe-packed-gemm
branch
from
August 19, 2026 02:33
0777291 to
c37b0ea
Compare
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
… that failed (#1373) Documentation only — records Phase 17 in the CPU-EP ledger. ## What lands here * **§37.1–37.2 — #1245**, the pure-Horner `exp8`. Full 7-model A/B across two independent passes, plus the accuracy contract: 1 ULP over **1,118,743,631** `f32` values by exhaustive sweep, and the story of the sweep that first reported `max_ulp = 8388646` because it included the deliberate flush-to-zero threshold. * **§37.3 — #1364**, the packed-panel `6x16` `gemm_bt` microkernel. Four independent passes. Phi-3.5-MoE prefill reaches **1.009** (parity with ORT) from 1.356; Qwen3-MoE **1.101** from 1.297. * **§37.5** — the unmerged two-pass online softmax, kept as a negative result (2.4× worse) with the reason it loses: §29 already removed the copy that would have made the extra pass expensive. ## §37.4 is the part that matters Choosing the packed kernel's row gate produced two negative results, and the second one is a methodology finding, not a perf finding. A four-band gate looked fine single-threaded and lost ~29% on the 8-thread Phi-3.5-MoE median. Raising it to eight bands removed that, but then the 2-thread cells showed +6% to +11% — which reads as a real multi-threaded regression. It is not usable data. After raising the gate above `MAX_MC`, the two binaries are **traced to be on the identical code path at two threads** (every row block `rows=56..64`, `packed=false`, verified with an instrumented build) — and they still measured **~40% apart on the median**, in the favourable direction for the arm that had just been slower. That is a worse null arm than §36.3's: ~40% at **two** threads on a code path proved identical, against §36.3's ±20–30% at thirty-two. Low thread counts are not the quiet regime on a shared box; they are the regime where one noisy neighbour is a larger fraction of what is running. The consequence is stated plainly in the section: this document now has two independent measurements of its own instrument disagreeing with itself by 20–40%, and every cross-arm claim in it that rests on a single pass without a null control should be read as provisional. ## Scope No code. No behaviour change. Both PRs being documented are already merged (#1245 as `512885a7b`, #1364 as `307c7277d`). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
…h it (#1380) ## The problem Two phases in a row reported a delta and then had to reconstruct, *after the fact*, whether the instrument could resolve it — §36.3 from cells the change could not reach, §37.4 from a code path traced to be identical. That reconstruction is an argument, and it arrives too late to change what was measured. ## The change `ab.py --null-control` adds a third arm that is **the first arm's own binary, with the first arm's `--arm-env`, under a second name**, interleaved and order-alternated exactly like the real arms. It cannot measure the change. Its delta is the host's noise floor for that cell, *in that invocation*. The driver also gained a deltas table (it previously printed only per-arm medians, leaving the subtraction to the reader): ``` === deltas vs 'before' (median ratio; negative = arm is faster) === The 'null' column is the same binary as 'before'. Its delta is this host's noise floor for the cell, so a real delta no larger than it is not a result. moe_phi35moe_h2048_i6400_e4_t512 t=1 after: -25.48% > noise (0.19%) null: -0.19% moe_mixtral_h1024_i3584_e8_t512 t=1 after: -2.38% > noise (0.63%) null: -0.63% ``` Anything inside the control prints `WITHIN NOISE` instead of a number that looks publishable. ## What it measured (§38) Same-binary control arm, `|Δ|` of the median native/ort ratio against the identical binary: | cell | 1 thread | 2 threads | 8 threads | | --- | --- | --- | --- | | `moe_mixtral_h1024_i3584_e8_t512` | 0.63% | 0.22% | **19.54%** | | `moe_qwen3moe_h2048_i768_e16_t512` | 4.75% | 3.97% | **12.68%** | | `moe_phi35moe_h2048_i6400_e4_t512` | 0.19% | 4.28% | 3.03% | | `sm_bert_b8_s128` | 0.21% | — | — | | `sm_decode_h32_kv8192` | 1.32% | — | — | | `sm_prefill_h32_s512` | 0.22% | — | — | **Single-thread cells are tight** — under 5%, mostly under 1.5% — so the single-thread ratios this ledger publishes are resolvable at the effect sizes it claims. **Eight-thread cells are not**: a 12–20% floor retires the 5–15% movements several earlier phases tabulated at `t=8`, exactly as §36.3 warned. Both phase-17 merges are re-measured against their own control, in the same invocation — the first self-controlled cells in the document: | change | cell (1 thread) | Δ median ratio | floor | verdict | | --- | --- | --- | --- | --- | | #1364 packed panel | `moe_phi35moe…t512` | **−25.48%** | 0.19% | 134× the floor | | #1364 packed panel | `moe_qwen3moe…t512` | **−21.56%** | 4.75% | 4.5× the floor | | #1364 packed panel | `moe_mixtral…t512` | −2.38% | 0.63% | 3.8× the floor | | #1245 Horner `exp8` | `sm_bert_b8_s128` | **−10.28%** | 0.21% | 49× the floor | | #1245 Horner `exp8` | `sm_decode_h32_kv8192` | **−11.73%** | 1.32% | 8.9× the floor | | #1245 Horner `exp8` | `sm_prefill_h32_s512` | **−10.33%** | 0.22% | 47× the floor | ## It corrects my own §37.4 §37.4 reported ~40% between two distinct binaries traced to identical code paths at two threads, and read that as the two-thread noise floor. **That reading was wrong.** The same-binary control puts the two-thread floor at 0.22–4.28%, an order of magnitude tighter, and the absolute ratios differ between the sessions (mixtral `t=2` reads 1.35 here, 2.23 there, same fixture, same thread count). The 40% was a *session*: a load episode long enough to outlast the arm alternation. That is a more useful finding than a floor, because it is not something interleaving can fix — a whole invocation can be poisoned, and only a control arm inside that invocation exposes it. §38.3 states the narrower rule that follows: run the control in every invocation whose result will be published, and discard the **invocation**, not the cell, when the control moves more than the effect. ## Scope * Additive: without `--null-control` the driver behaves as before, except that multi-arm runs now also print a deltas table. * `ruff check` clean. (`ruff format` does not pass on this file before or after this change, so it is left alone rather than reformatted in a perf-tooling PR.) * No Rust, no behaviour change in the EP. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
Main already carries a section 39 (#1364's harness follow-up, merged while this branch was open), so this appended one collided with it. Renumber to 40/Phase 20 rather than merging two sections under the same number. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
✅ Benchmarks — No RegressionComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
The packed 6x16 microkernel landed in #1364 single-threaded only, behind a structural gate: `PACK_MIN_ROWS` (72) sits above `MAX_MC` (64), so no row block the multi-threaded driver produces can ever be wide enough to pack. That was right at the time — the row-block driver packs *inside* each block, so `t` threads would pack the same panel `t` times — but it left multi-threaded MoE GEMM on the unpacked path. Add a second driver for the multi-threaded case that makes the column panel the unit of work: each panel is packed exactly once, by whichever lane owns it, and that lane sweeps the whole row extent against it. It is the single-threaded packed algorithm with its outer loop handed to the pool. Panel buffers are per-lane (`for_each_init`), sized once at `nc*k` and reused across the panels that lane draws, so no panel is ever packed twice and concurrent sessions cannot share one. Lanes write disjoint column ranges of `c` through a `Send`/`Sync` wrapper, the same pattern and the same disjointness argument as `x86_sgemm::CPtr`. The obvious alternative — pack a panel serially and parallelise the row bands under it, so one pack is literally shared across row blocks — was built and measured first. It loses: it pays a fork-join per panel and makes the pack a partly serial term, and on the `{phi35,qwen3,mixtral} x {96,148,256,365} rows x {2..32} threads` grid it came in at geomean 0.61-1.04 against the unpacked driver, worse the wider the pool, at every panel width from 128KB to 4MB. Parallelising the pack over the `k` dimension as well as over micro-panels recovered part of it (0.61 -> 0.79 at 16 threads) but never enough. That result is recorded in the driver's doc comment so it is not rediscovered. Selection needs two conditions, both measured on that grid: * a band for every thread, since the unpacked driver slices `m` finer and keeps more threads busy on a short `m`; * `k >= 2048`, where one micro-panel reaches a quarter of the panel budget. Below it the per-panel costs stop being amortised and the unpacked kernel's strided `b` reads are already cache-resident: `mixtral` `k=1024` measured 0.90-1.02 and `qwen3` `k=768` 0.59-1.09, while every `k >= 2048` shape won. Kernel grid, gated, vs the unpacked driver (geomean, higher is better): t=2 1.178, t=4 1.208, t=8 1.268, t=16 1.237, t=32 1.014; by model phi35 1.377, qwen3 1.077, mixtral 1.101; worst single cell 0.80. End to end on the production MoE fixtures at t=512, pure-native both arms (`nm -C | grep -ci mlas` = 0), interleaved arms with a same-binary null control, median native/ort ratio: phi35moe -21.6% / -11.3% / -25.2% / -35.0% / -12.9% at 2/4/8/16/32 threads and qwen3moe -27.6% / -19.3% / -13.3% / -6.6% / +0.8%, every one of those outside its cell's noise floor except qwen at 32. Mixtral's ratio is noise-dominated on this host (the null arm moved up to 66%); on native time against the null it is -2.4% / +1.7% / -5.3% / -6.3% / -3.6%, i.e. parity to a small win, which is what the kernel grid predicts for a model with only one of its two GEMMs above the `k` gate. The new test drives the driver through a real pool at 2/3/4/8 threads and asserts, per pool, that the shapes still reach it — plus that at least one leaves a short final column panel, a hole an earlier version of the test had. Mutations that skip the short final panel, the row remainder, the last micro-panel of a panel, or the last band of a sweep each fail it and nothing else. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
The packed 6x16 microkernel landed in #1364 single-threaded only, behind a structural gate: `PACK_MIN_ROWS` (72) sits above `MAX_MC` (64), so no row block the multi-threaded driver produces can ever be wide enough to pack. That was right at the time — the row-block driver packs *inside* each block, so `t` threads would pack the same panel `t` times — but it left multi-threaded MoE GEMM on the unpacked path. Add a second driver for the multi-threaded case that makes the column panel the unit of work: each panel is packed exactly once, by whichever lane owns it, and that lane sweeps the whole row extent against it. It is the single-threaded packed algorithm with its outer loop handed to the pool. Panel buffers are per-lane (`for_each_init`), sized once at `nc*k` and reused across the panels that lane draws, so no panel is ever packed twice and concurrent sessions cannot share one. Lanes write disjoint column ranges of `c` through a `Send`/`Sync` wrapper, the same pattern and the same disjointness argument as `x86_sgemm::CPtr`. The obvious alternative — pack a panel serially and parallelise the row bands under it, so one pack is literally shared across row blocks — was built and measured first. It loses: it pays a fork-join per panel and makes the pack a partly serial term, and on the `{phi35,qwen3,mixtral} x {96,148,256,365} rows x {2..32} threads` grid it came in at geomean 0.61-1.04 against the unpacked driver, worse the wider the pool, at every panel width from 128KB to 4MB. Parallelising the pack over the `k` dimension as well as over micro-panels recovered part of it (0.61 -> 0.79 at 16 threads) but never enough. That result is recorded in the driver's doc comment so it is not rediscovered. Selection needs two conditions, both measured on that grid: * a band for every thread, since the unpacked driver slices `m` finer and keeps more threads busy on a short `m`; * `k >= 2048`, where one micro-panel reaches a quarter of the panel budget. Below it the per-panel costs stop being amortised and the unpacked kernel's strided `b` reads are already cache-resident: `mixtral` `k=1024` measured 0.90-1.02 and `qwen3` `k=768` 0.59-1.09, while every `k >= 2048` shape won. Kernel grid, gated, vs the unpacked driver (geomean, higher is better): t=2 1.178, t=4 1.208, t=8 1.268, t=16 1.237, t=32 1.014; by model phi35 1.377, qwen3 1.077, mixtral 1.101; worst single cell 0.80. End to end on the production MoE fixtures at t=512, pure-native both arms (`nm -C | grep -ci mlas` = 0), interleaved arms with a same-binary null control, median native/ort ratio: phi35moe -21.6% / -11.3% / -25.2% / -35.0% / -12.9% at 2/4/8/16/32 threads and qwen3moe -27.6% / -19.3% / -13.3% / -6.6% / +0.8%, every one of those outside its cell's noise floor except qwen at 32. Mixtral's ratio is noise-dominated on this host (the null arm moved up to 66%); on native time against the null it is -2.4% / +1.7% / -5.3% / -6.3% / -3.6%, i.e. parity to a small win, which is what the kernel grid predicts for a model with only one of its two GEMMs above the `k` gate. The new test drives the driver through a real pool at 2/3/4/8 threads and asserts, per pool, that the shapes still reach it — plus that at least one leaves a short final column panel, a hole an earlier version of the test had. Mutations that skip the short final panel, the row remainder, the last micro-panel of a panel, or the last band of a sweep each fail it and nothing else. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
## What Lets multi-threaded `gemm_bt` reach the packed 6x16 microkernel that #1364 landed single-threaded only. #1364's gate was structural: `PACK_MIN_ROWS` (72) sits above `MAX_MC` (64), so no row block the multi-threaded driver produces is ever wide enough to pack. That was the right call — the row-block driver packs *inside* each block, so `t` threads would pack the same panel `t` times — but it left multi-threaded MoE GEMM on the unpacked path. This adds `gemm_bt_panel_parallel`, selected before the row-block driver when `threads > 1`. The **column panel** is the unit of work: each panel is packed exactly once, by whichever lane owns it, and that lane sweeps the whole row extent against it. It is the single-threaded packed algorithm with its outer loop handed to the pool. Ownership of the panel buffer is per-lane and explicit: - **Identity** — created by `for_each_init`, so it belongs to one worker lane of one invocation. Nothing is cached or static, so concurrent sessions cannot collide. - **Capacity** — allocated once at exactly `nc*k` and reused across the panels that lane draws, never regrown. A short final panel uses a prefix and the sweep is told how many micro-panels are live. - **No duplicate pack** — panel indices are disjoint across lanes, so total pack work is exactly `n16*k` element moves however many threads run. Lanes write disjoint column ranges of `c` through a `Send`/`Sync` wrapper — same pattern and same disjointness argument as the existing `x86_sgemm::CPtr`. ## The design that lost The literal "one immutable pack shared across row blocks" shape — pack a panel serially, parallelise the row bands under it — was **built and measured first**. It loses, and the negative result is recorded in the driver's doc comment so it isn't rediscovered: | panel width | t=2 | t=4 | t=8 | t=16 | t=32 | |---|---|---|---|---|---| | 128KB | 0.83 | 0.56 | 0.35 | 0.20 | 0.40 | | 512KB | 1.04 | 0.84 | 0.76 | 0.61 | 0.77 | | 4MB | 1.00 | 0.74 | 0.64 | 0.49 | 0.72 | (geomean vs the unpacked driver, higher is better; 512KB/4MB rows are after also parallelising the pack over `k`, which recovered a lot and still wasn't enough). It pays a fork-join per panel and makes the pack a partly serial term — at `k=6400` a panel is a single micro-panel, so splitting the pack by micro-panel alone leaves it on one thread while the other 31 wait. Distributing whole panels keeps the pack fully parallel and the sweep barrier-free, and it still packs each panel exactly once. ## Gate Two conditions, both measured: 1. **A band for every thread** (`m/6 >= threads`). The unpacked driver slices `m` into `MR`-row blocks and keeps more threads busy on a short `m` than a `PMR`-row band sweep can; packing there trades a real loss (idle cores) for a speculative win. 2. **`k >= 2048`** — where one micro-panel reaches a quarter of the panel budget, `PANEL_BYTES / (PNR * 4 * 4)`. Below it the per-panel costs stop being amortised and the unpacked kernel's strided `b` reads are already cache-resident. Measured: `mixtral` `k=1024` 0.90–1.02, `qwen3` `k=768` 0.59–1.09, every `k >= 2048` shape a win. Admitting `k >= 1024` scored marginally better on grid geomean (1.194 vs 1.177) but turned mixtral's first GEMM into a consistent small loss. ## Measurements **Kernel grid** — `gemm_bt` in-process, best-of-7, `{phi35,qwen3,mixtral} x {96,148,256,365} rows x {2,4,8,16,32} threads`, gated, vs the unpacked driver (higher is better): | | t=2 | t=4 | t=8 | t=16 | t=32 | |---|---|---|---|---|---| | geomean | **1.178** | **1.208** | **1.268** | **1.237** | **1.014** | By model: phi35 **1.377**, qwen3 **1.077**, mixtral **1.101**. Worst single cell 0.80. **End to end — re-measured from a production default build.** The current task asked for the matrix to be rebuilt without reusing any MLAS-linked historical claim, so every number below comes from one fresh run and the earlier end-to-end table has been replaced rather than amended. Build: `cargo build --release -p onnx-genai-bench --features bench-native` — default features on, `mlas` never named. Proved MLAS-free four ways, on **both** arms: `nm -C` 0, `nm -D` 0, `ldd` 0, `strings -a | grep 'mlas-sys|MlasGemm|mlas_sys'` 0. Method: `bench_generic` runs ours and ORT **inside a single invocation**, so the `ours/ORT` ratio cannot be split by host drift; `ab.py` reverses arm order every other trial; `--null-control` re-runs the *baseline binary under a second name*, so every cell carries an A/A floor measured in the same invocation. 3 models x 1/2/4/8/16/32 threads x 5 trials x 3 arms = 270 invocations. Both metrics are given because they disagree, and the disagreement is the point. | model | t | Δ ours/ORT | A/A floor | Δ our native | A/A floor | verdict | | --- | --- | --- | --- | --- | --- | --- | | mixtral | 1 | +0.00% | 0.00% | -0.22% | 0.73% | parity | | mixtral | 2 | -5.80% | 1.14% | -5.14% | 2.03% | **faster** | | mixtral | 4 | -2.45% | 1.34% | -3.96% | 0.26% | **faster** | | mixtral | 8 | -5.66% | 11.95% | -4.05% | 0.74% | **faster** | | mixtral | 16 | +13.12% | 12.52% | -4.50% | 4.21% | **faster** | | mixtral | 32 | +9.38% | 1.58% | -2.46% | 2.20% | **faster** | | phi35moe | 1 | +3.12% | 0.00% | +5.36% | 0.22% | parity (re-run, see below) | | phi35moe | 2 | -20.38% | 1.05% | -17.71% | 0.06% | **faster** | | phi35moe | 4 | -14.74% | 0.08% | -28.33% | 15.02% | **faster** | | phi35moe | 8 | -40.10% | 11.66% | -35.87% | 13.91% | **faster** | | phi35moe | 16 | -38.64% | 3.73% | -38.42% | 2.50% | **faster** | | phi35moe | 32 | -11.93% | 0.43% | -12.03% | 0.39% | **faster** | | qwen3moe | 1 | -0.99% | 0.27% | -1.09% | 1.06% | **faster** | | qwen3moe | 2 | -18.26% | 0.27% | -19.88% | 1.37% | **faster** | | qwen3moe | 4 | -22.96% | 10.74% | -41.19% | 22.38% | **faster** | | qwen3moe | 8 | -21.84% | 6.73% | -47.98% | 36.76% | **faster** | | qwen3moe | 16 | -9.55% | 5.21% | -15.95% | 2.12% | **faster** | | qwen3moe | 32 | +4.17% | 7.16% | -9.52% | 9.84% | parity | **15 faster, 3 parity, 0 slower** on native time — including all six mixtral cells and all six qwen3moe cells, which is the parity-or-better-beyond-phi bar this was asked to clear. Two things in that table need stating plainly rather than being smoothed over: * **The two mixtral cells where the *ratio* regresses (+13.12% at t=16, +9.38% at t=32) are ORT-arm artifacts, not slowdowns.** Decomposing the same invocations: at t=32 our native time is **−2.46%** while the ORT arm drifted **+5.64%** (its own null moved −0.80%); at t=16 our native is **−4.50%** against a −0.45% ORT move and a 12.52% ratio floor. The ratio is the publishable metric when the ORT arm is stable, and at 16-32 threads on this host it is not. Reported as measured. * **phi35moe t=1 first measured +5.36% native against a 0.22% floor**, which would have been a real regression. It is not. The t=1 code path is provably unchanged — `gemm_bt_panel_parallel` is selected only under `threads > 1`, and the supporting refactor is behaviour-preserving (`panel_cols(k, n16, PANEL_BYTES)` computes the same value the inlined arithmetic did; `PMR`, `PNR`, `PANEL_BYTES`, `PACK_MIN_BANDS`, `PACK_MIN_ROWS` are all byte-identical to `main`, only visibility changed). Re-running that one cell at 9 trials confirms it: native p50 main 1589.0 ms, **this PR 1665.9 ms, null (main's own binary) 1667.0 ms** — the control moved with us, so per ledger §39.3 the invocation is discarded, not the cell. Ratio delta +0.64% against a 0.40% floor. ## Tests `gemm_bt_packed_mt_matches_naive_across_thread_counts` drives the driver through a real rayon pool at 2/3/4/8 threads over 8 shapes covering both remainders, one-panel and many-panel extents, panel counts above and below the thread count, and `k` tails. It asserts per pool that the shapes still reach this driver, and asserts that at least one shape leaves a **short final column panel** — a hole an earlier version of this test had (a mutation that dropped the short final panel survived it). Falsified: | mutation | result | |---|---| | short final panel skipped (`n16 / nc`) | fails, only this test | | row remainder dropped | fails, only this test | | last micro-panel of each panel skipped | fails, only this test | | last band of every sweep skipped | fails this + the single-threaded packed test | | panel width ignores thread count (perf-only) | stays green, correctly | | no-op control | stays green | ## Validation - `cargo test -p onnx-runtime-ep-cpu --lib` — 1448 passed, 0 failed, 17 ignored - `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` — 0 warnings - `cargo fmt --check` — clean - `cargo check -p onnx-runtime-ep-cpu --features mlas` — clean (the `mlas` build excludes this whole path) No ORT CPU fallback, no runtime or default MLAS: the bench binaries link zero MLAS symbols. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
A packed-panel
6x16microkernel forgemm_bt, the kernel that computesC = A·Btᵀdirectly on the[out][in]-major MoE expert weights (#1241). It isentered only from the single-threaded driver on prefill-shaped row extents;
decode, small-token MoE and every multi-threaded shape keep the existing
tile_4x2path unchanged.Why
MoE expert GEMMs were the largest remaining single-thread loss on the
production-native matrix — §31.5/§31.8 of the CPU-EP ledger put MoE prefill at
1.22–1.56x ORT at
t=512and attributed it to "microkernel efficiencyagainst MLAS's packed panels and wider register tiles". That attribution is
right, and it is fixable without handing anything to ORT's CPU EP.
tile_4x2keeps k-partials in the SIMD lanes. Two consequences: everyoutput cell ends in a horizontal reduction, and the
bside of the inner loopis strided by
k— six loads per eight FMAs.How
tile_4x2(unchanged, still the default)micro_6x16(new)baccesskpack_paneltransposes one column panel ofbtinto micro-panel-major order(
packed[s*k*PNR + p*PNR + t]). This is the only transpose in the path, it isbounded by the panel rather than the whole expert bank, and every row band in
the sweep reuses it.
block_packeddrives column panel → row band → micro-panel. Both remainders(
rows % 6,n % 16) fall through to the existing full-kdot, which isthe same code the unpacked path already uses for its edges.
gemm_bthands the whole row extent to one block when thepacked kernel is selected, so the pack is amortised once rather than once per
64-row
MAX_MCblock.The gate is structural, not tuned
PACK_MIN_ROWS = PMR * PACK_MIN_BANDS = 72, deliberately aboveMAX_MC(64),pinned by a
constassertion:The multi-threaded driver slices
minto2*threadsblocks capped atMAX_MC,and each block re-packs the same panel. Keeping the bar above that cap makes
"the panel is packed once per sweep" a property of the code rather than of the
shape, and means multi-threaded behaviour is byte-identical to
mainforevery shape. The assertion stops a future
MAX_MCbump from silently undoingthat.
Why I will not make a multi-threaded claim
I tried to. It does not survive its own control. A 2-thread A/B of two binaries
that I traced to be on the identical code path (every row block
rows=56..64,packed=false) still came out ~40% apart on the median. Multi-threadmeasurement on this shared 32-core box is not usable at this effect size, in
either direction, so this PR takes the position that costs nothing: MT is left
exactly as it was. Sharing one pack across threads (or making
gemm_bt_col_parallelthe packed path, since each of its tasks already owns allmrows) is the natural follow-up, and it needs a quiet machine.Evidence
Both binaries built
--no-default-features --features bench-native,cuda-13000,differing only in
matmul.rs. The arm is proved by the linker, not the flag:nm -C <bin> | grep -ci mlas= 0 on both. Arms interleaved innermost in onedriver process, arm order alternating per rep, thread-matched
--native-threads 1 --ort-intra-threads 1. Metric isnative_min / ort_min(lower is better; <1.000 beats ORT).
Single thread,
t=512— the shapes that reach the new kernelFinal gate, shipped binary (6 reps):
moe_phi35moe_h2048_i6400_e4_t512moe_qwen3moe_h2048_i768_e16_t512moe_mixtral_h1024_i3584_e8_t512Four independent passes, Δ on the min ratio:
Phi-3.5-MoE prefill crosses 1.000 in the shipped pass and Qwen3-MoE lands at
1.10, from a 1.30–1.54 loss. Mixtral moves least, and I believe that is real
rather than measurement: its expert bank is ~352 MB of fp32, so that shape is
closer to DRAM-bound than microkernel-bound and a better inner loop has less to
win.
t=1/t=32— unchanged by construction, and verifiedTraced with a temporary instrumented build: every
gemm_bt_blockatt=32hasrowsin 14–19, far below the gate, sopacked=falseand the executed codeis the unpacked path. Their run-to-run spread (±10% on min, ±3% on median, sign
flipping between passes) is therefore a direct read of this box's noise
floor. At
t=512the traced extents are 74–365 rows, all packed.Correctness
cargo test -p onnx-runtime-ep-cpu --lib→ 1434 passed, 0 failed, 17 ignored.gemm_bt_packed_kernel_matches_naive_single_threaded: 10 shapesagainst
naive_bton a 1-thread pool, covering both remainders,ktails, theshortest admitted
k, and multi-column-panel sweeps. It carries anon-vacuity assertion — if a policy change stops these shapes from reaching
the packed kernel, the test fails loudly rather than silently measuring the old
path.
corrupting the pack layout to
dst[t*k+p]. Each mutation fails only the newtest; the other five
gemm_bttests stay green.cargo fmt --checkclean,cargo check --features mlascompiles.Review
Reviewed by Opus (
claude-opus-4.8), verdict APPROVE-WITH-NITS, no blockingfindings. It independently re-ran the packed test with 7 additional adversarial
shapes (the
nc-floor branch atk>8192, single-micro-panel wide sweeps,whole-
mwidening atm=1000, maximal row/column remainders) underAddressSanitizer — all clean — and verified the panel-index bounds, the
jend-j0 ≡ 0 (mod PNR)contract, remainder exhaustiveness/disjointness, and thatno consumer of
gemm_btdepends on its accumulation order. Its one substantivenit — that the original
rows >= 48gate did not keep large-mmulti-threadedblocks out of the packed path, contrary to the comment — is what produced the
structural
> MAX_MCgate and theconstassertion above.Scope
tile_4x2,block,dot,gemm_bt_block_scalarand themlasarm untouched.