diff --git a/docs/benchmarks/2026-08-21-decode-worker-cpu-placement.md b/docs/benchmarks/2026-08-21-decode-worker-cpu-placement.md new file mode 100644 index 0000000000..7dbc665f6e --- /dev/null +++ b/docs/benchmarks/2026-08-21-decode-worker-cpu-placement.md @@ -0,0 +1,98 @@ +# The "t=8 wash" is worker-to-CPU placement, not the kernel + +**Date:** 2026-08-21 · **Owner:** Roy (CPU MatMul) · **Host:** AMD EPYC 9V74, +32 vCPU (16c x 2 SMT, **siblings adjacent**), AVX2/FMA/F16C, no AVX-512/VNNI. +Ledger entry: §24 of [`CPU_MATMUL_ASSIGNMENT.md`](../performance/CPU_MATMUL_ASSIGNMENT.md). +**Negative result. No kernel change.** Handed to the runtime owner as **#1680**. + +--- + +## 1. The premise, and why it does not reproduce + +The reported observation was that #1628's packed-nibble int4 acc4 win holds at +t=1, t=4 and t=16 but **washes out at t=8**, with the implied action being to +look at the kernel. + +A thread-count *label* is not a pool width. Re-measured against explicitly set +pool widths on the same binary and shapes: + +| pool width | 1 | 4 | 8 | 12 | 16 | +|---|---|---|---|---|---| +| acc4 speedup vs acc0 | 1.66x | 1.65x | **1.66x** | 1.238x | 0.993x | + +**There is no t=8 anomaly.** The win is flat through width 8 and collapses +*after* it. The shape of the real effect — flat, then a cliff between 8 and 12, +reaching parity at 16 — is not the shape that was reported, and tuning a kernel +against the reported shape would have been tuning against noise. + +Width 8 is exactly the physical core count of one half of this machine, which is +the tell. + +## 2. Hypotheses tried and discarded + +**Memory bandwidth.** A STREAM-style all-thread sweep read 83 GB/s against a +41 GB/s decode draw, which would make the loop comfortably compute-bound. That +figure **does not reconcile with §22 of the ledger**, which measured this host +at 31-36 GB/s within a CCX and ~56.6 GB/s across both — and §22's numbers are +the ones the ledger stands behind, because they were taken with the access +pattern the decode loop actually uses. Against §22, a 41 GB/s draw is 72% of the +across-CCX ceiling and *above* the within-CCX one. + +So bandwidth is **not** dismissed by the 83 GB/s sweep, and is not dismissed on +that basis. It is ruled out by §3's placement A/B, which holds shapes, bytes, +thread count and binary constant and changes **only which CPUs the workers sit +on**. A bandwidth ceiling is indifferent to that. The result is not. + +**Task grain.** Shard count and per-shard barrier time were read directly and +move as expected with width — no straggler, no grain cliff at 8 or 12. Discarded. + +## 3. Root cause: logical-order pinning + +`decode_spmd.rs::node_shards` pins worker *i* to `allowed_cpus()[i]` — **logical +order, no topology awareness**. + +On this host SMT siblings are adjacent: CPUs 0 and 1 are the two hardware +threads of physical core 0. So a 16-worker pool lands on CPUs 0-15, which is +**8 physical cores**, and every worker contends with a sibling for the same +execution units. A width-12 pool puts 4 of its 12 workers on shared cores, which +is exactly where the table above starts to bend. + +Verified two ways. + +**Observationally**, by reading `/proc//task/*/stat` field 39 (`processor`) +for every worker thread during a run: the 16 workers report CPUs 0 through 15. + +**Causally**, by changing nothing but placement on the same binary and shapes: + +| placement, 16 workers | speedup | +|---|---| +| default (`allowed_cpus()[i]`, CPUs 0-15 = 8 physical cores) | 0.982x | +| one worker per physical core (`taskset -c 0,2,4,...`) | **1.225x** | + +Same kernel, same data, same worker count. The only variable is which CPUs they +sit on, and it moves the result by 1.25x. + +## 4. Disposition + +**No kernel change was made.** The kernel is not implicated by any measurement +here, and distorting it to compensate for scheduler placement would bake a +host-topology artefact into shipped code — the specific failure mode the +directive named. Filed as **#1680** with the pool-width sweep, the `/proc` +evidence and the placement A/B, for whoever owns `decode_spmd.rs`. + +## 5. The part that affects everyone else's measurements + +**Every unpinned multi-thread number taken on this host above pool width 8 is +contaminated**, and the contamination is silent — it looks like a kernel that +stops scaling rather than like a placement bug. + +Two workable rules, either sufficient: + +- pin explicitly with `taskset -c 0,2,4,6,...` (even CPUs are distinct physical + cores on this machine), or +- keep pool width at or below 8, where the default placement happens to be + correct by accident. + +Both §23 and §25's records were taken under the first rule. Any older +multi-thread figure in this directory taken above width 8 without pinning should +be treated as unreconstructed. diff --git a/docs/benchmarks/2026-08-21-half-decode-layout-divergence.md b/docs/benchmarks/2026-08-21-half-decode-layout-divergence.md new file mode 100644 index 0000000000..f29b5708ba --- /dev/null +++ b/docs/benchmarks/2026-08-21-half-decode-layout-divergence.md @@ -0,0 +1,179 @@ +# `f16`/`bf16` decode: three operators, one backend, three different prices + +**Date:** 2026-08-21 · **Owner:** Roy (CPU MatMul) · **Host:** AMD EPYC 9V74, +32 vCPU (16c x 2 SMT), AVX2/FMA/F16C, no AVX-512/VNNI. +Ledger entry: §25 of [`CPU_MATMUL_ASSIGNMENT.md`](../performance/CPU_MATMUL_ASSIGNMENT.md). +Merged as `2e1cfb67c` (#1687), closing #1381. Residual split to **#1702**. + +All timings `taskset`-pinned to distinct physical cores — see the +[placement record](2026-08-21-decode-worker-cpu-placement.md). + +--- + +## 1. The enumeration + +#1381's dispatch comment claimed the divergence was already closed: "both +operators, both stored orders and both 16-bit formats reach the same GEMV +backend". Same **backend**, not same **kernel** — and the operator list was +incomplete. Enumerated with route counters rather than by reading cfgs: + +| operator | stored order | kernel taken | +|---|---|---| +| `Gemm` transB=1 | `[N,K]` | `gemv_half_nk` | +| `Gemm` transB=0 | `[K,N]` | `gemv_half_kn` | +| `MatMul` | `[K,N]` | `gemv_half_kn` | +| `FusedMatMulBias` | `[K,N]` | **no 16-bit GEMV at all** | + +The fourth row was found empirically. `count_half_decode_gemv()` lives in +`MatMulKernel::execute_with_backend`, and `fused_matmul_bias.rs` calls the free +`matmul_dense_prepacked_into`, so a probe over a decode step reads +`matmul_gemv=1 fused_gemv=0`. Its f16 GEMV is under +`cfg(any(target_os = "macos", target_os = "ios"))`. + +The MatMul-side transposed row had **no test at all**: the `decode_matmul` +fixtures only ever built B as `[k,n]`. That is the fifth unmeasured region +behind a gate this project has recorded (cf. ledger §11, §12, §14, §19). + +## 2. Root cause: `[K,N]` crosses a page every `p` + +Walking a single output column of a `[K,N]` weight strides `n*2` bytes between +consecutive `p` — **12 KB at n=6144**. The L2 streaming prefetcher does not +cross page boundaries, so it cannot run ahead: every few steps it restarts. + +Direct kernel A/B on identical bytes, same call, only the stored order differing: + +| shape | `gemv_half_kn` (us) | `gemv_half_nk` (us) | penalty | +|---|---|---|---| +| qkv 4096x6144 | 5022 | 1688 | **2.98x** | +| down 14336x4096 | 11015 | 7068 | **1.56x** | + +The penalty tracks the stride, as the mechanism predicts: qkv's 12 KB stride is +the worse of the two. + +## 3. The zero-memory alternative, tried first — negative + +Paying `2*K*N` bytes for a transpose is a real cost, so software prefetch was +tried before accepting it. `_mm_prefetch` at distance 12 into the strided inner +loop, on qkv: + +| variant | us | +|---|---| +| `kn`, no prefetch | **5022** | +| `kn`, `_mm_prefetch` distance 12 | 5580 | + +**Worse.** The stride penalty is not prefetch-recoverable here — the extra +requests add pressure without arriving early enough to matter. That negative is +what justifies spending the memory. + +## 4. Accuracy moves the same way, so there is no trade + +`kn` carries **one serial accumulator per column** across the whole contraction. +`nk` carries four, combined pairwise. Against an f64 oracle, `nk` is +**2.7-9.3x more accurate** across the tested shapes. + +The slow side was also the less accurate side. There was no tradeoff to weigh — +which is worth stating explicitly, because a layout change that reassociates a +reduction usually does have one. + +**The first version of this test was worthless and looked fine.** It built +operands from `*0.125` and `*0.0625` — exactly representable, so every partial +sum was exact — and reported **zero error and 100% bit-identity**. It could not +detect the effect it existed to measure, and it reported that confidently. +Rebuilt on hostile xorshift data, bit-identity is ~3%. + +This is the same class of failure as §23's instrumented baseline: **check that +the measurement can see the effect at all before believing a clean result.** + +## 5. Production-path A/B + +Two builds of `benches/half_decode_gemv_ab.rs` — one from an `origin/main` +worktree, one from the branch — through the shipped routing, `taskset`-pinned, +`steady_ms`: + +| dtype | shape | before | after | speedup | +|---|---|---|---|---| +| **f32 (null control)** | attn_out 1024x768 | 0.068 | 0.067 | 1.01x | +| **f32 (null control)** | square 2048x2048 | 0.094 | 0.100 | 0.94x | +| **f32 (null control)** | mlp 4096x11008 | 6.605 | 6.628 | 1.00x | +| **f32 (null control)** | lm_head 896x151936 | 14.996 | 14.833 | 1.01x | +| f16 | attn_out 1024x768 | 0.063 | 0.028 | **2.25x** | +| f16 | square 2048x2048 | 0.314 | 0.058 | **5.41x** | +| f16 | mlp 4096x11008 | 2.867 | 1.791 | **1.60x** | +| f16 | lm_head 896x151936 | 8.553 | 7.089 | **1.21x** | +| bf16 | attn_out 1024x768 | 0.077 | 0.026 | **2.96x** | +| bf16 | square 2048x2048 | 0.320 | 0.058 | **5.52x** | +| bf16 | mlp 4096x11008 | 2.782 | 1.719 | **1.62x** | +| bf16 | lm_head 896x151936 | 8.635 | 7.064 | **1.22x** | + +**The f32 rows are the null control** and move 0.94-1.01x, which is the evidence +that the harness and the host are quiet and that nothing outside the 16-bit +routing changed. + +**Read the small numbers, not the big one.** `square 2048x2048` at 5.41x/5.52x +is a shape nobody runs. The large model-shaped rows are 1.2-1.6x. `lm_head` — +which gains **least**, at 1.21x — is also the row that pays **most**: 272 MB +resident for one transposed weight. + +`max_rel` is unchanged on every row. Two bf16 rows keep a bit-identical digest: +bf16's coarser mantissa rounds the reassociation to the same bits on that data, +which is expected and is not evidence that the reassociation did not happen (see +§4). + +## 6. The memory-plan coupling — the part that could have gone badly + +`node_weight_transpose_cache_bytes` is the predictor `engine/load.rs` budgets +against under #1056. For `MatMul` it was `cfg(any(target_os = "macos", +target_os = "ios"))`. + +On x86 this transpose would therefore have been **completely invisible to the +memory plan** — gigabytes of retained weight buffers that the plan did not know +about, on the most common operator in the graph. The predictor was rewritten so +the x86 case is predicted (`Gemm` transB=0, and `MatMul` with a constant 16-bit +B, at `2*numel`), with the Apple arm preserved. + +**Rule this establishes: any change that makes a kernel retain a weight-scaled +buffer must update that predictor in the same commit.** + +`FusedMatMulBias` is deliberately **excluded** from the predictor, because it +takes no x86 16-bit GEMV — budgeting it would over-reserve `2*K*N` for every +fused projection that never allocates one. That exclusion **inverts the moment +the GEMV is enabled**; #1702 carries the warning. + +Two pre-existing guard tests encoded the opposite decision ("a transposed +variant would cost a permanent `2*K*N` bytes"). They were not deleted. They now +assert the stronger invariant they were reaching for: no **unbudgeted** copy, +and never an f32 widening, checked against the predictor itself. + +## 7. Admission is no longer numerically neutral + +Declining the transpose cache now changes **which kernel runs**, and therefore +output bits. Three contract comments still asserted neutrality — including one +that described the exact opposite of the call directly beneath it. All three +corrected. Adversarial review caught that the commit message disclosed this but +the comments a maintainer actually reads did not. + +## 8. A real bug, introduced and fixed here + +`WEIGHT_TRANSPOSE_F16` stores raw `u16` keyed `(addr, k, n)` with **no dtype +discriminator** — safe only while exactly one dtype used it. + +Routing bf16 through the same cache let a bf16 weight hit an f16 entry left +behind at a recycled address: `-0.000021640852 != -0.8984375`. It reproduced +**only in company** — never when the test ran alone — because it needs a prior +allocation to recycle. + +Guarding the *view* dtype is not sufficient; the **key** needs the +discriminator. Fixed by adding `tag: u16` to `WeightTransposeKey`, with +`the_two_16_bit_formats_do_not_share_a_cache_entry` as the regression. + +Related standing trap: tests depending on the cache verdict must use +`weight_transpose::CacheEnabledScope` (thread-local RAII), never +`set_cache_enabled` (process-global) — the documented #983/#1033/#1056 "passes +alone, fails in company" failure. + +## 9. Still divergent, deliberately + +`FusedMatMulBias` takes no 16-bit GEMV on x86: **2845 us** on qkv against +`MatMul`'s **1830 us** after this change, same weights and shape. It is a +separate mechanism with its own memory-plan consequence and a real bias-epilogue +difference, so it is not folded in here. Split to **#1702**. diff --git a/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md new file mode 100644 index 0000000000..35e8b8f85e --- /dev/null +++ b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md @@ -0,0 +1,128 @@ +# int4 `accuracy_level = 0` decode: the route the production default actually +# takes, and the correction to its attribution + +**Date:** 2026-08-21 · **Owner:** Roy (CPU MatMul) · **Host:** AMD EPYC 9V74, +32 vCPU (16c x 2 SMT), AVX2/FMA/F16C, **no AVX-512/VNNI**. +Ledger entry: §23 of [`CPU_MATMUL_ASSIGNMENT.md`](../performance/CPU_MATMUL_ASSIGNMENT.md). +Merged as `99f105d52` (#1679), for #1676. + +Every timing here is `taskset -c 0,2,4,...`-pinned to distinct physical cores. +Unpinned multi-thread numbers on this host measure worker placement, not the +kernel — see the [placement record](2026-08-21-decode-worker-cpu-placement.md). + +--- + +## 1. The route, by counter + +The question was where the 1.84x acc0 gap sits. The first thing to establish is +which kernel the production default even runs. Route counters instrumented from +operator entry through to the innermost arm, over a real decode step at +`accuracy_level = 0`: + +| counter | count | +|---|---| +| `entry_bits4` | 95 | +| `percolumn` | 95 | +| `nblock` | **0** | +| `block_simd` | 129,499,136 | + +`nblock = 0`. The register-blocked N-blocked kernel added by #1104 — measured +there at 1.46x on a 14B model and proven byte-identical — **was never reached at +the production default**, from the day it merged. + +The cause is one line. #1104 shipped it behind +`ONNX_GENAI_CPU_MM_INT4_NBLK`, defaulted `false`, with the comment that this was +"until the win is measured, exactly like the toggles that preceded it". The +measurement did not happen. Nothing in the tree asserted which route production +took, so the toggle's expiry condition was invisible. + +**Fix:** default `true`, plus +`acc0_decode_reaches_the_nblocked_kernel_by_default`, which reads the counters +and fails if the default route changes. The default is now a checked property. + +## 2. The first attribution was wrong, and the instrument is why + +The initial A/B reported 3.14x (t=1) and 4.80x (t=4) for the N-blocked route. +Those numbers are **void**. + +The route probe is an `AtomicU64::fetch_add`. In the per-column arm it sits +inside the block loop and fired **129,499,136 times** per measurement window; in +the N-blocked arm it fired 95 times. The probe inflated the *baseline it was +measuring against* by roughly 3x, and — because it did so consistently across +repetitions — produced a stable, self-consistent, wrong answer. + +This is distinct from §18's probe, which was genuine kernel overhead: a +per-block `is_x86_feature_detected!` that broke inlining and forced the +accumulator through memory, so removing it made the *shipping* kernel faster. +This one never affected production at all. It only corrupted its own baseline. + +Rebuilt with the probe out of the timed path, counters read once at the end: + +| route | t=1 (ms) | t=4 (ms) | vs per-column | +|---|---|---|---| +| per-column (previous default) | 56.528 | 28.518 | 1.00x | +| N-blocked, group of 1 | 55.156 | 27.861 | 1.02x | +| N-blocked, group of 2 | 47.679 | 23.885 | 1.19x | +| N-blocked, group of 4 | **38.114** | **19.238** | **1.48x** | + +Ratios are from the t=1 column; t=4 agrees to within 0.01x on every row, which +is the check that this is a per-core kernel effect and not a scheduling one. + +## 3. What the win is, and what it is not + +The group-of-1 row is the control that matters, and it is the reason the obvious +explanation is wrong. + +The per-column path carries a horizontal reduction on the critical path: +`extractf128` -> `movehl` -> `shuffle`, each dependent on the last, once per +32-weight block, every four FMAs. The N-blocked kernel removes it — it keeps the +scale in a vector accumulator and reduces once per column instead. That is a +real latency chain, it is textbook, and it is the thing you would name if asked +to predict the win. + +**It is worth 1.02x. Nothing measurable.** Group-of-1 is the N-blocked kernel +with the reduction restructured and *no* column grouping, and it lands on top of +the baseline. The block loop carries enough independent work for the +out-of-order engine to hide the chain entirely. + +The entire win is **four-column activation reuse**: 1.45x from group 1 to group +4 (55.156 -> 38.114). Each activation block is loaded once and used against four +weight columns. + +Two independent confirmations that this is the right decomposition: the +group-2 row sits where a load-amortisation model predicts (1.19x, roughly the +square-root-ish midpoint rather than half the gain), and the group-4 total, +1.48x, matches #1104's independently measured 1.46x on entirely different +hardware and a different model. + +## 4. Numerics: the trade, stated + +The N-blocked kernel applies the block scale as a separated correction rather +than folding it per block. Against an f64 oracle, worst cell measured: + +| route | worst relative error vs f64 | +|---|---| +| per-column | 6.739e-6 | +| N-blocked, group of 4 | **2.422e-5** | + +**3.59x worse.** It is shipped anyway, and disclosed rather than buried, on +three grounds: it remains inside the pinned accuracy envelope; the envelope +guard was sized to this *measurement* (8x headroom) rather than to hope, so a +future regression trips it; and a 1.48x decode win for a 3.59x relative error +increase inside an envelope is a defensible trade **only if both numbers are +visible to whoever inherits it**. + +`accuracy_level = 0`'s exact contract is unaffected — the contract is the +envelope, and this stays inside it. No reduced-precision diversion is involved; +this is f32 accumulation throughout. + +## 5. What this does not close + +The acc0 gap against ORT is not closed by this. The dormant kernel was a free +1.48x that had been sitting in the tree unclaimed, so it is the correct first +move, but the remaining gap after it is a separate mechanism and is not +attributed here. + +**Negative-result note for the next person:** do not spend time on the +horizontal reduction in the acc0 int4 path. It is measured at 1.02x. Column +grouping is where the arithmetic intensity is. diff --git a/docs/performance/CPU_MATMUL_ASSIGNMENT.md b/docs/performance/CPU_MATMUL_ASSIGNMENT.md index 0a4d829b63..c1b270d6d5 100644 --- a/docs/performance/CPU_MATMUL_ASSIGNMENT.md +++ b/docs/performance/CPU_MATMUL_ASSIGNMENT.md @@ -1895,3 +1895,202 @@ progress on acc0. Full record: [`docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md`](../benchmarks/2026-08-21-int4-acc4-execution-regime.md). + +### 23. The register-blocked int4 decode kernel shipped default-off and stayed dormant for its whole life (**fixed**) + +`accuracy_level = 0` is the production default, and route counters instrumented +from operator entry through the kernel said it never reached the N-blocked +kernel: `entry_bits4=95 percolumn=95 nblock=0 block_simd=129,499,136`. +#1104 built that kernel, measured 1.46x on a 14B model, proved it +byte-identical, and then shipped it behind a **default-off** env toggle "until +the win is measured, exactly like the toggles that preceded it". The +measurement never happened. Nothing asserted which route production took, so a +finished, proven kernel sat unreachable in the tree while the default path took +the per-column loop. + +The repair is one line of default (`unwrap_or(false)` -> `unwrap_or(true)`) plus +`acc0_decode_reaches_the_nblocked_kernel_by_default`, which makes the default +route a **checked property** rather than a comment. Merged as `99f105d52`. + +**The first attribution I produced was wrong, and the instrument is what was +wrong.** The route probe's `fetch_add` fired 129 million times in the per-column +arm and zero in the N-blocked arm, so the counter inflated the baseline it was +supposed to measure by ~3x and produced a self-consistent 3.14x/4.80x. Rebuilt +without the probe in the timed path: + +| route | t=1 | t=4 | vs per-column | +|---|---|---|---| +| per-column (previous default) | 56.528 | 28.518 | 1.00x | +| N-blocked, group of 1 | 55.156 | 27.861 | 1.02x | +| N-blocked, group of 2 | 47.679 | 23.885 | 1.19x | +| N-blocked, group of 4 | 38.114 | 19.238 | **1.48x** | + +**The group-of-1 row rules out the explanation the structure suggests.** +Restructuring the reduction — keeping the scale in a vector accumulator and +reducing once per column instead of once per 32-weight block — is worth +**1.02x, i.e. nothing measurable**, even though the per-column path's hreduce +(`extractf128`/`movehl`/`shuffle`, each dependent on the last) sits on the +critical path every four FMAs. The block loop has enough independent work for +the out-of-order engine to hide it. The entire win is the **four-column +activation reuse**: 1.45x from group 1 to group 4, which also matches #1104's +independently measured 1.46x. + +**The numerics move the wrong way and are shipped anyway, disclosed.** The +N-blocked kernel's separated correction is **3.59x worse** relatively against +f64 on the worst cell measured (2.422e-5 vs 6.739e-6). It stays within the +pinned envelope, the envelope guard was sized to the *measurement* (8x) rather +than to hope, and the tradeoff is stated rather than buried — a 1.48x decode win +for a 3.59x relative error increase inside an envelope is a defensible trade +only if both numbers are on the table. + +Two lessons. An instrument in the timed path is a **measurement error, not +overhead** — distinct from §18's probe, which was real kernel overhead that +broke inlining and whose removal made the shipping kernel genuinely faster; this +one changed nothing in production and only corrupted its own baseline. And +"default off until measured" is a decision that **expires silently** unless a +test asserts the default. + +Full record: +[`docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md`](../benchmarks/2026-08-21-int4-acc0-dormant-nblock.md). + +### 24. The t=8 "wash" is worker-to-CPU placement, not the kernel (**negative result; runtime-owned**) + +The premise handed to me was that #1628's int4 acc4 win vanishes at t=8 while +holding at t=1/4/16, and that the kernel should be looked at. **It does not +reproduce.** Measured against explicit pool widths rather than a thread-count +label, the win is flat through width 8 and collapses *after* it: + +| pool width | 1 | 4 | 8 | 12 | 16 | +|---|---|---|---|---|---| +| speedup | 1.66x | 1.65x | 1.66x | 1.238x | 0.993x | + +So there is no t=8 anomaly to tune for. There is a **width-12-and-above** +collapse, and it is not the kernel. + +Two hypotheses were tried and discarded before the right one. Memory bandwidth: +a STREAM-style all-thread sweep read 83 GB/s against a 41 GB/s decode draw — +but that figure **does not reconcile with §22**, which measured this host at +31-36 GB/s within a CCX and ~56.6 GB/s across both, and §22's numbers are the +ones this file stands behind. Against those, 41 GB/s is 72% of the across-CCX +ceiling and *above* the within-CCX one, so bandwidth cannot be dismissed by my +83 GB/s sweep and is not dismissed here on that basis. It is ruled out by the +placement A/B below instead, which holds shapes, bytes, thread count and binary +constant and moves **only** which CPUs the workers sit on — a bandwidth ceiling +does not care about that, and the result does. Task grain: the shard count and +barrier time move as expected. + +**Root cause: `decode_spmd.rs::node_shards` pins worker *i* to +`allowed_cpus()[i]` in logical order.** On this host SMT siblings are adjacent +(CPUs 0 and 1 are the two threads of core 0), so 16 workers land on CPUs 0-15, +which is **8 physical cores**, and half of them contend for a sibling's +execution units. Verified by reading `/proc//task/*/stat`, and confirmed +decisively by comparison: default placement 0.982x versus one-worker-per- +physical-core 1.225x on the same binary and the same shapes. + +**No kernel change was made**, which is the point. Filed as **#1680** with the +measurement and the placement evidence, for the runtime owner. Every timing in +§23 and §25 is `taskset`-pinned to even CPUs as a consequence — an unpinned +multi-thread number on this host is measuring the scheduler, not the kernel. + +Full record: +[`docs/benchmarks/2026-08-21-decode-worker-cpu-placement.md`](../benchmarks/2026-08-21-decode-worker-cpu-placement.md). + +### 25. `f16`/`bf16` decode diverged by *layout*, and the slow side was also the less accurate one (**fixed**) + +#1381's dispatch comment claimed the divergence was closed: "both operators, +both stored orders and both 16-bit formats reach the same GEMV backend". Same +**backend**, not same **kernel**. Enumerated with route counters: + +| operator | stored order | kernel taken | +|---|---|---| +| `Gemm` transB=1 | `[N,K]` | `gemv_half_nk` | +| `Gemm` transB=0 | `[K,N]` | `gemv_half_kn` | +| `MatMul` | `[K,N]` | `gemv_half_kn` | +| `FusedMatMulBias` | `[K,N]` | **no 16-bit GEMV at all** | + +The fourth row was found empirically, not inferred — `count_half_decode_gemv()` +lives in `MatMulKernel::execute_with_backend` and `fused_matmul_bias.rs` calls +the free `matmul_dense_prepacked_into`, so a probe read `matmul_gemv=1 +fused_gemv=0`. It matters because the optimizer fuses `MatMul + Add(bias)`. The +MatMul-side transposed row had **no test**: `decode_matmul` only ever built B as +`[k,n]` — the fifth unmeasured-region-behind-a-gate this file records (cf. §11, +§12, §14, §19). + +**`[K,N]` crosses a page every `p`.** The stride between consecutive `p` is +`n*2` bytes — 12 KB at n=6144 — and the L2 prefetcher does not cross page +boundaries. Direct kernel A/B on identical bytes, same call, only the stored +order differing: + +| shape | `gemv_half_kn` us | `gemv_half_nk` us | penalty | +|---|---|---|---| +| qkv 4096x6144 | 5022 | 1688 | **2.98x** | +| down 14336x4096 | 11015 | 7068 | **1.56x** | + +**The zero-memory alternative was tried first and is a negative result.** +`_mm_prefetch` at distance 12 into the strided inner loop made `kn` *worse* +(5580 vs 5022 us on qkv). The stride penalty is not prefetch-recoverable, which +is what justifies paying `2*K*N` bytes for a transpose. + +**Accuracy moves the same way, so there is no trade to weigh.** `kn` carries one +serial accumulator per column across the whole contraction; `nk` carries four +combined pairwise. Against an f64 oracle, `nk` is **2.7-9.3x** more accurate. +The first version of that test used `*0.125` operands — exactly representable, +every partial sum exact — and reported zero error and 100% bit-identity. It +could not detect the effect it existed to measure. Hostile data shows ~3% +bit-identity. That is the same failure mode as the instrument in §23: **check +that the measurement can see the effect at all.** + +Production-path A/B, two builds, `taskset`-pinned, `steady_ms`: + +| dtype | shape | before | after | speedup | +|---|---|---|---|---| +| **f32 (null control)** | attn_out 1024x768 | 0.068 | 0.067 | 1.01x | +| **f32 (null control)** | mlp 4096x11008 | 6.605 | 6.628 | 1.00x | +| f16 | attn_out 1024x768 | 0.063 | 0.028 | **2.25x** | +| f16 | square 2048x2048 | 0.314 | 0.058 | **5.41x** | +| f16 | mlp 4096x11008 | 2.867 | 1.791 | **1.60x** | +| f16 | lm_head 896x151936 | 8.553 | 7.089 | **1.21x** | +| bf16 | attn_out 1024x768 | 0.077 | 0.026 | **2.96x** | +| bf16 | mlp 4096x11008 | 2.782 | 1.719 | **1.62x** | +| bf16 | lm_head 896x151936 | 8.635 | 7.064 | **1.22x** | + +**The best row is not the interesting one.** `square` at 5.4x is a shape nobody +runs; the large model-shaped rows (`mlp`, `lm_head`) are 1.2-1.6x, and +`lm_head` — which gains least at 1.21x — is also the row that pays most, +**272 MB** resident for one weight. The small `attn_out` projection sits between +them at 2.25x/2.96x. + +**The memory-plan coupling is the part that could have gone badly.** +`node_weight_transpose_cache_bytes` is what `engine/load.rs` budgets against +under #1056, and it was `cfg(macos/ios)` for `MatMul`. On x86 this transpose +would have been **completely invisible to the plan**, which would have +under-budgeted by gigabytes. Any change that makes a kernel retain a +weight-scaled buffer must update that predictor. `FusedMatMulBias` is +deliberately excluded — it has no x86 16-bit GEMV, so budgeting it would +over-reserve every fused projection. + +**Two guard tests encoded the opposite decision** ("a transposed variant would +cost a permanent 2*K*N bytes"). They were not deleted: they now assert the +stronger invariant they were reaching for — no **unbudgeted** copy, and never an +f32 widening, checked against the predictor itself. + +**Admission is no longer numerically neutral on this path.** Declining the cache +changes *which kernel* runs and therefore output bits. Three contract comments +still claimed neutrality and were corrected; adversarial review caught that the +commit message disclosed it but the comments a maintainer actually reads did +not. + +**A real bug, introduced and fixed here.** The f16 transpose cache stores raw +`u16` keyed `(addr, k, n)` with **no dtype** — safe only while one dtype used +it. Routing bf16 through it let a bf16 weight hit an f16 entry left at a +recycled address: `-0.000021640852 != -0.8984375`, reproducible **only in +company**. Guarding the view dtype is not enough; the *key* needs the +discriminator. + +**Still divergent, deliberately:** `FusedMatMulBias` takes no 16-bit GEMV on +x86 — 2845 us on qkv against MatMul's 1830 us after this change. It is a +separate mechanism with its own memory-plan consequence and is not folded in +here. Merged as `2e1cfb67c`. + +Full record: +[`docs/benchmarks/2026-08-21-half-decode-layout-divergence.md`](../benchmarks/2026-08-21-half-decode-layout-divergence.md).