From 09eca60c81620d8e9f5bb8c4ca4954d8b6451dc5 Mon Sep 17 00:00:00 2001 From: Justin Chu Date: Sun, 23 Aug 2026 14:39:51 +0000 Subject: [PATCH 1/3] =?UTF-8?q?docs(perf):=20the=201.84x=20acc0=20gap=20is?= =?UTF-8?q?=20stale=20=E2=80=94=20re-measured=20it=20is=201.12x=20at=20t?= =?UTF-8?q?=3D1/4/8?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The acc0 int4 decode gap against ORT has been the top remaining CPU MatMulNBits target on the strength of a `1.84x` figure. Re-measured on `e189244ba` with a matched core budget, a checked realized width and three independent launches per width, it is **1.120x at t=1, 1.148x at t=4 and 1.120x at t=8** — flat across the measurable range. The old figure was not mislabelled. It was a correct measurement of a tree that no longer exists: six merges landed after `e9754e7ef` published it, three of them direct acc0 kernel work (#1667 broke the serial f32 reduction chain, 5.75x t=1; #1679 enabled the register-blocked kernel *at accuracy_level = 0*; #1783 folded the zero-point unpack). The control is the ORT arm, which reproduces to +4.4% (30.632 -> 31.99 ms) on the same host, binary, graph and statistic — so the comparison is sound and the whole movement is on our side: native 56.307 -> 35.36 at t=1 and 14.091 -> 4.57 at t=8. A yesterday's note on this same figure said it was `t=1`-only and that the production-width gap was unmeasured; both were true, and both missed that the `t=1` number was itself three kernel merges old. Three harness defects had to be fixed before the number meant anything, all in `acc0_gap_matrix.py`: * **The arms were not getting the same machine.** `ONNX_GENAI_CPU_DECODE_THREADS=w` confines the whole native process to `w` CPUs; the script pinned ORT to all 16 at every thread count. Measured in both configurations, the asymmetry is only 1-2% on a quiet host, which is recorded so it is not re-litigated. * **The realized width was assumed, not checked.** A cell is now refused unless the binary reports `as_requested` (width 1's `path=flat` excepted -- it builds no pool by design). * **A pre-check cannot see a competitor that arrives mid-cell.** A sibling `cargo test` started during this matrix and four cells that had passed `wait_quiet` were measured against it. Every arm now runs inside a `LoadWatch` sampling the instantaneous runnable count, refused above `threads + slack`. t=16 remains unresolved: every cell at that width was contaminated, and it is the same width whose launch distribution spans 514% with no known mechanism. Full record: docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- .../benches/acc0_gap_matrix.py | 182 ++++++++++++++-- .../2026-08-21-int4-acc0-dormant-nblock.md | 15 +- .../2026-08-21-int4-acc4-execution-regime.md | 45 ++-- .../2026-08-23-acc0-gap-vs-ort-by-width.md | 205 ++++++++++++++++++ docs/performance/CPU_MATMUL_ASSIGNMENT.md | 57 +++-- 5 files changed, 444 insertions(+), 60 deletions(-) create mode 100644 docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md diff --git a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py index fcb9bee7a4..7c4c6616a4 100755 --- a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py +++ b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py @@ -52,6 +52,19 @@ Separating the two requires re-running both arms interleaved on a quiet host, which is what `wait_quiet` and `competing_load` below exist to guarantee. Until that has been done this cell has no published ratio. + +Two preconditions added 2026-08-23, both of which this script previously +assumed rather than checked +--------------------------------------------------------------------------- +* **The two arms must get the same machine.** `ONNX_GENAI_CPU_DECODE_THREADS=w` + confines the *whole native process* to `w` CPUs (it prints so). Pinning ORT + to all 16 even CPUs, as this script did at every thread count, gave ORT a + 16-core machine while native had `w` -- more L3 and more memory controllers, + on a workload that is bandwidth-bound by construction. See `native_pin`. +* **The native width must be non-vacuous.** The realized width is read back + from the binary and a cell is refused unless it equals the request. Timings + cannot detect this: a sweep that silently runs one width in every row looks + perfectly stable, because it is -- it is the same configuration each time. """ import argparse import json @@ -59,6 +72,7 @@ import re import subprocess import sys +import threading import time HERE = os.path.dirname(os.path.abspath(__file__)) @@ -71,7 +85,27 @@ } # Even CPUs are distinct physical cores on this SMT host (siblings adjacent). -PIN = ",".join(str(c) for c in range(0, 32, 2)) +EVEN = list(range(0, 32, 2)) +PIN = ",".join(str(c) for c in EVEN) + + +def native_pin(threads): + """The CPUs the *native* arm actually gets at this thread count. + + `ONNX_GENAI_CPU_DECODE_THREADS=w` does not merely size a pool: it confines + the whole process to `w` CPUs, and says so on stderr -- + + CPU decode budget 4 confined the process to 4 CPUs [0, 2, 4, 6] + + So pinning both arms to all 16 even CPUs, as this script used to, hands ORT + a 16-core machine while native has `w`. That is not a thread-count + comparison, it is a machine-size comparison, and at every `w < 16` it + flatters ORT with more L3 and more memory controllers than native can + reach. It went unnoticed because the only published gap cell was `t=1`, + where ORT's `intra_op_num_threads=1` means one thread regardless, and + because at `t=16` the two pins coincide exactly. + """ + return ",".join(str(c) for c in EVEN[:threads]) def weight_bytes(model, block): @@ -107,17 +141,37 @@ def native(binary, model, block, acc, threads, sessions, tokens, reps, extra=Non if extra: env.update(extra) r = sh(f"taskset -c {PIN} {binary}", env) - for line in r.stdout.splitlines(): + steady, width_line = None, None + for line in r.stdout.splitlines() + r.stderr.splitlines(): if line.strip().startswith("steady"): f = line.split() - return {"ms_token": float(f[1]), "p90": float(f[2]), - "tps": float(f[3]), "spread": float(f[4])} - sys.stderr.write(r.stdout + r.stderr) - raise RuntimeError("native arm produced no steady row") - - -def ort(model, block, acc, threads, sessions, tokens, reps): - cmd = (f"taskset -c {PIN} python3 ort_matmulnbits_baseline.py " + steady = {"ms_token": float(f[1]), "p90": float(f[2]), + "tps": float(f[3]), "spread": float(f[4])} + if line.strip().startswith("decode_width"): + width_line = line.strip() + if steady is None: + sys.stderr.write(r.stdout + r.stderr) + raise RuntimeError("native arm produced no steady row") + # Non-vacuity, checked rather than assumed. A width sweep whose rows all + # report the same width is the failure that produced the retracted + # "t=1 == t=2" claim, and it is invisible in the timings themselves -- + # they look stable, they are just all the same configuration. The binary + # reads the realized width back out of the pool, so ask it. + if width_line is None: + raise RuntimeError("native arm did not report decode_width") + steady["decode_width"] = width_line + if not width_line.endswith("as_requested"): + # Width 1 legitimately builds no pool at all (`allowed.len() == 1` + # declines), so `path=flat` there is correct, not a reduction. Any + # other mismatch invalidates the row's label. + if not (threads == 1 and "path=flat" in width_line): + raise RuntimeError(f"native width vacuous: {width_line}") + return steady + + +def ort(model, block, acc, threads, sessions, tokens, reps, pin=None): + pin = pin or native_pin(threads) + cmd = (f"taskset -c {pin} python3 ort_matmulnbits_baseline.py " f"--model {model} --block {block} --accuracy {acc} --threads {threads} " f"--tokens {tokens} --reps {reps} --sessions {sessions}") r = subprocess.run(cmd, shell=True, capture_output=True, text=True, cwd=BENCH, @@ -126,7 +180,7 @@ def ort(model, block, acc, threads, sessions, tokens, reps): if not m: sys.stderr.write(r.stdout + r.stderr) raise RuntimeError("ORT arm produced no throughput") - return {"tps": float(m.group(1)), "spread": float(m.group(2))} + return {"tps": float(m.group(1)), "spread": float(m.group(2)), "pin": pin} def competing_load(): @@ -181,6 +235,61 @@ def wait_quiet(threshold=3.0, limit=900): return os.getloadavg()[0], competing_load() +class LoadWatch: + """Peak runnable count *during* an arm, not merely before it. + + `wait_quiet` is a pre-check, and a pre-check cannot see a competitor that + starts after the cell does. That is not hypothetical either: a sibling + agent's `cargo test` began mid-matrix and four cells that had passed the + pre-check were measured against a saturated box, at spreads of 20-63%. + + Two details matter. The instantaneous **runnable** count (field 4 of + `/proc/loadavg`) is used rather than load average, which is a 1-minute + exponential average that both lags a job that just started and stays high + long after one ends. And the threshold scales with the thread count, + because at `t` threads our own arm legitimately contributes ~`t` runnable + threads -- a constant like "runnable > 4" would refuse every honest cell + at `t >= 8`. + + It is also worth being explicit that this is a *necessary*, not a + sufficient, condition. Ten launches at width 16 split into a fast and a + slow mode 1.8x apart in wall time while burning identical CPU-seconds + (14.4 vs 14.1 CPU-s per wall-s): the affected threads were never + descheduled, they just retired fewer instructions per cycle. No load, + CPU-efficiency or context-switch guard can see that. + """ + + def __init__(self, period=1.0): + self.period = period + self.peak = 0 + self._stop = threading.Event() + self._thread = None + + @staticmethod + def runnable(): + try: + with open("/proc/loadavg") as f: + return int(f.read().split()[3].split("/")[0]) + except Exception: + return -1 + + def _loop(self): + while not self._stop.is_set(): + self.peak = max(self.peak, self.runnable()) + self._stop.wait(self.period) + + def __enter__(self): + self.peak = self.runnable() + self._thread = threading.Thread(target=self._loop, daemon=True) + self._thread.start() + return self + + def __exit__(self, *exc): + self._stop.set() + self._thread.join(timeout=2 * self.period) + return False + + def main(): ap = argparse.ArgumentParser() ap.add_argument("--binary", required=True) @@ -193,6 +302,15 @@ def main(): ap.add_argument("--reps", type=int, default=3) ap.add_argument("--out", default="acc0_gap.json") ap.add_argument("--aa", action="store_true", help="per-cell interleaved A/A") + ap.add_argument("--slack", type=int, default=4, + help="peak runnable count above `--threads` that still " + "counts as a quiet host during a cell") + ap.add_argument("--ort-pin", choices=["matched", "wide", "both"], + default="matched", + help="CPUs for the ORT arm: the same `t` the native process " + "confines itself to (matched, the only comparison), all " + "16 physical cores (wide, what this script used to do), " + "or both so the asymmetry is quantified") args = ap.parse_args() rows = [] @@ -214,29 +332,55 @@ def main(): for pid, pcpu, cmd in busy[:3]: sys.stderr.write(f" pid={pid} cpu={pcpu:.0f}% {cmd}\n") # Interleaved: native, ORT, native again (the A/A partner). - a1 = native(args.binary, model, args.block, args.acc, t, s, - args.tokens, args.reps) - o = ort(model, args.block, args.acc, t, s, args.tokens, args.reps) - aa = "" - if args.aa: - a2 = native(args.binary, model, args.block, args.acc, t, s, + # Every arm runs inside a LoadWatch so a competitor that + # arrives mid-cell is caught, not just one that was already + # there when `wait_quiet` returned. + with LoadWatch() as watch: + a1 = native(args.binary, model, args.block, args.acc, t, s, args.tokens, args.reps) - aa = f"{a2['tps'] / a1['tps']:.3f}" + o = None + o_wide = None + if args.ort_pin in ("matched", "both"): + o = ort(model, args.block, args.acc, t, s, args.tokens, + args.reps, pin=native_pin(t)) + if args.ort_pin in ("wide", "both"): + o_wide = ort(model, args.block, args.acc, t, s, + args.tokens, args.reps, pin=PIN) + if o is None: + o = o_wide + aa = "" + if args.aa: + a2 = native(args.binary, model, args.block, args.acc, t, + s, args.tokens, args.reps) + aa = f"{a2['tps'] / a1['tps']:.3f}" + if watch.peak > t + args.slack: + sys.stderr.write( + f"WARNING {model} t={t} s={s}: peak runnable " + f"{watch.peak} > {t} + {args.slack} during the cell; " + f"cell is UNTRUSTED\n") + busy = busy or [("-", 0.0, f"peak runnable {watch.peak}")] nat_bw = a1["tps"] * wb / 1e9 ort_bw = o["tps"] * wb / 1e9 row = {"model": model, "threads": t, "sessions": s, "native": a1, "ort": o, "ratio": a1["tps"] / o["tps"], + "ort_wide": o_wide, + "ratio_wide": (a1["tps"] / o_wide["tps"]) if o_wide else None, "native_gbs": nat_bw, "ort_gbs": ort_bw, "weight_bytes": wb, "loadavg_at_start": load, + "peak_runnable": watch.peak, "trusted": not busy, "competitors": [c[2] for c in busy], "aa": float(aa) if aa else None} rows.append(row) flag = "" if not busy else " !CONTENDED" + wide = "" + if o_wide is not None and row["ratio_wide"] is not None: + wide = (f" wide_ort={o_wide['tps']:.1f} " + f"ratio_wide={row['ratio_wide']:.3f}") print(f"{model:>6} {t:>3} {s:>2} {a1['tps']:>9.1f} " f"{a1['spread']:>8.1f} {o['tps']:>9.1f} {o['spread']:>8.1f} " f"{row['ratio']:>7.3f} {nat_bw:>9.1f} {ort_bw:>9.1f} {aa:>6}" - f"{flag}") + f"{wide}{flag}") sys.stdout.flush() with open(os.path.join(HERE, args.out), "w") as f: json.dump(rows, f, indent=1) diff --git a/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md index 0e589813fe..f562b54ecf 100644 --- a/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md +++ b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md @@ -19,12 +19,15 @@ which kernel the production default even runs. Route counters instrumented from operator entry through to the innermost arm, over a real decode step at `accuracy_level = 0`: -> **Note (2026-08-23):** the `1.84x` inherited here is a **`threads = 1`** -> figure — the only acc0 cell with an ORT baseline — and the acc0 gap at -> production width is unmeasured. The route findings below are categorical -> (which kernel runs) and are unaffected by that; only the *sizing* of the gap -> is. See the scope note in -> [2026-08-21-int4-acc4-execution-regime.md](2026-08-21-int4-acc4-execution-regime.md). +> **Note (2026-08-23, revised):** the `1.84x` inherited here **no longer +> holds** — re-measured on `e189244ba` the acc0 gap is **1.12x at t=1, 1.15x at +> t=4, 1.12x at t=8**. Part of that closure is the work this very document +> motivated: enabling the register-blocked kernel at `accuracy_level = 0` +> (#1679) was one of three acc0 merges that landed after the 1.84x was taken. +> The route findings below are categorical (which kernel runs) and stand +> unchanged; only the *sizing* of the gap moved, and it moved because the +> dormant kernel was woken up. See +> [2026-08-23-acc0-gap-vs-ort-by-width.md](2026-08-23-acc0-gap-vs-ort-by-width.md). | counter | count | |---|---| diff --git a/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md b/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md index 4f87b2ebcd..d13157b263 100644 --- a/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md +++ b/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md @@ -475,31 +475,38 @@ than the 1.79x row.** I am not claiming parity at t=16 on this evidence. ### The default path is untouched and is now the worse gap This kernel is gated to `accuracy_level = 4`. Production default is -`accuracy_level = 0`, and measuring both sides there on the same host: +`accuracy_level = 0`, and measuring both sides there on the same host +(**this table is superseded — see the note directly below it**): | threads | ORT acc0 | native acc0 | gap | |---|---|---|---| | 1 | 30.632 ms/tok | 56.307 | **1.84x** | | 8 | — | 14.091 | — | -> **Scope note (2026-08-23).** The `1.84x` is a **`threads = 1`** figure and is -> the only acc0 cell with an ORT baseline — the t=8 row has no ORT number, so -> **the acc0 gap at production width is unmeasured**. This is not a pedantic -> label: on the acc4 table in this same document the gap moves from 3.01x at -> t=1 to 1.47x at t=16, so width is among the largest effects here and 1.84x -> should not be assumed to survive to t=8/16. At t=1 the native side also runs -> `path=flat`, confined by the decode budget to one CPU with no pool built -> (see -> [2026-08-23-acc4-decode-width-remeasurement.md](2026-08-23-acc4-decode-width-remeasurement.md)), -> so this row compares native-serial against ORT-single-thread specifically. -> Measuring ORT acc0 at t=4/8 is the prerequisite for sizing acc0 work. - -So the honest summary is that we have moved the *opt-in* path from 3.01x to -1.79x and left the *default* path at 1.84x, where it was. Those two numbers -being nearly equal is a coincidence of this shape, not a shared cause: the acc4 -gap is now dominated by the t=8 scaling anomaly, while the acc0 gap is a -different kernel entirely. **Closing acc0 is a separate, larger piece of work -and nothing here should be read as progress on it.** +> **Superseded 2026-08-23 — the table above is stale, not merely unlabelled.** +> Re-measured on `e189244ba` with a matched core budget and three independent +> launches per width, the acc0 gap is **1.120x at t=1** (native 35.36 ms, ORT +> 31.99), **1.148x at t=4**, **1.120x at t=8**. The rows above were correct when +> taken; they predate #1667 (broke the serial f32 reduction chain in the int4 +> decode GEMV, 5.75x t=1), #1679 (enabled the register-blocked kernel *at +> `accuracy_level = 0`*) and #1783 (folded the zero-point unpack), plus three +> merges that change what a width means (#1728, #1794, #1746). +> +> The control is the ORT arm: it reproduces to **+4.4%** (30.632 → 31.99 ms) on +> the same host and statistic, so the comparison is sound and the whole +> movement is on our side — native went 56.307 → 35.36 at t=1 and 14.091 → 4.57 +> at t=8. An earlier note here said the figure was `t=1`-only and that the +> production-width gap was unmeasured. Both were true and both missed that the +> `t=1` number was three kernel merges out of date. Full record: +> [2026-08-23-acc0-gap-vs-ort-by-width.md](2026-08-23-acc0-gap-vs-ort-by-width.md). + +So the honest summary *at the time of writing* was that we had moved the +*opt-in* path from 3.01x to 1.79x and left the *default* path at 1.84x, where +it was. Nothing in this document was progress on acc0, and that remains true of +this document. What is no longer true is the conclusion drawn from it: acc0 was +closed to ~1.12x by the three kernel merges listed in the note above, within +two days of this being written, and the figure here outlived the tree it +measured. **Do not rank work off this table.** Two implementation details cost more than they saved and are recorded so they are not re-tried: diff --git a/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md b/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md new file mode 100644 index 0000000000..d3108d41ac --- /dev/null +++ b/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md @@ -0,0 +1,205 @@ +# The acc0 int4 decode gap against ORT is ~1.1x, not 1.84x + +**Date:** 2026-08-23 · **Host:** AMD EPYC 9V74, 16 physical / 32 logical, single +NUMA, AVX2+FMA+F16C (no AVX-512) · **Main:** `e189244ba` · **Harness:** +`crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs` + +`benches/ort_matmulnbits_baseline.py` + +## Result + +llama3-8B projection chain, `block_size = 32`, `accuracy_level = 0` (the +production default), one session, `taskset` to physical cores, three +independent launches per width, four arms per cell interleaved. + +| width | native ms/token | ORT ms/token | **gap (ORT tok/s ÷ native tok/s)** | launches | across-launch spread | +|---:|---:|---:|---:|---:|---:| +| 1 | 35.36 | 31.99 | **1.120x** | 2 | 1.4% | +| 4 | 8.96 | 8.18 | **1.148x** | 4 | **17.1%** | +| 8 | 4.57 | 4.20 | **1.120x** | 3 | 5.0% | +| 16 | — | — | **not resolvable** | — | — | + +The previously published figure for this cell was **1.84x** at `t = 1` +(`docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md`), carried into the +ledger as the reason acc0 was the top remaining CPU MatMulNBits target. + +**That figure is not mislabelled, it is stale.** It is a real measurement of a +tree that no longer exists. + +## Why the old number moved, and the control that proves it + +The `1.84x` was `native 56.307 ms` against `ORT 30.632 ms`. Measuring both +sides again on current main: + +| arm | then | now | change | +|---|---:|---:|---| +| ORT, `t=1` | 30.632 ms | 31.99 ms | **+4.4%** — reproduces | +| native, `t=1` | 56.307 ms | 35.36 ms | **1.59x faster** | +| native, `t=8` | 14.091 ms | 4.57 ms | **3.08x faster** | + +**The ORT arm reproducing is the control.** It is the same binary, the same +graph, the same host and the same statistic, and it lands within 4.4% of a +number taken two days ago. So the harness, the shapes and the definition are +comparable, and the movement is entirely on our side. Had ORT moved too, this +would be a host or harness finding rather than a kernel one, and nothing below +could be claimed. + +The commit that published `56.307` is `e9754e7ef` (#1628). Every one of these +landed **after** it, verified with `git merge-base --is-ancestor`: + +| commit | PR | what it did | +|---|---|---| +| `8aed77a17` | #1667 | broke the serial f32 reduction chain in the int4 decode GEMV (**5.75x t=1**) | +| `99f105d52` | #1679 | enabled the register-blocked int4 kernel **at `accuracy_level = 0`** (1.17–1.69x) | +| `4e17f2251` | #1783 | folded the int4 zero-point unpack | +| `c3f0b0afa` | #1728 | sized the SPMD decode grain from its own pool, not ambient Rayon | +| `0652fdd2e` | #1794 | sized the persistent decode pool by physical cores | +| `6fdc04d75` | #1746 | gave the reserved dispatcher CPU a compute lane | + +The first three are direct acc0 kernel work and the last three change what a +width means. A gap figure that predates six such merges is not evidence about +today's tree, and **no amount of re-labelling would have fixed it** — the +2026-08-23 scope note added to it (#1843) said the number was `t=1`-only and +that the production-width gap was unmeasured. Both statements were true. Both +missed that the `t=1` number itself was three kernel merges out of date. + +**This is the more dangerous half of the "stale number" failure class.** A +number that is *wrong* gets challenged. A number that was *right when taken* +reads as evidence forever, because its provenance looks impeccable — and it +keeps a closed problem at the top of the priority list while the actual top +item goes unexamined. + +## What the gap actually is now + +Native scales slightly better than ORT across the range that is measurable: + +| width | native ms/token | native vs its own t=1 | ORT ms/token | ORT vs its own t=1 | +|---:|---:|---:|---:|---:| +| 1 | 35.36 | 1.00x | 31.99 | 1.00x | +| 4 | 8.96 | 3.95x | 8.18 | 3.91x | +| 8 | 4.57 | 7.73x | 4.20 | 7.62x | + +The `t=4` row is the weakest of the three: its gap spans 1.087–1.284 across +four launches (17.1%) and one of its A/A pairs came back at 0.868, so treat it +as "about the same as its neighbours" rather than as a distinct 1.15x. The +`t=1` and `t=8` rows agree to three digits (1.120x both) at 1.4% and 5.0%. + +So the gap is flat at ~1.12x from t=1 to t=8: it is **not** a scaling problem, +and there is no width at which acc0 collapses. The remaining ~11% is a kernel +efficiency difference, and it is small enough that it now sits below several +other open items rather than above them. + +## Method, and three things that had to be fixed before the number meant anything + +### 1. The two arms were not getting the same machine + +`ONNX_GENAI_CPU_DECODE_THREADS=w` does not merely size a pool. It confines the +whole process: + +``` +onnx-genai: CPU decode budget 4 confined the process to 4 CPUs [0, 2, 4, 6] +onnx-genai: CPU decode budget 4 bounded the global Rayon pool (prefill/MLAS parallelism capped at 4 workers) +``` + +`acc0_gap_matrix.py` pinned **both** arms to all 16 even CPUs at every thread +count. At `t=4` that gave ORT four threads free to roam sixteen cores — more +L3, more memory controllers — while native had four. On a workload that is +bandwidth-bound by construction, that is a machine-size comparison wearing a +thread-count label. It went unnoticed because the only published gap cell was +`t=1`, where ORT's `intra_op_num_threads=1` means one thread regardless, and +because at `t=16` the two pins coincide exactly. + +Both pins were measured so the size of the effect is data rather than +argument, and **the honest answer is that it is small**: ORT gains +**1–2%** from the wider pin on a quiet host (t=8: 1.120x matched vs 1.145x +wide; t=4: 1.148x vs 1.160x; t=1: 1.120x vs 1.136x). The concern was legitimate +and the fix is correct, but it does not move any conclusion. Recording that +here so the next person does not re-litigate it. + +### 2. The width had to be checked, not assumed + +Each native arm reports the width read back out of the pool, and a cell is +refused unless it matches the request: + +``` +decode_width requested=4 realized=4 path=spmd-pool as_requested +``` + +Timings cannot detect a vacuous sweep. A harness that silently runs one width +in every row looks perfectly stable, because it *is* — it is the same +configuration each time. That is precisely how `t=1 ≡ t=2` survived into three +documents (#1740, #1837). + +Width 1 legitimately reports `path=flat`: at `allowed.len() == 1`, +`build_from_env` declines to build a pool at all, so the t=1 column is +**serial-on-the-dispatcher**, not "a pool with one worker". The comparison at +that width is native-serial against ORT-single-thread. + +### 3. A pre-check cannot see a competitor that arrives mid-cell + +`wait_quiet` samples before the cell. During this matrix a sibling agent +started a `cargo test` on the same crate — a full 32-thread saturation — and +four cells that had passed the pre-check were measured against it, at intra-run +spreads of 20–63%. They are discarded. + +The guard is now a `LoadWatch` sampling the **instantaneous runnable count** +(field 4 of `/proc/loadavg`, not load average, which is a one-minute +exponential average that lags a job that just started and stays high long after +one ends) every second for the duration of every arm, refusing the cell if the +peak exceeds `width + 4`. The threshold has to scale with width: at width `w` +our own arm legitimately contributes ~`w` runnable threads, so a constant +"runnable > 4" would refuse every honest cell at `w >= 8`. + +**This is necessary, not sufficient.** Ten launches at width 16 split into a +fast and a slow mode 1.8x apart in wall time while burning *identical* +CPU-seconds (14.4 vs 14.1 CPU-s per wall-s): the affected threads were never +descheduled, they just retired fewer instructions per cycle. No load, +CPU-efficiency, or context-switch guard can see that. + +### 4. The wide-pin arm turned out to be a contention detector + +One `t=1` cell passed the load pre-check (runnable 3, no process above 150% +CPU) and still produced: + +``` +w=1 native 68.488 ms ort_matched 60.489 ms (both confined to CPU 0) +w=1 native_aa 61.077 ms ort_wide 32.409 ms (roams 16 CPUs) +``` + +Every arm confined to CPU 0 ran ~2x slow; the one arm free to roam was normal. +That is a single busy CPU, and no host-level guard can see it — the box really +was quiet in aggregate. The cell is discarded, and it was caught **only** +because an arm on a different pin was measured beside it. + +Worth generalising: at narrow widths the confinement lands on specific CPUs +(`w=1` → CPU 0), and CPU 0 is also where the harness, the shell and assorted +daemons live. **A single-CPU pin is the most fragile cell in any width sweep**, +and it is the one every speedup is quoted against. + +## What is still unresolved + +**Width 16 could not be measured, again.** Six cells at that width span native +5.719–12.486 ms/token with A/A ratios from 0.969 to 1.295, all taken while the +sibling `cargo test` was running. This is the same width whose launch +distribution spans 1.476–9.064 ms (514%) with no identified mechanism — see +[2026-08-23-acc4-decode-width-remeasurement.md](2026-08-23-acc4-decode-width-remeasurement.md). + +The contaminated cells do put ORT at 2.29–3.17 ms against native 5.72–6.00 at +that width, which would be a ~1.6x gap if it survived, and native was in its +slow mode for all of them. **That is a hypothesis, not a result**, and it is +the one cell that would matter most: `t=16` is closest to an unconfined +production process. It needs a dedicated quiet-host study with the launch +distribution treatment, not another row in a matrix. + +## Reproduce + +```bash +cargo build --release -p onnx-runtime-ep-cpu --bench int4_decode_loop_ab +python3 crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py \ + --binary target/release/deps/int4_decode_loop_ab- \ + --models llama --threads 1,4,8 --sessions 1 --acc 0 --block 32 \ + --tokens 192 --reps 2 --aa --ort-pin both +``` + +`--ort-pin matched` is the default and is the only setting that compares like +with like; `both` adds the wide-pin arm, which costs one extra ORT run per cell +and is worth it at narrow widths for the reason in §4. diff --git a/docs/performance/CPU_MATMUL_ASSIGNMENT.md b/docs/performance/CPU_MATMUL_ASSIGNMENT.md index 94496b5ee4..e265983963 100644 --- a/docs/performance/CPU_MATMUL_ASSIGNMENT.md +++ b/docs/performance/CPU_MATMUL_ASSIGNMENT.md @@ -1940,22 +1940,47 @@ t=1 and t=4 rows have headroom and are unaffected; **t=8 is the widest row worth arguing from.** **The default path did not move and is now the bigger target.** This kernel is -gated to `accuracy_level = 4`; production default is 0, where native is 56.307 -ms/token against ORT's 30.632 — **1.84x**, unchanged. **That figure is -`threads = 1` only.** It is the sole acc0 cell with an ORT baseline: the source -table has one further acc0 row (t=8, native 14.091) with no ORT number beside -it, so **the acc0 gap at production width is unmeasured**. That matters more -than a missing label usually would, because the acc4 table immediately above -shows the gap moving from 3.01x at t=1 to 1.47x at t=16 on the same shape — -width is one of the largest effects in this section. Do not quote 1.84x as -"the acc0 gap" without the `t=1` qualifier, and do not assume it survives to -t=8/16. Note also that at t=1 the native side runs `path=flat`, confined by -the decode budget to a single CPU with no pool built at all (§20), so this is -a serial-vs-ORT-single-thread comparison specifically. That it nearly equals -the new acc4 gap is a coincidence of this shape and this width, not a shared -cause. Nothing here is progress on acc0. - -Full record: +gated to `accuracy_level = 4`; production default is 0, where native was 56.307 +ms/token against ORT's 30.632 — **1.84x** at the time this was written. + +> **Re-measured 2026-08-23: the acc0 gap is ~1.12x, and 1.84x is stale.** On +> `e189244ba`, llama / block 32 / one session / matched core budget, three +> independent launches per width: **t=1 1.120x** (native 35.36 ms, ORT 31.99, +> 1.4% across launches), **t=4 1.148x** (17.1%, the weakest row), **t=8 1.120x** +> (5.0%). The gap is flat across the measurable range, so +> acc0 is neither a scaling problem nor the top target any more. +> +> The old figure was **not mislabelled — it was a correct measurement of a tree +> that no longer exists.** The control that establishes this is the ORT arm: +> it reproduces to **+4.4%** (30.632 → 31.99 ms), same binary, same graph, same +> host, same statistic, so the harness and the definition are comparable and +> the entire movement is on our side. Native went 56.307 → 35.36 at t=1 +> (1.59x) and 14.091 → 4.57 at t=8 (3.08x). Six merges landed after +> `e9754e7ef` published the number, three of them direct acc0 kernel work +> (#1667 broke the serial f32 reduction chain, 5.75x t=1; #1679 enabled the +> register-blocked kernel *at acc0*; #1783 folded the zero-point unpack) and +> three of which change what a width means (#1728, #1794, #1746). +> +> The 2026-08-23 scope note this replaces said the figure was `t=1`-only and +> that the production-width gap was unmeasured. Both were true; both missed +> that the `t=1` number was itself three kernel merges out of date. **A number +> that was right when taken is more durable than one that was wrong** — its +> provenance looks impeccable, so it is never challenged, and it keeps a closed +> problem at the top of the list while the real top item goes unexamined. The +> lesson is not "label the width", it is **re-measure before ranking work off a +> figure you did not take today**. +> +> Still true and unchanged: at t=1 the native side runs `path=flat`, confined by +> the decode budget to a single CPU with no pool built at all (§20), so that row +> compares native-serial against ORT-single-thread. **t=16 remains unresolved** +> — every cell at that width was taken against a sibling `cargo test` and +> discarded; the contaminated cells hint at ~1.6x and native was in the slow +> mode of its 514% bimodality for all of them, so it is a hypothesis, not a +> result, and it is the cell closest to an unconfined production process. +> Full record: +> [`docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md`](../benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md). + +Full record for the acc4 table above: [`docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md`](../benchmarks/2026-08-21-int4-acc4-execution-regime.md). ### 23. The register-blocked int4 decode kernel shipped default-off and stayed dormant for its whole life (**fixed**) From b09396e5cfd2f5419df426e9541ef6256dcb9943 Mon Sep 17 00:00:00 2001 From: Roy Date: Sun, 23 Aug 2026 16:13:11 +0000 Subject: [PATCH 2/3] docs(perf): measure the old-vs-new acc0 kernel A/B, and retract the t=8 3.08x MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review of the previous commit made the right objection: the ORT arm reproducing to +4.4% shows the *ORT* ruler did not move and says nothing about the native one, which changed repeatedly over the same window (#1722 is literally titled "make the acc0 native and ORT arms measure one quantity"). So the inference is replaced with a measurement. `e9754e7ef`'s bench is rebuilt in a second worktree and run beside current main's -- same host, same environment, `PROBE_REPS=1` on both so neither gets a rep loop the other lacks, arms interleaved and the order alternated: | width | kernel-only, measured | published pair implies | verdict | |------:|----------------------:|-----------------------:|---------| | 1 | 1.64x [1.61-1.88], 12 cells | 1.59x | movement is kernel | | 8 | 1.82x [1.78-1.89], 6 cells | 3.08x | 3.08x RETRACTED | Both old figures reproduce to within 0.4% -- but only **unpinned**. The old bench never called `EpFactory::initialize`, so it never ran `bound_process_to_decode_budget()`; that function, physical-core selection included, already existed at `e9754e7ef` and production always called it. The old t=8 row therefore measured eight decode workers scattered over 32 logical CPUs onto SMT siblings -- a topology no served session ever ran in. #1766 added the call. Pinned to eight physical cores the same old binary gives 8.430 ms against its unpinned 14.115, and forced onto `0-7`, 16.121. 1.67x of the claimed 3.08x was placement, not kernel work. This is the effect 2026-08-21-decode-worker-cpu-placement.md (#1680) already recorded, landing on a number I quoted two days later. Two corrections that look right and are not, recorded so they are not re-applied: the ~11% warmup/spawn handicap of §27 is in `tokens_s_total`, and both published figures are `ms_token` -- the old ORT harness docstring names "the native harness's `steady` column-2 median" as its comparand and the reproductions land on it. Deducting 11% yields a number neither tree produces. The asymmetry that *is* real, old ORT `min` over reps against old native single-shot, biases in ORT's favour. The gap conclusion is unchanged and is measured on today's tree, both arms, matched pins -- but the headline table is rebuilt to fix three defects: - it mixed statistics (native median latency over ORT throughput-equivalent) so its columns did not yield its own gap figure. Both sides are now `tokens_s_total`, with the mixed variant shown and labelled; - `t=4` was quoted as 1.148x when its A/A null spans 0.868-1.150, so the gap is inside its own noise floor there. Now "~1.15x, does not resolve"; - "three independent launches per width" was false (3/5/3 cells across two invocations) and the `t=1` 1.4% spread depended on an undisclosed post-hoc discard. Retained-cell figures are published beside the headline (1.112x [0.927-1.128], n=3) and the discard rule is stated prospectively. `acc0_gap_matrix.py` gains `--launches`, a per-width `--tokens 1:64,4:192` map and a `gap` column in ORT/native orientation beside `ratio`, so the Reproduce block names a command that produces the published table. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- .../benches/acc0_gap_matrix.py | 104 +++++- .../2026-08-21-int4-acc0-dormant-nblock.md | 10 +- .../2026-08-21-int4-acc4-execution-regime.md | 50 ++- .../2026-08-23-acc0-gap-vs-ort-by-width.md | 312 +++++++++++++++--- docs/performance/CPU_MATMUL_ASSIGNMENT.md | 82 +++-- 5 files changed, 456 insertions(+), 102 deletions(-) diff --git a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py index 7c4c6616a4..281c131c50 100755 --- a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py +++ b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py @@ -65,11 +65,36 @@ from the binary and a cell is refused unless it equals the request. Timings cannot detect this: a sweep that silently runs one width in every row looks perfectly stable, because it is -- it is the same configuration each time. + +Distribution, orientation and token budget +------------------------------------------ +* `--launches N` repeats the whole matrix `N` times. Every arm is already a + fresh process, so this distributes over **process launches**, which is the + larger source of variance and the one `--reps` cannot see: reps live inside + one process and inherit whatever state it started in. The per-width summary + at the end medians the gap over *paired* cells -- native and ORT run seconds + apart inside one launch, so pairing cancels drift that a ratio of two + independently-medianed columns would keep -- and prints the cell count and + the full range beside it, because a median of two cells and a median of six + are not the same claim. +* `--tokens` accepts either a single count or a per-width map, + `1:64,4:192,8:384`. A fixed count is not width-neutral: at `t=1` a token + costs ~16x what it costs at `t=16`, so a flat budget spends the most wall + time on the narrow widths and still gives them the fewest samples relative + to their variance. +* Two ratio columns are printed, and this is deliberate. `ratio` is + **native/ORT** (below 1.0 means we are behind); `gap` is its reciprocal, the + "Nx behind ORT" figure the documents quote. Printing both means neither + orientation has to be inferred from which side of 1.0 a number happens to + sit -- which is exactly the kind of inference that produced a published + reciprocal once already. """ import argparse +import collections import json import os import re +import statistics import subprocess import sys import threading @@ -298,8 +323,19 @@ def main(): ap.add_argument("--sessions", default="1,2,4") ap.add_argument("--block", type=int, default=32) ap.add_argument("--acc", type=int, default=0) - ap.add_argument("--tokens", type=int, default=24) + ap.add_argument("--tokens", default="24", + help="measured tokens per session. Either a single number, " + "or a per-width map `1:64,4:192,8:384` -- a fixed token " + "count makes a t=1 cell 16x longer in wall time than a " + "t=16 cell, so the narrow widths get the fewest samples " + "exactly where the variance is worst") ap.add_argument("--reps", type=int, default=3) + ap.add_argument("--launches", type=int, default=1, + help="independent repetitions of the whole matrix. Every " + "arm is a fresh process, so this distributes over " + "process launches rather than over reps inside one " + "process -- the launch-to-launch spread is the larger " + "of the two and is invisible to --reps") ap.add_argument("--out", default="acc0_gap.json") ap.add_argument("--aa", action="store_true", help="per-cell interleaved A/A") ap.add_argument("--slack", type=int, default=4, @@ -313,16 +349,28 @@ def main(): "or both so the asymmetry is quantified") args = ap.parse_args() + # `--tokens 192` or `--tokens 1:64,4:192,8:384`. + if ":" in args.tokens: + tokens_for = {} + for pair in args.tokens.split(","): + key, value = pair.split(":") + tokens_for[int(key)] = int(value) + else: + fixed = int(args.tokens) + tokens_for = collections.defaultdict(lambda: fixed) + rows = [] - hdr = (f"{'model':>6} {'t':>3} {'s':>2} {'nat_tps':>9} {'nat_sp%':>8} " - f"{'ort_tps':>9} {'ort_sp%':>8} {'ratio':>7} {'nat_GB/s':>9} " + hdr = (f"{'model':>6} {'L':>2} {'t':>3} {'s':>2} {'nat_tps':>9} {'nat_sp%':>8} " + f"{'ort_tps':>9} {'ort_sp%':>8} {'ratio':>7} {'gap':>7} {'nat_GB/s':>9} " f"{'ort_GB/s':>9} {'A/A':>6}") print(hdr) print("-" * len(hdr)) - for model in args.models.split(","): + for launch in range(args.launches): + for model in args.models.split(","): wb = weight_bytes(model, args.block) for t in [int(x) for x in args.threads.split(",")]: + tokens = tokens_for[t] for s in [int(x) for x in args.sessions.split(",")]: load, busy = wait_quiet() if busy: @@ -337,21 +385,21 @@ def main(): # there when `wait_quiet` returned. with LoadWatch() as watch: a1 = native(args.binary, model, args.block, args.acc, t, s, - args.tokens, args.reps) + tokens, args.reps) o = None o_wide = None if args.ort_pin in ("matched", "both"): - o = ort(model, args.block, args.acc, t, s, args.tokens, + o = ort(model, args.block, args.acc, t, s, tokens, args.reps, pin=native_pin(t)) if args.ort_pin in ("wide", "both"): o_wide = ort(model, args.block, args.acc, t, s, - args.tokens, args.reps, pin=PIN) + tokens, args.reps, pin=PIN) if o is None: o = o_wide aa = "" if args.aa: a2 = native(args.binary, model, args.block, args.acc, t, - s, args.tokens, args.reps) + s, tokens, args.reps) aa = f"{a2['tps'] / a1['tps']:.3f}" if watch.peak > t + args.slack: sys.stderr.write( @@ -361,8 +409,14 @@ def main(): busy = busy or [("-", 0.0, f"peak runnable {watch.peak}")] nat_bw = a1["tps"] * wb / 1e9 ort_bw = o["tps"] * wb / 1e9 - row = {"model": model, "threads": t, "sessions": s, + row = {"model": model, "launch": launch, "threads": t, + "sessions": s, "tokens": tokens, "native": a1, "ort": o, "ratio": a1["tps"] / o["tps"], + # `ratio` is native/ORT (below 1 means we are behind); + # `gap` is its reciprocal, the "Nx behind ORT" the docs + # quote. Both are printed so neither orientation has to + # be inferred from which side of 1.0 a number sits. + "gap": o["tps"] / a1["tps"], "ort_wide": o_wide, "ratio_wide": (a1["tps"] / o_wide["tps"]) if o_wide else None, "native_gbs": nat_bw, "ort_gbs": ort_bw, @@ -377,15 +431,43 @@ def main(): if o_wide is not None and row["ratio_wide"] is not None: wide = (f" wide_ort={o_wide['tps']:.1f} " f"ratio_wide={row['ratio_wide']:.3f}") - print(f"{model:>6} {t:>3} {s:>2} {a1['tps']:>9.1f} " + print(f"{model:>6} {launch:>2} {t:>3} {s:>2} {a1['tps']:>9.1f} " f"{a1['spread']:>8.1f} {o['tps']:>9.1f} {o['spread']:>8.1f} " - f"{row['ratio']:>7.3f} {nat_bw:>9.1f} {ort_bw:>9.1f} {aa:>6}" + f"{row['ratio']:>7.3f} {row['gap']:>7.3f} " + f"{nat_bw:>9.1f} {ort_bw:>9.1f} {aa:>6}" f"{wide}{flag}") sys.stdout.flush() with open(os.path.join(HERE, args.out), "w") as f: json.dump(rows, f, indent=1) print() + # Per-width summary over launches. The gap is medianed over *paired* cells + # -- native and ORT run seconds apart inside one launch, so pairing cancels + # drift that a ratio of two independently-medianed columns would keep. The + # cell count and the full range are printed beside it because a median of + # two cells and a median of six are not the same claim, and a range that + # straddles the A/A null means the cell did not resolve the gap at all. + if args.launches > 1 or args.aa: + print(f"{'model':>6} {'t':>3} {'s':>2} {'gap_med':>8} {'gap_min':>8} " + f"{'gap_max':>8} {'cells':>6} {'aa_min':>7} {'aa_max':>7}") + for model in args.models.split(","): + for t in [int(x) for x in args.threads.split(",")]: + for s in [int(x) for x in args.sessions.split(",")]: + kept = [r for r in rows + if r["model"] == model and r["threads"] == t + and r["sessions"] == s and r["trusted"]] + if not kept: + print(f"{model:>6} {t:>3} {s:>2} " + f"{'no trusted cells':>40}") + continue + gaps = sorted(r["gap"] for r in kept) + aas = sorted(r["aa"] for r in kept if r["aa"] is not None) + print(f"{model:>6} {t:>3} {s:>2} " + f"{statistics.median(gaps):>8.3f} {gaps[0]:>8.3f} " + f"{gaps[-1]:>8.3f} {len(gaps):>6} " + f"{(f'{aas[0]:.3f}' if aas else '-'):>7} " + f"{(f'{aas[-1]:.3f}' if aas else '-'):>7}") + print() print("weight bytes per token per session:", {m: f"{weight_bytes(m, args.block) / 1e6:.1f} MB" for m in args.models.split(",")}) diff --git a/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md index f562b54ecf..83a1cd1c42 100644 --- a/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md +++ b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md @@ -14,14 +14,16 @@ kernel — see the [placement record](2026-08-21-decode-worker-cpu-placement.md) ## 1. The route, by counter -The question was where the 1.84x acc0 gap sits. The first thing to establish is -which kernel the production default even runs. Route counters instrumented from +The question was where the acc0 gap sat, which at the time was measured at +1.84x (**stale — see the note below**). The first thing to establish is which +kernel the production default even runs. Route counters instrumented from operator entry through to the innermost arm, over a real decode step at `accuracy_level = 0`: > **Note (2026-08-23, revised):** the `1.84x` inherited here **no longer -> holds** — re-measured on `e189244ba` the acc0 gap is **1.12x at t=1, 1.15x at -> t=4, 1.12x at t=8**. Part of that closure is the work this very document +> holds** — re-measured on `e189244ba` the acc0 gap is **1.12x at t=1 and +> 1.12x at t=8**; the `t=4` cell reads ~1.15x but sits inside its own A/A null +> (0.868-1.150) and does not resolve. Part of that closure is the work this very document > motivated: enabling the register-blocked kernel at `accuracy_level = 0` > (#1679) was one of three acc0 merges that landed after the 1.84x was taken. > The route findings below are categorical (which kernel runs) and stand diff --git a/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md b/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md index d13157b263..9a73f81a10 100644 --- a/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md +++ b/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md @@ -484,29 +484,47 @@ This kernel is gated to `accuracy_level = 4`. Production default is | 8 | — | 14.091 | — | > **Superseded 2026-08-23 — the table above is stale, not merely unlabelled.** -> Re-measured on `e189244ba` with a matched core budget and three independent -> launches per width, the acc0 gap is **1.120x at t=1** (native 35.36 ms, ORT -> 31.99), **1.148x at t=4**, **1.120x at t=8**. The rows above were correct when -> taken; they predate #1667 (broke the serial f32 reduction chain in the int4 -> decode GEMV, 5.75x t=1), #1679 (enabled the register-blocked kernel *at +> Re-measured on `e189244ba` with a matched core budget, arms interleaved and +> the gap medianed over paired cells, the acc0 gap is **1.120x at t=1** +> (range 1.112–1.128) and **1.120x at t=8** (1.089–1.145). The `t=4` cell reads +> ~1.15x but does **not** resolve: its A/A null spans 0.868–1.150, so the gap +> is inside its own noise floor there. The rows above were correct when taken; +> they predate #1667 (broke the serial f32 reduction chain in the int4 decode +> GEMV, 5.75x t=1), #1679 (enabled the register-blocked kernel *at > `accuracy_level = 0`*) and #1783 (folded the zero-point unpack), plus three -> merges that change what a width means (#1728, #1794, #1746). +> merges that change what a width means (#1728, #1794, #1746) and two that +> changed the ruler (#1722, #1766). > -> The control is the ORT arm: it reproduces to **+4.4%** (30.632 → 31.99 ms) on -> the same host and statistic, so the comparison is sound and the whole -> movement is on our side — native went 56.307 → 35.36 at t=1 and 14.091 → 4.57 -> at t=8. An earlier note here said the figure was `t=1`-only and that the -> production-width gap was unmeasured. Both were true and both missed that the -> `t=1` number was three kernel merges out of date. Full record: +> **The `t=1` and `t=8` rows are stale for different reasons, and only one of +> them is the kernel.** An earlier version of this note argued from the ORT arm +> reproducing (+4.4%, 30.632 → 31.99 ms) that "the whole movement is on our +> side". That control is not sufficient — it shows the *ORT* ruler held still +> and says nothing about the native one. Rebuilding `e9754e7ef`'s bench and +> running it beside current main's, same host, `PROBE_REPS=1` on both, arms +> interleaved: **t=1 is 1.64x kernel** [1.61–1.88, 12 paired cells], matching +> the 1.59x the published pair implies. **t=8 is 1.82x kernel** [1.78–1.89, +> 6 cells], *not* the 3.08x that pair implies. +> +> The missing 1.67x at t=8 is worker placement. `14.091` reproduces today to +> 0.2% — but only **unpinned**. The old bench never called +> `EpFactory::initialize`, so `bound_process_to_decode_budget()` never ran and +> its eight workers scattered across 32 logical CPUs onto SMT siblings; that +> function already existed at `e9754e7ef` and production always called it, so +> the row measured a topology no served session used. #1766 added the call. +> The same old binary pinned to eight physical cores gives 8.430 ms, and forced +> onto `0-7` gives 16.121. This is the effect +> [2026-08-21-decode-worker-cpu-placement.md](2026-08-21-decode-worker-cpu-placement.md) +> already documented, landing on this table's own `t=8` row. Full record: > [2026-08-23-acc0-gap-vs-ort-by-width.md](2026-08-23-acc0-gap-vs-ort-by-width.md). So the honest summary *at the time of writing* was that we had moved the *opt-in* path from 3.01x to 1.79x and left the *default* path at 1.84x, where it was. Nothing in this document was progress on acc0, and that remains true of -this document. What is no longer true is the conclusion drawn from it: acc0 was -closed to ~1.12x by the three kernel merges listed in the note above, within -two days of this being written, and the figure here outlived the tree it -measured. **Do not rank work off this table.** +this document. What is no longer true is the conclusion drawn from it: acc0 is +~1.12x on current main, and the figure here outlived the tree it measured. +Roughly 1.6-1.8x of that closure is the three kernel merges listed in the note +above; at `t=8` the rest is the worker placement the old bench never applied. +**Do not rank work off this table.** Two implementation details cost more than they saved and are recorded so they are not re-tried: diff --git a/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md b/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md index d3108d41ac..281e1c8d49 100644 --- a/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md +++ b/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md @@ -8,15 +8,61 @@ NUMA, AVX2+FMA+F16C (no AVX-512) · **Main:** `e189244ba` · **Harness:** ## Result llama3-8B projection chain, `block_size = 32`, `accuracy_level = 0` (the -production default), one session, `taskset` to physical cores, three -independent launches per width, four arms per cell interleaved. - -| width | native ms/token | ORT ms/token | **gap (ORT tok/s ÷ native tok/s)** | launches | across-launch spread | -|---:|---:|---:|---:|---:|---:| -| 1 | 35.36 | 31.99 | **1.120x** | 2 | 1.4% | -| 4 | 8.96 | 8.18 | **1.148x** | 4 | **17.1%** | -| 8 | 4.57 | 4.20 | **1.120x** | 3 | 5.0% | -| 16 | — | — | **not resolvable** | — | — | +production default), one session, `taskset` to physical cores, four arms per +cell interleaved. + +The gap is `ORT tok/s ÷ native tok/s` computed **within** each launch, then +medianed across launches. It is a paired statistic on purpose: the two arms run +seconds apart on the same machine, so pairing cancels drift that a ratio of +two independently-medianed columns would keep. + +| width | native tok/s | ORT tok/s | **gap** | gap range | trusted cells | A/A range | +|---:|---:|---:|---:|---:|---:|---:| +| 1 | 27.9 | 31.2 | **1.120x** | 1.112–1.128 | 2 of 3 | 1.025–1.036 | +| 4 | 107.2 | 122.3 | **~1.15x** | 1.087–1.284 | 4 of 5 | 0.868–1.150 | +| 8 | 211.0 | 238.0 | **1.120x** | 1.089–1.145 | 3 of 3 | 0.997–1.028 | +| 16 | — | — | **not resolvable** | — | 0 of 3 | — | + +**Read the `t=4` row as "about the same as its neighbours", not as a distinct +1.15x.** Its A/A null — two identical native arms in the same launch — spans +0.868 to 1.150, so the whole gap sits inside its own noise floor at that width. +Quoting it to four digits would be false precision, and an earlier draft of this +document did exactly that. Only `t=1` and `t=8`, whose A/A nulls are within +3.6% and 2.8% of unity, resolve the gap at all. + +Counts are cells, not launches: `t=4` pools five cells across two script +invocations, of which four are trusted. `t=1` is discussed under +[the discard](#the-t1-discard-is-post-hoc-and-here-is-what-it-costs). + +**On statistics, because these columns are not interchangeable.** tok/s above is +`tokens_s_total` on both arms — wall-derived over every measured token — and the +gap divides one by the other, so it is like for like. The native binary +*additionally* reports a median per-token latency (`ms_token`), and the two +native views do not agree at width: + +| width | native `ms_token` (median latency) | native `1000 ÷ tok/s` | divergence | +|---:|---:|---:|---:| +| 1 | 35.61 ms | 35.84 ms | 0.6% | +| 4 | 8.95 ms | 9.33 ms | 4.2% | +| 8 | 4.56 ms | 4.74 ms | 3.9% | + +`tokens_s_total` is wall-derived and carries the tail; `ms_token` is a median +and discards it. The tail grows with width, which is why the divergence does. + +**A latency-over-latency gap cannot be formed today at all**: the current ORT +harness reports only `tokens_s_total`, so there is no ORT median to divide by. +Substituting native's `ms_token` into the numerator anyway — a mixed quantity, +given only to show the conclusion is not sensitive to the choice — yields +**1.113 / 1.098 / 1.085** at t=1 / 4 / 8, i.e. 0.6–2.4% below the throughput +gap and mildly *declining* with width rather than flat. An earlier draft of this +table did exactly that substitution without saying so, printed native `ms_token` +beside ORT `1000/tps`, and called the result a gap; its columns did not yield +its own gap figure. The headline above is throughput on both sides. + +Note the old published `1.84x` was itself a latency-style ratio (native +`ms_token` over ORT's min-over-reps per-`Run` median), so the then-versus-now +comparison spans a statistic change. It does not matter here: on either +definition today's figure is ~1.1x. The previously published figure for this cell was **1.84x** at `t = 1` (`docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md`), carried into the @@ -25,23 +71,120 @@ ledger as the reason acc0 was the top remaining CPU MatMulNBits target. **That figure is not mislabelled, it is stale.** It is a real measurement of a tree that no longer exists. -## Why the old number moved, and the control that proves it - -The `1.84x` was `native 56.307 ms` against `ORT 30.632 ms`. Measuring both -sides again on current main: - -| arm | then | now | change | -|---|---:|---:|---| -| ORT, `t=1` | 30.632 ms | 31.99 ms | **+4.4%** — reproduces | -| native, `t=1` | 56.307 ms | 35.36 ms | **1.59x faster** | -| native, `t=8` | 14.091 ms | 4.57 ms | **3.08x faster** | - -**The ORT arm reproducing is the control.** It is the same binary, the same -graph, the same host and the same statistic, and it lands within 4.4% of a -number taken two days ago. So the harness, the shapes and the definition are -comparable, and the movement is entirely on our side. Had ORT moved too, this -would be a host or harness finding rather than a kernel one, and nothing below -could be claimed. +## Why the old number moved: measured, not inferred + +The `1.84x` was `native 56.307 ms` against `ORT 30.632 ms`, and `t=8` carried a +companion `native 14.091 ms`. + +An earlier draft of this document argued from a control: ORT re-measures at +31.99 ms, within 4.4% of its published 30.632, therefore the harness is +comparable and *"the movement is entirely on our side."* **That argument does +not hold, and the reviewer who pushed on it was right to.** The ORT arm +reproducing shows the *ORT* ruler did not move. It says nothing about the +native ruler, which sits in a different binary and changed repeatedly over the +same window — including `81e611c03` (#1722), whose title is literally *"make +the acc0 native and ORT arms measure one quantity"*. + +So the inference was replaced with a measurement. `e9754e7ef`'s tree is checked +out in a second worktree, its `int4_decode_loop_ab` is built, and it is run +**beside** current main's on the same host, in the same environment, with +`PROBE_REPS=1` on both so neither gets a rep loop the other lacks, arms +interleaved within each launch and the launch order alternated. + +### First: both published figures reproduce to within 0.4%, which identifies them + +| published | rebuilt `e9754e7ef` today | reps | delta | +|---|---:|---:|---:| +| `56.307 ms` (t=1) | **56.519** (56.402 / 56.519 / 56.878) | 3 | +0.4% | +| `14.091 ms` (t=8) | **14.115** (14.105 / 14.115 / 14.196) | 3 | +0.2% | + +Both reproduce **unpinned**, in the `steady` **`ms_token`** column. That pins +down what the old numbers are, which matters because the two obvious +corrections both turn out to land elsewhere: + +- **The documented ~11% warmup handicap does not apply here.** §27 of the + ledger (#1712) records that the native arm's clock started before thread + spawn and three warmup steps — at `tokens = 24`, 27 steps of work charged + against 24 counted tokens. That bias is in **`tokens_s_total`**, the + wall-derived column. `ms_token` is a median over per-token samples collected + *after* the warmup steps, and it never carried it. The published figures are + `ms_token`: `ort_matmulnbits_baseline.py`'s own docstring at that commit names + its comparand as *"the native harness's `steady` column-2 median"*, and the + reproduction above lands on it to 0.4%. Deducting 11% from `56.307` would give + a number that no run of either tree produces. +- **There is a real statistic asymmetry, and it points the other way.** Old ORT + reported `min` over reps of a per-`Run` median; old native reported a single + rep's median with no rep loop at all. Best-of-N against single-shot flatters + ORT, so it made the old gap look *worse* than it was. Calling the two arms + "the same statistic", as the first draft did, was wrong; the direction of the + error is the one that does not help the argument. + +### Then: what actually moved, per width + +| width | factor | measured | what it is | +|---:|---|---:|---| +| 1 | old unpinned vs old pinned | pinned is **1.7% slower** | placement, negligible | +| 1 | **old → new, paired, 12 interleaved cells** | **1.64x** [1.61–1.88] | **kernel** | +| 1 | end-to-end, unpinned both | 1.60x (56.519 → 35.361) | — | +| 8 | old unpinned → old pinned to 8 physical cores | **1.67x** (14.115 → 8.430) | **benchmark defect** | +| 8 | **old → new, paired, 6 interleaved cells** | **1.82x** [1.78–1.89] | **kernel** | +| 8 | end-to-end, unpinned both | 3.06x (14.115 → 4.619) | — | + +**At `t=1` the apparent movement is kernel.** Measured old-vs-new is 1.60–1.64x +against the 1.59x the published pair implies. The ruler question was worth +asking and the answer is that it does not bite at this width. + +**At `t=8` it is not, and I am retracting the `3.08x` this document previously +claimed.** 1.67x of it is a defect in the *old benchmark*, and the mechanism is +specific: the old bench never called `EpFactory::initialize`, so it never ran +`bound_process_to_decode_budget()` and its process was never confined. That +function — with its physical-core `select_budget_cpus` — **already existed at +`e9754e7ef`**; only the bench was missing the call, which `11cb8e5f3` (#1766) +added. So the old bench measured a thread topology **no served session ever ran +in**: eight decode workers scattered across 32 logical CPUs, landing on SMT +siblings, against a full-width prefill/MLAS Rayon pool. + +The SMT mechanism is directly measurable. Forcing the same binaries onto +`0-7` — four physical cores plus their siblings — against `0,2,...,14`, eight +distinct physical cores: + +| binary | 8 physical cores | 4 cores + SMT siblings | unpinned | penalty | +|---|---:|---:|---:|---:| +| `e9754e7ef` | 8.430 ms | 16.121 ms | **14.115 ms** | 1.91x | +| current main | 4.664 ms | 7.988 ms | **4.619 ms** | 1.71x | + +The old binary's unpinned run sits between its two pinned extremes, which is +what "the scheduler chose for us" looks like. Today's binary is **unaffected by +the absence of a pin** (4.619 unpinned vs 4.664 pinned, 0.99x) because it +confines itself. That is the correct behaviour and it is what production always +did — but it is a *measurement* correction, not a speedup delivered to anyone, +and reporting it inside a kernel-improvement figure was the error. + +**This effect was already documented, on this host, by me, and I still walked +into it.** [2026-08-21-decode-worker-cpu-placement.md](2026-08-21-decode-worker-cpu-placement.md) +(#1680, ledger §24) concluded that an apparent "t=8 wash" was worker-to-CPU +placement rather than the kernel, and the dormant-nblock record opens with +"unpinned multi-thread numbers on this host measure worker placement, not the +kernel". The `14.091` figure is an unpinned multi-thread number on this host. +Knowing the failure mode, having written it down, and citing it in a +neighbouring document was not enough to stop me quoting a 3.08x off exactly +that kind of number two days later. + +The net effect on this document's conclusion is nil: the **gap** rows at the top +are measured today, on both arms, under matched pins. What changes is the +credit — **1.8x of kernel improvement at `t=8`, not 3.1x.** + +### Why the paired ratio survives a busy host + +The absolute levels above move with load; the ratios do not. One launch caught a +competitor mid-cell and came back at `old 13.263 ms / new 7.007 ms` — both arms +about 1.6x slow — and its ratio was **1.893**, inside the range of the five +clean launches. Interleaving the arms seconds apart inside one launch is what +buys that: whatever slows one arm has usually not left by the time the other +runs. It is the same reason the gap column at the top is a paired median rather +than a ratio of independently-medianed columns. + +## Six merges landed between the old number and this one The commit that published `56.307` is `e9754e7ef` (#1628). Every one of these landed **after** it, verified with `git merge-base --is-ancestor`: @@ -56,11 +199,19 @@ landed **after** it, verified with `git merge-base --is-ancestor`: | `6fdc04d75` | #1746 | gave the reserved dispatcher CPU a compute lane | The first three are direct acc0 kernel work and the last three change what a -width means. A gap figure that predates six such merges is not evidence about -today's tree, and **no amount of re-labelling would have fixed it** — the -2026-08-23 scope note added to it (#1843) said the number was `t=1`-only and -that the production-width gap was unmeasured. Both statements were true. Both -missed that the `t=1` number itself was three kernel merges out of date. +width means. Two more changed the **ruler** rather than the tree, and belong +beside them: + +| commit | PR | what it did to the measurement | +|---|---|---| +| `81e611c03` | #1722 | made the native and ORT acc0 arms measure one quantity (barrier; removed the warmup/spawn charge from `tokens_s_total`) | +| `11cb8e5f3` | #1766 | made the benches call `EpFactory::initialize`, so a `t=N` row finally runs in the topology a served session runs in — **this is the 1.67x above** | + +A gap figure that predates eight such merges is not evidence about today's tree, +and **no amount of re-labelling would have fixed it** — the 2026-08-23 scope +note added to it (#1843) said the number was `t=1`-only and that the +production-width gap was unmeasured. Both statements were true. Both missed that +the `t=1` number itself was three kernel merges out of date. **This is the more dangerous half of the "stale number" failure class.** A number that is *wrong* gets challenged. A number that was *right when taken* @@ -70,23 +221,55 @@ item goes unexamined. ## What the gap actually is now -Native scales slightly better than ORT across the range that is measurable: +Native scales slightly better than ORT across the range that is measurable. +Both columns are `1000 ÷ tok/s`, so the scaling ratios are like for like: -| width | native ms/token | native vs its own t=1 | ORT ms/token | ORT vs its own t=1 | +| width | native | native vs its own t=1 | ORT | ORT vs its own t=1 | |---:|---:|---:|---:|---:| -| 1 | 35.36 | 1.00x | 31.99 | 1.00x | -| 4 | 8.96 | 3.95x | 8.18 | 3.91x | -| 8 | 4.57 | 7.73x | 4.20 | 7.62x | - -The `t=4` row is the weakest of the three: its gap spans 1.087–1.284 across -four launches (17.1%) and one of its A/A pairs came back at 0.868, so treat it -as "about the same as its neighbours" rather than as a distinct 1.15x. The -`t=1` and `t=8` rows agree to three digits (1.120x both) at 1.4% and 5.0%. - -So the gap is flat at ~1.12x from t=1 to t=8: it is **not** a scaling problem, -and there is no width at which acc0 collapses. The remaining ~11% is a kernel -efficiency difference, and it is small enough that it now sits below several -other open items rather than above them. +| 1 | 35.84 ms | 1.00x | 32.00 ms | 1.00x | +| 4 | 9.33 ms | 3.84x | 8.17 ms | 3.91x | +| 8 | 4.74 ms | 7.56x | 4.20 ms | 7.62x | + +**`t=1` is a different code path from the other two rows**, and any scaling +figure quoted against it should say so. At `allowed.len() == 1`, +`build_from_env` declines to build a pool at all, so the native `t=1` column is +serial-on-the-dispatcher (`path=flat`, read back from the runtime and checked +per cell) rather than a one-worker pool. It is "vs serial", not "vs one +worker". Sebastian measured the same thing from the other side on the acc4 path +(#1740): at `total_workers <= 1`, `dispatch_output_rows` short-circuits and the +spawned worker receives no dispatch at all, 0% busy over a six-second window. + +So the gap is flat at ~1.12x at the two widths that resolve it: it is **not** a +scaling problem, and there is no width at which acc0 collapses. The remaining +~11–12% is a kernel efficiency difference, and it is small enough that it now +sits below several other open items rather than above them. + +### The `t=1` discard is post-hoc, and here is what it costs + +Three `t=1` cells were taken and **two are in the headline**. The discarded one +is the CPU-0 contention cell described under method §4 below: both arms confined +to CPU 0 ran ~2x slow (native 68.488, ORT 60.489) while the roaming arm beside +them was normal (32.409). + +The discard rule was **not pre-registered** — it was written after seeing that +cell — so the retained-cell figures have to be published beside it: + +| `t=1` | gap | range | spread | cells | +|---|---:|---:|---:|---:| +| headline (CPU-0 cell discarded) | 1.120x | 1.112–1.128 | **1.4%** | 2 | +| all trusted cells retained | 1.112x | 0.927–1.128 | **18.1%** | 3 | + +The median barely moves, so the conclusion survives either way. What does not +survive is the *precision*: "1.4% across launches" and "`t=1` and `t=8` agree to +three digits" are both artefacts of the discard, and the honest version of the +`t=1` row is `1.11–1.12x` with one cell in three thrown out for a documented, +independently-corroborated reason. + +**The rule, stated in advance for next time:** a cell is refused if any arm's +wide-pin counterpart beats its matched-pin counterpart by more than the +1–2% measured pin asymmetry, because that is a per-CPU contention signature the +host-level guard cannot see. Applying that rule prospectively would have +rejected this cell without anyone having to look at its gap. ## Method, and three things that had to be fixed before the number meant anything @@ -192,14 +375,47 @@ distribution treatment, not another row in a matrix. ## Reproduce +The gap matrix: + ```bash cargo build --release -p onnx-runtime-ep-cpu --bench int4_decode_loop_ab python3 crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py \ --binary target/release/deps/int4_decode_loop_ab- \ --models llama --threads 1,4,8 --sessions 1 --acc 0 --block 32 \ - --tokens 192 --reps 2 --aa --ort-pin both + --tokens 1:64,4:192,8:384 --reps 2 --launches 3 --aa --ort-pin both ``` `--ort-pin matched` is the default and is the only setting that compares like with like; `both` adds the wide-pin arm, which costs one extra ORT run per cell -and is worth it at narrow widths for the reason in §4. +and is worth it at narrow widths for the reason in §4. `--launches 3` is what +produces the distribution — reps inside one process cannot see the +launch-to-launch spread, which is the larger of the two. The per-width token +map exists because a flat count spends the most wall time on the narrowest +width and still gives it the fewest samples relative to its variance. + +The published table came from three invocations of this script rather than one +(`t=4` pools two of them), which is why its cell counts are 3 / 5 / 3 rather +than a uniform `--launches 3`. The `gap` column it prints is the one quoted +here; `ratio` beside it is the reciprocal, native/ORT. + +The old-versus-new kernel A/B: + +```bash +git worktree add ../old-tree e9754e7ef +git -C ../old-tree submodule update --init crates/onnx-runtime-cpuinfo/vendor/cpuinfo +cargo build --release --manifest-path ../old-tree/Cargo.toml \ + -p onnx-runtime-ep-cpu --bench int4_decode_loop_ab +``` + +Then run both binaries alternately with `PROBE_REPS=1` on each (the old tree +has no rep loop, so matching it is the only way both report one statistic), +`PROBE_BLOCK=32 PROBE_ACCURACY=0 PROBE_SESSIONS=1 PROBE_LAYERS=1`, +`ONNX_GENAI_CPU_DECODE_THREADS=w`, comparing the `steady` row's **`ms_token`** +column. Do not compare `tokens_s_total` across the two trees: the old binary +starts its wall clock before thread spawn and three warmup steps, which is the +~11% bias #1722 removed, and it lives entirely in that column. + +To reproduce the placement finding, run the *old* binary at `t=8` three ways — +`taskset -c 0,2,4,6,8,10,12,14` (eight physical cores), `taskset -c 0-7` (four +cores plus SMT siblings), and unpinned. The unpinned run is the one that +returns 14.09 ms. diff --git a/docs/performance/CPU_MATMUL_ASSIGNMENT.md b/docs/performance/CPU_MATMUL_ASSIGNMENT.md index e265983963..f67e8ebc90 100644 --- a/docs/performance/CPU_MATMUL_ASSIGNMENT.md +++ b/docs/performance/CPU_MATMUL_ASSIGNMENT.md @@ -1939,27 +1939,62 @@ an ORT number. Neither is resolvable by the measurement that produced it. The t=1 and t=4 rows have headroom and are unaffected; **t=8 is the widest row worth arguing from.** -**The default path did not move and is now the bigger target.** This kernel is -gated to `accuracy_level = 4`; production default is 0, where native was 56.307 -ms/token against ORT's 30.632 — **1.84x** at the time this was written. +**The default path was the bigger target when this was written, and is not any +more.** This kernel is gated to `accuracy_level = 4`; production default is 0, +where native measured 56.307 ms/token against ORT's 30.632 — **1.84x** at the +time. That figure is now stale; the re-measurement replaces it. > **Re-measured 2026-08-23: the acc0 gap is ~1.12x, and 1.84x is stale.** On -> `e189244ba`, llama / block 32 / one session / matched core budget, three -> independent launches per width: **t=1 1.120x** (native 35.36 ms, ORT 31.99, -> 1.4% across launches), **t=4 1.148x** (17.1%, the weakest row), **t=8 1.120x** -> (5.0%). The gap is flat across the measurable range, so -> acc0 is neither a scaling problem nor the top target any more. +> `e189244ba`, llama / block 32 / one session / matched core budget, arms +> interleaved and the gap medianed over paired cells: **t=1 1.120x** (native +> 27.9 tok/s, ORT 31.2; range 1.112–1.128 over 2 retained cells of 3), +> **t=8 1.120x** (range 1.089–1.145, 3 cells), **t=4 ~1.15x but unresolved** — +> its A/A null spans 0.868–1.150, so the gap sits inside its own noise floor at +> that width. Flat across the measurable range, so acc0 is neither a scaling +> problem nor the top target any more. > > The old figure was **not mislabelled — it was a correct measurement of a tree -> that no longer exists.** The control that establishes this is the ORT arm: -> it reproduces to **+4.4%** (30.632 → 31.99 ms), same binary, same graph, same -> host, same statistic, so the harness and the definition are comparable and -> the entire movement is on our side. Native went 56.307 → 35.36 at t=1 -> (1.59x) and 14.091 → 4.57 at t=8 (3.08x). Six merges landed after -> `e9754e7ef` published the number, three of them direct acc0 kernel work -> (#1667 broke the serial f32 reduction chain, 5.75x t=1; #1679 enabled the -> register-blocked kernel *at acc0*; #1783 folded the zero-point unpack) and -> three of which change what a width means (#1728, #1794, #1746). +> that no longer exists.** An earlier draft argued this from the ORT arm alone +> (it reproduces to +4.4%, 30.632 → 31.99 ms). **That control is not +> sufficient**: it shows the *ORT* ruler did not move and says nothing about the +> native ruler, which changed repeatedly over the same window. So the inference +> was replaced with a direct A/B — `e9754e7ef`'s bench rebuilt and run beside +> current main's, same host, same environment, `PROBE_REPS=1` on both, arms +> interleaved: +> +> | width | kernel-only, measured | published pair implies | verdict | +> |---:|---:|---:|---| +> | 1 | **1.64x** [1.61–1.88], 12 paired cells | 1.59x | apparent movement **is** kernel | +> | 8 | **1.82x** [1.78–1.89], 6 paired cells | 3.08x | **3.08x retracted** | +> +> **The `t=8` overclaim is worth knowing about.** Both old figures reproduce +> today to within 0.4% — but only **unpinned**. The old bench never called +> `EpFactory::initialize`, so it never ran `bound_process_to_decode_budget()` +> and its process was never confined; that function, physical-core selection +> included, **already existed at `e9754e7ef`**, and only the bench was missing +> the call (added by #1766 `11cb8e5f3`). So the old `t=8` row measured eight +> decode workers scattered over 32 logical CPUs onto SMT siblings — a topology +> **no served session ever ran in**. Pinning the same old binary to eight +> physical cores gives 8.430 ms against its unpinned 14.115, and forcing it onto +> `0-7` gives 16.121: 1.67x of the claimed 3.08x was that, not kernel work. +> Today's binary is pin-insensitive (4.619 unpinned vs 4.664 pinned) because it +> confines itself. +> +> Two corrections that look right and are not, recorded so they are not +> re-applied here: the ~11% warmup/spawn handicap of §27 is in +> **`tokens_s_total`**, and both published figures are **`ms_token`** (the old +> ORT harness's docstring names "the native harness's `steady` column-2 median" +> as its comparand, and the reproductions land on it to 0.4%), so deducting 11% +> yields a number neither tree produces. The asymmetry that *is* real — +> old ORT took `min` over reps, old native was single-shot — biases in **ORT's** +> favour, making the old gap look worse rather than better. +> +> Eight merges landed after `e9754e7ef` published the number: three direct acc0 +> kernel work (#1667 broke the serial f32 reduction chain, 5.75x t=1; #1679 +> enabled the register-blocked kernel *at acc0*; #1783 folded the zero-point +> unpack), three changing what a width means (#1728, #1794, #1746), and two +> changing the ruler (#1722 made the two arms measure one quantity; #1766 put +> the benches in the production decode topology). > > The 2026-08-23 scope note this replaces said the figure was `t=1`-only and > that the production-width gap was unmeasured. Both were true; both missed @@ -1972,12 +2007,13 @@ ms/token against ORT's 30.632 — **1.84x** at the time this was written. > > Still true and unchanged: at t=1 the native side runs `path=flat`, confined by > the decode budget to a single CPU with no pool built at all (§20), so that row -> compares native-serial against ORT-single-thread. **t=16 remains unresolved** -> — every cell at that width was taken against a sibling `cargo test` and -> discarded; the contaminated cells hint at ~1.6x and native was in the slow -> mode of its 514% bimodality for all of them, so it is a hypothesis, not a -> result, and it is the cell closest to an unconfined production process. -> Full record: +> compares native-serial against ORT-single-thread and every scaling figure +> quoted against it is "vs serial", not "vs a one-worker pool". **t=16 remains +> unresolved** — every cell at that width was taken against a sibling +> `cargo test` and discarded; the contaminated cells hint at ~1.6x and native +> was in the slow mode of its 514% bimodality for all of them, so it is a +> hypothesis, not a result, and it is the cell closest to an unconfined +> production process. Full record: > [`docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md`](../benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md). Full record for the acc4 table above: From 9b9c355ee5b59168738f8aa137f1b9a0f6282bed Mon Sep 17 00:00:00 2001 From: Justin Chu Date: Sun, 23 Aug 2026 17:24:42 +0000 Subject: [PATCH 3/3] docs(perf): restore the t=16 acc0 cells and condition the re-ranking on them Second adversarial review found that this record wrote off width 16 as "0 of 3 cells, all taken against a sibling cargo test". That is false, and false in the direction that flattered the conclusion: two of the three t=16 cells passed the load guard cleanly (runnable 6, no competitor recorded) and read 1.831 and 1.456, median 1.643x. The width still does not resolve, but for the honest reason -- its A/A null spans 0.969-1.295 (against 3.6% at t=1 and 2.8% at t=8) and both arms show 20-55% intra-run spread -- not because the data was contaminated. The cell the guard did refuse reads 1.585, between the two retained, so the discard is not load-bearing either way. Because t=16 is the width closest to an unconfined production process, the re-ranking claim ("acc0 is no longer the top CPU MatMulNBits target") is now stated as conditional on it, in all four places it appears, with the quiet-host t=16 study named as the thing that settles it. Also from the review: - Headline cell counts corrected: t=4 is 4 trusted of 6 taken (was "4 of 5"), and the published table came from two script invocations, not three. - The "trusted cells" column split into harness-trust vs editorial retention. All three t=1 cells passed the guard; one was dropped afterwards by me, and conflating the two hid that. - The t=1 placement probe's samples are now published rather than summarised, including the 114.94 ms outlier on one pinned rep -- that excursion is the old binary's bimodality and is why the t=1 A/B range reaches 1.88x. - t=16's ORT arm spread (55.4%) disclosed alongside native's; the denominator is no better behaved than the numerator at that width. - The four checksum falsifier constants in int4_decode_loop_ab's module doc were stale. Corrected, with a note that they drift with reduction reassociation and that the *pattern* -- block 16 moves under ONNX_GENAI_CPU_MM_INT4_GEBP=0, block 32 does not -- is the route evidence. - acc0_gap_matrix.py: a --tokens map missing a --threads width died with a bare KeyError after the first cell had already waited out the load guard. It now refuses the run at parse time with the missing widths named. - Reproduce block updated to the two invocations actually run, so the commands as written pass the new validation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- .../benches/acc0_gap_matrix.py | 15 +- .../benches/int4_decode_loop_ab.rs | 9 +- .../2026-08-21-int4-acc0-dormant-nblock.md | 4 +- .../2026-08-21-int4-acc4-execution-regime.md | 6 + .../2026-08-23-acc0-gap-vs-ort-by-width.md | 163 ++++++++++++++---- docs/performance/CPU_MATMUL_ASSIGNMENT.md | 21 ++- 6 files changed, 168 insertions(+), 50 deletions(-) diff --git a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py index 281c131c50..ab3149222e 100755 --- a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py +++ b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py @@ -350,11 +350,22 @@ def main(): args = ap.parse_args() # `--tokens 192` or `--tokens 1:64,4:192,8:384`. + threads = [int(x) for x in args.threads.split(",")] if ":" in args.tokens: tokens_for = {} for pair in args.tokens.split(","): key, value = pair.split(":") tokens_for[int(key)] = int(value) + # A per-width map that is missing a width would otherwise die with a + # bare KeyError partway through a matrix that has already spent + # minutes in `wait_quiet`. Fail before any measurement starts. + missing = sorted(set(threads) - set(tokens_for)) + if missing: + ap.error( + f"--tokens map has no entry for width(s) {missing}; it must " + f"cover every --threads value. Got {sorted(tokens_for)}, need " + f"{sorted(threads)}. Use a scalar (--tokens 192) for one " + f"token count at every width.") else: fixed = int(args.tokens) tokens_for = collections.defaultdict(lambda: fixed) @@ -369,7 +380,7 @@ def main(): for launch in range(args.launches): for model in args.models.split(","): wb = weight_bytes(model, args.block) - for t in [int(x) for x in args.threads.split(",")]: + for t in threads: tokens = tokens_for[t] for s in [int(x) for x in args.sessions.split(",")]: load, busy = wait_quiet() @@ -451,7 +462,7 @@ def main(): print(f"{'model':>6} {'t':>3} {'s':>2} {'gap_med':>8} {'gap_min':>8} " f"{'gap_max':>8} {'cells':>6} {'aa_min':>7} {'aa_max':>7}") for model in args.models.split(","): - for t in [int(x) for x in args.threads.split(",")]: + for t in threads: for s in [int(x) for x in args.sessions.split(",")]: kept = [r for r in rows if r["model"] == model and r["threads"] == t diff --git a/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs b/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs index 88f2e81d82..0696c5a9a7 100644 --- a/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs +++ b/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs @@ -36,8 +36,13 @@ //! not a multiple of 32, so at `m = 1` the GEBP gate is already satisfied //! and block 16 runs the *fused prefill* kernel `quant_prefill_gebp`. //! Falsifier: `ONNX_GENAI_CPU_MM_INT4_GEBP=0` changes the block-16 checksum -//! (844.702358 -> 844.714874) and its time, and leaves block 32 untouched -//! (979.199155 either way). Any block-16 row taken here is a GEBP row. +//! (844.536810 -> 844.551163) and its time, and leaves block 32 untouched +//! (978.949310 either way). Any block-16 row taken here is a GEBP row. +//! These constants are numerics-sensitive and drift whenever a kernel +//! reassociates its reduction (#1667, #1783 both moved them in the fourth +//! decimal place); it is the *pattern* -- block 16 moves, block 32 does not +//! -- that is the route evidence, so re-derive them rather than reading a +//! mismatch as a route change. //! Set `ONNX_GENAI_CPU_MM_INT4_GEBP=0` to reach the decode kernel at 16. //! - `PROBE_ACCURACY` -- `accuracy_level` (default 0). **4 is the only value //! that reaches the packed-nibble kernel**, so without this axis that route diff --git a/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md index 83a1cd1c42..622c533147 100644 --- a/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md +++ b/docs/benchmarks/2026-08-21-int4-acc0-dormant-nblock.md @@ -23,7 +23,9 @@ operator entry through to the innermost arm, over a real decode step at > **Note (2026-08-23, revised):** the `1.84x` inherited here **no longer > holds** — re-measured on `e189244ba` the acc0 gap is **1.12x at t=1 and > 1.12x at t=8**; the `t=4` cell reads ~1.15x but sits inside its own A/A null -> (0.868-1.150) and does not resolve. Part of that closure is the work this very document +> (0.868-1.150) and does not resolve, and **`t=16` reads ~1.64x on two +> guard-passing cells against a 0.969-1.295 null, so it does not resolve +> either** and remains the open row. Part of that closure is the work this very document > motivated: enabling the register-blocked kernel at `accuracy_level = 0` > (#1679) was one of three acc0 merges that landed after the 1.84x was taken. > The route findings below are categorical (which kernel runs) and stand diff --git a/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md b/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md index 9a73f81a10..c49d1a6e8b 100644 --- a/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md +++ b/docs/benchmarks/2026-08-21-int4-acc4-execution-regime.md @@ -516,6 +516,12 @@ This kernel is gated to `accuracy_level = 4`. Production default is > [2026-08-21-decode-worker-cpu-placement.md](2026-08-21-decode-worker-cpu-placement.md) > already documented, landing on this table's own `t=8` row. Full record: > [2026-08-23-acc0-gap-vs-ort-by-width.md](2026-08-23-acc0-gap-vs-ort-by-width.md). +> +> **`t=16` is the open row.** Two guard-passing cells there read 1.831 and 1.456 +> (median **1.643x**), but that width's A/A null spans 0.969–1.295 and both arms +> show 20–55% intra-run spread, so it does not resolve. It is the width closest +> to an unconfined production process, and a confirmed ~1.64x there would +> reverse the re-ranking below. So the honest summary *at the time of writing* was that we had moved the *opt-in* path from 3.01x to 1.79x and left the *default* path at 1.84x, where diff --git a/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md b/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md index 281e1c8d49..4dc34a816c 100644 --- a/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md +++ b/docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md @@ -16,23 +16,37 @@ medianed across launches. It is a paired statistic on purpose: the two arms run seconds apart on the same machine, so pairing cancels drift that a ratio of two independently-medianed columns would keep. -| width | native tok/s | ORT tok/s | **gap** | gap range | trusted cells | A/A range | +| width | native tok/s | ORT tok/s | **gap** | gap range | cells (trusted/taken) | A/A range | |---:|---:|---:|---:|---:|---:|---:| -| 1 | 27.9 | 31.2 | **1.120x** | 1.112–1.128 | 2 of 3 | 1.025–1.036 | -| 4 | 107.2 | 122.3 | **~1.15x** | 1.087–1.284 | 4 of 5 | 0.868–1.150 | -| 8 | 211.0 | 238.0 | **1.120x** | 1.089–1.145 | 3 of 3 | 0.997–1.028 | -| 16 | — | — | **not resolvable** | — | 0 of 3 | — | - -**Read the `t=4` row as "about the same as its neighbours", not as a distinct -1.15x.** Its A/A null — two identical native arms in the same launch — spans -0.868 to 1.150, so the whole gap sits inside its own noise floor at that width. -Quoting it to four digits would be false precision, and an earlier draft of this -document did exactly that. Only `t=1` and `t=8`, whose A/A nulls are within -3.6% and 2.8% of unity, resolve the gap at all. - -Counts are cells, not launches: `t=4` pools five cells across two script -invocations, of which four are trusted. `t=1` is discussed under -[the discard](#the-t1-discard-is-post-hoc-and-here-is-what-it-costs). +| 1 | 27.9 | 31.2 | **1.120x** | 1.112–1.128 | 3/3, **2 retained** | 1.025–1.036 | +| 4 | 107.2 | 122.3 | ~1.15x | 1.087–1.284 | 4/6 | **0.868–1.150** | +| 8 | 211.0 | 238.0 | **1.120x** | 1.089–1.145 | 3/3 | 0.997–1.028 | +| 16 | — | — | ~1.64x, **does not resolve** | 1.456–1.831 | 2/3 | **0.969–1.295** | + +Two columns of that table are doing different jobs and must not be read the same +way. *Trusted/taken* is the harness's own verdict — it refuses a cell whose peak +runnable count exceeded `width + slack`. *Retained* is editorial: all three `t=1` +cells passed the harness, and one was dropped afterwards by me, which is +[disclosed and costed below](#the-t1-discard-is-post-hoc-and-here-is-what-it-costs). + +**Only `t=1` and `t=8` resolve the gap.** The others fail on their own A/A null — +two identical native arms in the same launch, which is the noise floor any real +effect has to clear: + +- **`t=4`** reads ~1.15x against an A/A of 0.868–1.150. The gap is inside its own + noise floor. Read it as "about the same as its neighbours", not as a distinct + 1.15x; an earlier draft of this document quoted `1.148x` to four digits. +- **`t=16`** reads 1.64x against an A/A of 0.969–1.295 — ±30%, an order of + magnitude worse than at `t=8`. **This is the row that matters most and it is + the row we cannot measure**, because `t=16` is the closest cell to an + unconfined production process. See [below](#what-is-still-unresolved); an + earlier draft of this document wrongly wrote it off as contaminated. + +`t=1` and `t=8` clear their nulls by 3.6% and 2.8% respectively. + +**So "the gap is ~1.12x" is a statement about `t=1` and `t=8` only.** It is not +established at `t=16`, where the two cells that did pass the guard both point +higher. **On statistics, because these columns are not interchangeable.** tok/s above is `tokens_s_total` on both arms — wall-derived over every measured token — and the @@ -123,7 +137,7 @@ corrections both turn out to land elsewhere: | width | factor | measured | what it is | |---:|---|---:|---| -| 1 | old unpinned vs old pinned | pinned is **1.7% slower** | placement, negligible | +| 1 | old unpinned vs old pinned | pinned is **1.9% slower** | placement, negligible | | 1 | **old → new, paired, 12 interleaved cells** | **1.64x** [1.61–1.88] | **kernel** | | 1 | end-to-end, unpinned both | 1.60x (56.519 → 35.361) | — | | 8 | old unpinned → old pinned to 8 physical cores | **1.67x** (14.115 → 8.430) | **benchmark defect** | @@ -134,6 +148,25 @@ corrections both turn out to land elsewhere: against the 1.59x the published pair implies. The ruler question was worth asking and the answer is that it does not bite at this width. +The `t=1` placement row above is a four-arm probe of its own, and its samples +are given here rather than summarised, because one of them is an outlier that +matters: + +| arm | median | samples | +|---|---:|---| +| old, pinned `taskset -c 0` | 57.703 | 57.306 / 57.703 / **114.94** | +| old, unpinned | 56.651 | 56.454 / 56.651 / 56.784 | +| new, pinned | 35.007 | 34.965 / 35.007 / 35.556 | +| new, unpinned | 35.383 | 35.057 / 35.383 / 35.414 | + +Placement is worth 1.9% at this width in either binary — nothing like the 1.67x +it is worth at `t=8` — which is the expected result: a single-threaded process +has no SMT sibling to collide with. **The `114.94` is the old binary's +bimodality**, a 2.0x excursion on one pinned rep of three with the other two +within 0.7% of each other, and it is the reason the `t=1` A/B range extends to +1.88x. It is reported, not dropped; the median is unaffected by it, which is +why the median is what the table quotes. + **At `t=8` it is not, and I am retracting the `3.08x` this document previously claimed.** 1.67x of it is a defect in the *old benchmark*, and the mechanism is specific: the old bench never called `EpFactory::initialize`, so it never ran @@ -239,10 +272,18 @@ worker". Sebastian measured the same thing from the other side on the acc4 path (#1740): at `total_workers <= 1`, `dispatch_output_rows` short-circuits and the spawned worker receives no dispatch at all, 0% busy over a six-second window. -So the gap is flat at ~1.12x at the two widths that resolve it: it is **not** a -scaling problem, and there is no width at which acc0 collapses. The remaining -~11–12% is a kernel efficiency difference, and it is small enough that it now -sits below several other open items rather than above them. +So the gap is flat at ~1.12x at the two widths that resolve it, and across those +widths acc0 is **not** a scaling problem. The remaining ~11–12% is a kernel +efficiency difference small enough to sit below several other open items rather +than above them. + +**That re-ranking is conditional on `t=16`, and `t=16` is not in the table +above.** Its two guard-passing cells read ~1.64x — the width closest to an +unconfined production process is also the only one pointing at a large gap, and +its A/A null is too wide to call it either way. If a quiet-host study confirms +1.64x there, acc0 goes back to the top of the list. "acc0 is no longer the top +target" is therefore a claim about `t=1` and `t=8`, held provisionally, with the +`t=16` study as the thing that settles it. ### The `t=1` discard is post-hoc, and here is what it costs @@ -360,18 +401,48 @@ and it is the one every speedup is quoted against. ## What is still unresolved -**Width 16 could not be measured, again.** Six cells at that width span native -5.719–12.486 ms/token with A/A ratios from 0.969 to 1.295, all taken while the -sibling `cargo test` was running. This is the same width whose launch -distribution spans 1.476–9.064 ms (514%) with no identified mechanism — see -[2026-08-23-acc4-decode-width-remeasurement.md](2026-08-23-acc4-decode-width-remeasurement.md). +**Width 16, and it is the row that matters.** `t=16` is the closest cell in this +matrix to an unconfined production process, and it is the one width where the gap +does not resolve. -The contaminated cells do put ORT at 2.29–3.17 ms against native 5.72–6.00 at -that width, which would be a ~1.6x gap if it survived, and native was in its -slow mode for all of them. **That is a hypothesis, not a result**, and it is -the one cell that would matter most: `t=16` is closest to an unconfined -production process. It needs a dedicated quiet-host study with the launch -distribution treatment, not another row in a matrix. +An earlier draft of this document dismissed it as contaminated — "every cell taken +against a sibling `cargo test`". **That was wrong, and it was wrong in the +direction that flattered us.** Two of the three `t=16` cells passed the load guard +cleanly (runnable 6, no competitor recorded), and those two cells read: + +| `t=16` | gap | A/A | native spread | ORT spread | +|---|---:|---:|---:|---:| +| launch 1 (runnable 6) | 1.831 | 1.295 | 17.8% | **55.4%** | +| launch 2 (runnable 6) | 1.456 | 0.969 | 27.7% | **19.6%** | +| **median** | **1.643** | — | — | — | + +So the honest statement is **~1.64x at `t=16`, from two cells the harness +accepted** — not "no data". The reason it still does not resolve is the A/A null, +not contamination: two *identical* native arms in the same launch differ by up to +29.5%, against 3.6% at `t=1` and 2.8% at `t=8`. A 1.64x gap measured on an +instrument with a ±30% null is not a 1.64x result. The intra-run spreads say the +same thing from inside each cell, and **both arms are unstable at this width** — +the ORT arm's own rep-to-rep spread reaches 55.4%, so the denominator is no +better behaved than the numerator. + +This is the same width whose launch distribution spans 1.476–9.064 ms/token +(514%) with no identified mechanism — see +[2026-08-23-acc4-decode-width-remeasurement.md](2026-08-23-acc4-decode-width-remeasurement.md) +— and the two cells above sit in different modes of it, which is the obvious +candidate for why the null is so wide. + +The third `t=16` cell — the one the guard *did* refuse, at runnable 9 against a +sibling `resch-dispatch` debug build — reads 1.585, i.e. **between** the two +retained cells. So the refusal is not load-bearing for the ~1.6x figure either +way; it is the null, not the discard, that stops this width resolving. + +**What this costs the headline.** "The gap is ~1.12x and flat" is established at +`t=1` and `t=8` and **is not established at `t=16`**, where the available +evidence points to roughly 1.6x. The re-ranking conclusion — that acc0 is no +longer the top CPU MatMulNBits target — rests on the two widths that resolve, and +**a confirmed 1.64x at `t=16` would overturn it.** That makes a dedicated +quiet-host study of `t=16`, with launch distributions and a pre-registered A/A +acceptance threshold, the first thing to do next rather than a footnote. ## Reproduce @@ -379,12 +450,24 @@ The gap matrix: ```bash cargo build --release -p onnx-runtime-ep-cpu --bench int4_decode_loop_ab -python3 crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py \ - --binary target/release/deps/int4_decode_loop_ab- \ +BIN=target/release/deps/int4_decode_loop_ab- + +# invocation 1 — widths 1, 4, 8 +python3 crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py --binary "$BIN" \ --models llama --threads 1,4,8 --sessions 1 --acc 0 --block 32 \ --tokens 1:64,4:192,8:384 --reps 2 --launches 3 --aa --ort-pin both + +# invocation 2 — widths 4, 16 +python3 crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py --binary "$BIN" \ + --models llama --threads 4,16 --sessions 1 --acc 0 --block 32 \ + --tokens 4:192,16:384 --reps 2 --launches 3 --aa --ort-pin both ``` +The `--tokens` map must name every width in `--threads`; the script refuses the +run up front if it does not, rather than dying with a `KeyError` after the first +cell has already waited out the load guard. Pass a scalar (`--tokens 192`) for +one token count at every width. + `--ort-pin matched` is the default and is the only setting that compares like with like; `both` adds the wide-pin arm, which costs one extra ORT run per cell and is worth it at narrow widths for the reason in §4. `--launches 3` is what @@ -393,10 +476,10 @@ launch-to-launch spread, which is the larger of the two. The per-width token map exists because a flat count spends the most wall time on the narrowest width and still gives it the fewest samples relative to its variance. -The published table came from three invocations of this script rather than one -(`t=4` pools two of them), which is why its cell counts are 3 / 5 / 3 rather -than a uniform `--launches 3`. The `gap` column it prints is the one quoted -here; `ratio` beside it is the reciprocal, native/ORT. +The published table came from **two** invocations of this script rather than one +(widths 1/4/8, then widths 4/16), which is why its cell counts are 3 / 6 / 3 / 3 +rather than a uniform `--launches 3`. The `gap` column it prints is the one +quoted here; `ratio` beside it is the reciprocal, native/ORT. The old-versus-new kernel A/B: @@ -419,3 +502,7 @@ To reproduce the placement finding, run the *old* binary at `t=8` three ways — `taskset -c 0,2,4,6,8,10,12,14` (eight physical cores), `taskset -c 0-7` (four cores plus SMT siblings), and unpinned. The unpinned run is the one that returns 14.09 ms. + +The `t=1` placement probe (the four-arm table above) alternates old/new x +pinned/unpinned at `ONNX_GENAI_CPU_DECODE_THREADS=1`, three reps, with the pinned +arms under `taskset -c 0`. diff --git a/docs/performance/CPU_MATMUL_ASSIGNMENT.md b/docs/performance/CPU_MATMUL_ASSIGNMENT.md index f67e8ebc90..0d5511ee00 100644 --- a/docs/performance/CPU_MATMUL_ASSIGNMENT.md +++ b/docs/performance/CPU_MATMUL_ASSIGNMENT.md @@ -1951,7 +1951,8 @@ time. That figure is now stale; the re-measurement replaces it. > **t=8 1.120x** (range 1.089–1.145, 3 cells), **t=4 ~1.15x but unresolved** — > its A/A null spans 0.868–1.150, so the gap sits inside its own noise floor at > that width. Flat across the measurable range, so acc0 is neither a scaling -> problem nor the top target any more. +> problem nor the top target any more — **conditional on t=16**, which does not +> resolve and points at ~1.6x (see below). > > The old figure was **not mislabelled — it was a correct measurement of a tree > that no longer exists.** An earlier draft argued this from the ORT arm alone @@ -2008,12 +2009,18 @@ time. That figure is now stale; the re-measurement replaces it. > Still true and unchanged: at t=1 the native side runs `path=flat`, confined by > the decode budget to a single CPU with no pool built at all (§20), so that row > compares native-serial against ORT-single-thread and every scaling figure -> quoted against it is "vs serial", not "vs a one-worker pool". **t=16 remains -> unresolved** — every cell at that width was taken against a sibling -> `cargo test` and discarded; the contaminated cells hint at ~1.6x and native -> was in the slow mode of its 514% bimodality for all of them, so it is a -> hypothesis, not a result, and it is the cell closest to an unconfined -> production process. Full record: +> quoted against it is "vs serial", not "vs a one-worker pool". **t=16 does not +> resolve, and it is the row that matters** — two of its three cells passed the +> load guard cleanly and read 1.831x and 1.456x (median **1.643x**), but the +> width's A/A null spans 0.969–1.295 (±30%, against 3.6% at t=1 and 2.8% at t=8) +> and native sits in different modes of its 514% launch bimodality across the two +> cells. So it is ~1.6x on an instrument too loose to call it, not "no data" — +> an earlier draft wrote the width off as wholly contaminated, which was wrong in +> the direction that flattered us. **The re-ranking below is conditional on +> that**: a confirmed 1.64x at t=16, the width closest to an unconfined +> production process, would put acc0 back at the top. A dedicated quiet-host +> study of t=16 with launch distributions and a pre-registered A/A threshold is +> the next action. Full record: > [`docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md`](../benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md). Full record for the acc4 table above: