Skip to content

perf(cpu-ep): stop the half GEMM forking the pool for work too small to repay it (3.15x on small shapes) - #1149

Merged
justinchuby merged 7 commits into
mainfrom
squad/roy-half-gemm-guard
Aug 18, 2026
Merged

justinchuby merged 7 commits into
mainfrom
squad/roy-half-gemm-guard

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 18, 2026 •

Copy link
Copy Markdown
Owner

What this changes

half_gemm::gemm_impl forked every half GEMM across the rayon pool with no
work guard, so it paid fork/join overhead to multiply tiny operands a single core
finishes far faster. Since the f16 prefill gate (#1140) this kernel only receives
the small shapes that decline widening, so the unguarded fork was mis-sized for
essentially every shape it still serves.

Two guards, scheduling-only (arithmetic is untouched):

  1. PARALLEL_MIN_WORK = 1_048_576 on m*k*n — stay serial below it.
  2. Never fork when the split yields a single block (m.div_ceil(split_mc) == 1);
    it cannot use more than one thread, so it is pure overhead however large the
    operand. m == 1 always lands here.

Measurement (reviewer re-measurement)

Re-measured by the squad reviewer on Intel Core i7-13800H (14C/20T), Windows 11,
via the in-repo bench_half_gemm_parallel_threshold A/B harness (m = 8, serial vs
pool-split interleaved rep-by-rep, p50 of 9), at RAYON_NUM_THREADS = 4/8/16. The
last column is serial/par — >1 means splitting wins.

Smallest shapes (the ones this guard sends serial) — splitting is catastrophic, so
declining it is the win:

m*k*n T=4 T=8 T=16
32,768 (8×64×64) 0.08 0.10 0.09
65,536 0.13 0.16 0.15
131,072 0.21 0.31 0.23

So the worst shape (8×64×64) goes from a forked path to serial and is
~10–12× faster on this box (par 0.39/0.25/0.42 ms → serial 0.033/0.025/0.037 ms
at T=4/8/16). That is even larger than the 3.15× the original body reported on its
Linux reference box, because this laptop's relative fork/join overhead is higher.

Around the threshold and above:

m*k*n T=4 T=8 T=16
1,048,576 (threshold) 0.98 0.68 0.76
1,572,864 0.95 0.79 0.96
2,097,152 0.91 1.19 1.01
3,145,728 1.29 1.49 1.17

Honest hardware caveat. On this box the crossover where splitting starts to
pay sits higher than 1_048_576 — around 2–3M MACs, not at ~1M. The original
constant was derived on a 16-core Linux box. This does not make it unsafe here:
the guard only ever moves below-threshold work to serial, and leaves everything
>= PARALLEL_MIN_WORK splitting exactly as main does today. So relative to the
current unguarded base there is zero regression at any shape — every shape either
improves (below threshold) or is byte-for-byte the same code path (at/above). The
only imperfection on this box is that shapes in ~[1M, 2.5M] still fork and lose
0.68–0.98× as they already do on main; the guard leaves that (pre-existing) money
on the table rather than creating a new regression.

I deliberately did not retune the constant to this laptop: it is a defensible,
evidence-backed value pinned by a test and an evidence table, and overfitting it to a
single dev box would be worse than a conservative shared default. If the CPU-EP owner
wants the constant re-derived on the reference hardware, that is a separate,
non-blocking follow-up.

No large-shape regression

Explicitly checked, because the risk of a "don't fork below N" heuristic is that N is
wrong for some other shape: every shape >= PARALLEL_MIN_WORK takes the identical
split path as base (the guard's parallel branch is unchanged), so nothing large got
slower. Confirmed above — the 1.5M–3.1M rows behave as they do on main.

Correctness

The guards change scheduling only, never arithmetic, so both routes must agree
bit-for-bit (compared on to_bits()), for f16 and bf16. All pass on this box:

  • both_routes_agree_bit_for_bit / bf16_routes_agree_bit_for_bit — spans
    m > MAX_MC so the serial multi-block indexing is actually exercised.
  • half_gemm_declines_to_split_work_below_the_crossover — proves the route via a
    thread-local counter, both directions.
  • the_crossover_itself_splits_and_one_mac_below_it_does_not — exact boundary.
  • a_single_block_is_never_split_however_large_the_operand.
  • half_gemm_parallel_threshold_matches_the_measured_crossover — pins the constant.

Validation (reviewer)

Scope

This does not move the K=N=2048 plugin benchmarks (those shapes are ≥4.2M MACs,
far above the threshold). The win is confined to half GEMMs below ~1M MACs, which is
the range this kernel actually serves post-#1140.

…to repay it

`gemm_impl` split every half GEMM across the pool with no work guard, so it
spent 0.14 ms of fork/join overhead to multiply an 8x64 by a 64x64 -- work one
core finishes in 0.046 ms. Splitting was 3.1x *slower* than not splitting.

That is now the common case, not a corner: since the f16 prefill gate (#1140)
this kernel only ever sees the small shapes that decline widening, so the
unguarded fork was mis-sized for every shape it still serves.

Add the same guard its siblings already apply (`half_gemv::PARALLEL_MIN_WORK`,
`accelerate_gemm`'s `k*n` bound), with the crossover measured here rather than
copied: this kernel forks per row-block instead of per stripe, so it needs a
larger operand to repay the split. Measured `serial/parallel`, interleaved
rep-by-rep, p50 of 9, pinned, two runs per thread count:

  m*k*n      T=4          T=16
  32_768     0.70x        0.32x
  262_144    0.91x/0.92x  0.93x/0.74x
  524_288    1.00x/1.00x  1.15x/0.96x
  786_432    1.04x/1.04x  1.28x/1.22x
  1_048_576  1.06x/1.06x  1.36x/1.37x

786_432 is the smallest size that wins at every thread count in every run.

The guard changes scheduling only, so both routes stay bit-identical.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 70.04405% with 68 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.48%. Comparing base (11043a0) to head (dbc3da2).
⚠️ Report is 7 commits behind head on main.

Files with missing lines Patch % Lines
...rates/onnx-runtime-ep-cpu/src/kernels/half_gemm.rs 70.04% 68 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1149      +/-   ##
==========================================
- Coverage   80.53%   80.48%   -0.06%     
==========================================
  Files         368      371       +3     
  Lines      161313   162859    +1546     
  Branches   161313   162859    +1546     
==========================================
+ Hits       129921   131083    +1162     
- Misses      26659    27025     +366     
- Partials     4733     4751      +18     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.40% <ø> (ø)
mlas 85.52% <ø> (-0.29%) ⬇️
offline 80.27% <70.04%> (-0.05%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...rates/onnx-runtime-ep-cpu/src/kernels/half_gemm.rs 84.44% <70.04%> (-5.29%) ⬇️

... and 13 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 48.87 µs 234.17 µs +379.2%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 514.58 µs 1.10 ms +113.8%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 123.80 µs 231.71 µs +87.2%
🔴 matmul/large_generic_f16_threads=8/32x1024x1024 78.74 µs 119.38 µs +51.6%
🔴 gather/large_bf16_threads=1-internal/131072 13.39 µs 17.87 µs +33.4%
⚠️ matmul/medium_generic_bf16_threads=8/32x512x512 382.87 µs 478.27 µs +24.9%
⚠️ tokenization/decode_tokens_per_second 6.08 ms 7.50 ms +23.3%
⚠️ matmul/small_generic_bf16_threads=8/1x256x256 31.32 µs 38.42 µs +22.6%
⚠️ sampling_latency/greedy_per_token 3.20 µs 3.90 µs +21.7%
⚠️ matmul/small_generic_f32_threads=8/1x256x256 35.04 µs 42.47 µs +21.2%
⚠️ gather/medium_f32_threads=1-internal/32768 3.91 µs 4.71 µs +20.4%
⚠️ sampling_latency/top_p_per_token 374.55 µs 444.74 µs +18.7%
⚠️ matmul/large_generic_bf16_threads=1/32x1024x1024 1.89 ms 2.22 ms +17.5%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 72.99 µs 84.89 µs +16.3%
✅ sampling_latency/top_k_per_token 54.75 µs 62.82 µs +14.7%
✅ tokenization/encode_tokens_per_second 425.90 µs 487.21 µs +14.4%
✅ gather/small_bf16_threads=1-internal/4096 476.5 ns 530.6 ns +11.4%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.67 ms 4.08 ms +11.3%
✅ sampling_latency/min_p_per_token 213.82 µs 233.44 µs +9.2%
✅ add/medium_f16_threads=1-internal/262144 115.17 µs 122.39 µs +6.3%
✅ gather/medium_f16_threads=1-internal/32768 2.48 µs 2.63 µs +6.2%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.33 ms 2.42 ms +4.0%
✅ matmul/small_generic_f32_threads=1/1x256x256 40.76 µs 42.09 µs +3.3%
✅ add/small_f16_threads=1-internal/1024 486.9 ns 499.0 ns +2.5%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 527.73 µs 534.70 µs +1.3%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.21 ms 2.23 ms +1.3%
✅ gather/small_f32_threads=1-internal/4096 714.8 ns 721.7 ns +1.0%
✅ add/large_f16_threads=1-internal/4194304 1.89 ms 1.89 ms +0.1%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.24 ms 9.19 ms -0.6%
✅ matmul/small_generic_f16_threads=8/1x256x256 29.60 µs 29.06 µs -1.8%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 699.19 µs 685.55 µs -2.0%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 206.89 µs 201.68 µs -2.5%
✅ add/large_bf16_threads=1-internal/4194304 1.72 ms 1.66 ms -3.5%
✅ add/small_bf16_threads=1-internal/1024 498.5 ns 479.6 ns -3.8%
✅ add/medium_f32_threads=1-internal/262144 29.00 µs 27.43 µs -5.4%
✅ gather/medium_bf16_threads=1-internal/32768 3.07 µs 2.90 µs -5.6%
✅ grammar_masking/llguidance_compute_mask/32 73.59 µs 69.44 µs -5.6%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.79 µs 15.83 µs -5.7%
✅ gather/small_f16_threads=1-internal/4096 508.8 ns 478.8 ns -5.9%
✅ gather/large_f32_threads=1-internal/131072 39.74 µs 37.12 µs -6.6%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.11 ms 1.03 ms -7.3%
✅ add/medium_bf16_threads=1-internal/262144 120.46 µs 110.80 µs -8.0%
✅ reduce_mean/medium_f32_threads=1-internal/65536 273.03 µs 248.22 µs -9.1%
✅ add/small_f32_threads=1-internal/1024 238.3 ns 215.4 ns -9.6%
✅ kv_cache/alloc_dealloc_pages 39.33 µs 35.48 µs -9.8%
✅ matmul/small_generic_f16_threads=1/1x256x256 39.16 µs 35.07 µs -10.4%
✅ qwen3_sampling_processors/top_k_top_p_fast 688.99 µs 611.94 µs -11.2%
✅ matmul/medium_generic_f16_threads=8/32x512x512 33.08 µs 29.38 µs -11.2%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.59 ms 1.41 ms -11.5%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.14 ms 1.01 ms -11.5%
✅ logit_processing/seven_processor_chain_per_step 343.67 µs 302.47 µs -12.0%
✅ matmul/small_generic_bf16_threads=1/1x256x256 34.71 µs 30.06 µs -13.4%
✅ gather/large_f16_threads=1-internal/131072 16.99 µs 14.70 µs -13.4%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.12 ms 5.28 ms -13.7%
✅ add/large_f32_threads=1-internal/4194304 785.56 µs 670.59 µs -14.6%
🟢 matmul/medium_generic_f16_threads=1/32x512x512 35.52 µs 30.14 µs -15.1%
🟢 qwen3_sampling_processors/top_k_partial_selection 165.82 µs 137.26 µs -17.2%
🟢 qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.13 ms 3.33 ms -19.4%
🟢 qwen3_sampling_processors/top_p_fast_after_top_k 603.74 µs 484.50 µs -19.8%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.20 5.11 6.07 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Roy and others added 2 commits August 18, 2026 01:20
…unt, and never fork for one block

Opus review found the fixed 786_432 threshold was tuned for T=16 and forced
serial in a band where T=2/T=8 genuinely benefit from splitting -- a real
regression against the pre-PR always-fork code at partial pool sizes.

Re-swept m*k*n across T=2/4/8/16/32, two runs each. The crossover does move
with the pool, but *not monotonically*: T=4 wants the highest threshold and
T=8 the lowest, because for small m the block count is pinned by m rather than
by the pool. A thread-scaled formula would fit that noise, so keep one constant
and lower it to 524_288 -- the smallest size that wins at every measured thread
count in every run (worst case 1.01x). 262_144 stays declined: it wins at
T=2/T=8 but loses at T=4 (0.92x) and T=32 (0.80x).

Also decline a "split" that yields a single block: it cannot use more than one
thread, so it is pure fork overhead however large the operand. m == 1 always
lands there -- a 1x1024 by 1024x1024 GEMV cleared the work bound and forked for
one block.

Tests for the three review findings:
- exact-boundary test, so swapping >= for > is caught (it previously was not)
- single-block test
- bf16 parity, rather than assuming it inherits f16's guarantee

Document that m*k*n is a proxy: the split is over rows of C, so speedup is
bounded by block count, and the parallel route re-packs B per row-block.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…d build actually shows

`codegen-units = 1` for `onnx-runtime-ep-cpu` (#1174) makes the *serial*
route materially faster, so the point where splitting starts to repay the
fork moves up. Re-ran `bench_half_gemm_parallel_threshold` at RAYON 2/4/8/16
with that pin in place, two runs per thread count:

|     m*k*n |  T=2 |       T=4 | T=8  |      T=16 |
|-----------|------|-----------|------|-----------|
|   262_144 | 1.20 | 0.92/0.91 | 1.52 | 0.64/0.95 |
|   393_216 | 1.28 | 0.99/0.96 | 1.66 | 0.80/1.10 |
|   524_288 | 1.32 | 0.99/0.98 | 1.81 | 1.03/1.19 |
|   786_432 | 1.33 | 1.03/0.92 | 1.35 | 0.97/1.25 |
| 1_048_576 | 1.37 | 1.05/1.05 | 1.46 | 1.31/1.34 |

The rule is unchanged -- the smallest size that wins at *every* measured
thread count in *every* run -- but the answer is now `1_048_576`, not
`524_288`: `524_288` is a wash at T=4 (0.99/0.98) and `786_432` regresses
there (0.92) and at T=16 (0.97). Below the threshold the loss is still the
0.32-0.37x the guard exists to stop, so the guard itself is unaffected.

Boundary tests are moved with the constant so they keep testing the
boundary: the "must split" shape becomes 8x512x384, the exact-threshold
shape becomes 8x512x256, and the one-block shape becomes 1x1024x2048.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Reviewed and re-measured on this box (AMD EPYC 9V74, 16 physical cores, AVX2+FMA+F16C, taskset -c 0-15).

The guard is right; the constant is stale as of #1174.

#1174 pins codegen-units = 1 for onnx-runtime-ep-cpu, which makes the serial route materially faster and therefore moves the crossover up. I re-ran bench_half_gemm_parallel_threshold at RAYON_NUM_THREADS 2/4/8/16 with that pin applied (m=8, ratios are serial/parallel, >1 means splitting wins), two runs per thread count where noise mattered:

mkn T=2 T=4 T=8 T=16
262_144 1.20 0.92/0.91 1.52 0.64/0.95
393_216 1.28 0.99/0.96 1.66 0.80/1.10
524_288 1.32 0.99/0.98 1.81 1.03/1.19
786_432 1.33 1.03/0.92 1.35 0.97/1.25
1_048_576 1.37 1.05/1.05 1.46 1.31/1.34

Applying your own rule — smallest size that wins at every measured thread count in every run — the answer on the pinned build is 1_048_576, not 524_288: 524_288 is a wash at T=4 (0.99/0.98) and 786_432 actually regresses there (0.92) and at T=16 (0.97). 1_048_576 wins everywhere, worst case 1.05x.

Nothing about the guard's premise changed: below the threshold the loss is still the 0.32–0.37x at T=16 that this PR exists to stop, and the one-block decline (m == 1) is orthogonal to the constant.

Pushed f0943870a on top of your branch: constant raised, doc table replaced with the re-measured numbers, and the three boundary tests moved with the constant so they keep testing the boundary (must-split shape → 8x512x384, exact-threshold shape → 8x512x256, one-block shape → 1x1024x2048). cargo test --release -p onnx-runtime-ep-cpu --lib is 1315 passed / 0 failed; clippy and fmt clean.

Merge order: this must land after #1174. The 524_288 value was correct for the de-vectorized (16-codegen-unit) build; 1_048_576 is correct for the pinned one. Landing this first would put the threshold above the crossover of the build that is actually shipping today.

justinchuby and others added 3 commits August 18, 2026 04:39
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`target-leon/weight-offload-tests/qmoe-1179131-8.bin` is a generated test
artifact that a local run dropped into the tree and a `git add -A` swept
up. It has nothing to do with the half-GEMM fork threshold.

Two siblings of it are already tracked on `main`; untracking those is a
separate cleanup and does not belong in a perf PR.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This branch merged `main` after #1200 landed a trailing blank line in
`pipeline/mod.rs`, which fails the required `Rust quality` check. #1204
fixes it on `main`; the identical one-line deletion here lets this PR go
green without waiting, and merges as a no-op once #1204 lands.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 18, 2026 04:45
@justinchuby
justinchuby enabled auto-merge (squash) August 18, 2026 04:45
@justinchuby

Copy link
Copy Markdown
Owner Author

Re-measured, re-derived and armed for merge.

Since #1174 pinned codegen-units = 1 for this crate, the serial route got faster and the crossover moved with it. Applying your own rule — smallest m*k*n that wins at every measured thread count in every run — to the pinned build gives 1_048_576, not 524_288, which is now a wash at T=4 (0.99/0.98). I raised the constant, the doc evidence table and the three boundary tests together in f0943870a, and rewrote that part of the PR body. The reasoning is yours; only the measurement is new.

Two housekeeping items also carried here:

Marked ready with squash auto-merge armed. Ordering constraint stands: this must land after #1174 (314ccab81), which it now does.

@justinchuby

Copy link
Copy Markdown
Owner Author

Reviewer verification — reproduced (and then some), merging

Validated on Intel Core i7-13800H (14C/20T), Windows 11, stacked on the just-merged #1183.

  • Small-shape win reproduced and larger here: via bench_half_gemm_parallel_threshold at T=4/8/16, the smallest shape (8×64×64) is ~10–12× faster going serial instead of forked (serial/par = 0.08–0.10), exceeding the claimed 3.15×. Declining the fork is unambiguously right for the shapes this kernel now serves.
  • No regression at any shape: the guard only routes below-threshold work to serial; everything >= PARALLEL_MIN_WORK takes the byte-identical split path as main. Checked 1.5M–3.1M rows — they behave exactly as base. Zero large-shape regression.
  • Honest hardware caveat (now in the body): on this box the true crossover is ~2–3M MACs, above the constant's 1_048_576 (derived on a 16-core Linux box), so shapes in ~[1M, 2.5M] still fork-and-lose as they already do on main. That is a pre-existing miss, not a new regression. I did not retune the constant to one dev box — it stays the evidence-backed shared default. Re-deriving it on reference hardware is a non-blocking follow-up.
  • Correctness: all 6 bit-for-bit / boundary / single-block / bf16 guard tests pass. Scheduling-only change.
  • No regression: cargo test --lib = 1352 passed / 0 failed / 17 ignored; clippy --all-targets -D warnings clean.

Merging by squash on this evidence.

@justinchuby
justinchuby merged commit 2d09da4 into main Aug 18, 2026
7 of 17 checks passed
@justinchuby
justinchuby deleted the squad/roy-half-gemm-guard branch August 18, 2026 14:54
justinchuby added a commit that referenced this pull request Aug 18, 2026
… (1.66x -> 1.36x vs ORT) (#1218)

## What this changes

`erf_ps`'s big-branch `R` polynomial and the `exp_ps` polynomial it
feeds are
switched from Horner's rule to **Estrin's scheme**. Horner is a six-deep
dependent
FMA chain; Estrin regroups it as `a^3·(c0 a^3 + c1 a^2 + c2 a + c3) +
(c4 a^2 + c5 a + c6)`,
three deep for two extra multiplies, whose halves the out-of-order
engine runs in
parallel. Only the two chains actually on `erf`'s critical path are
converted; the
small branch stays Horner (it runs in parallel and is off-path).

Rebased onto current `main` so it **preserves #1227** (`map_ps`'s
`#[target_feature(enable = "avx2,fma")]`). The original branch predated
#1227 and its
raw diff would have moved that attribute back onto `map_bias_ps`; that
hunk is
dropped. Only `simd_activations.rs`'s erf/exp polynomials and one
accuracy test
change here.

## Numerics — measured, not assumed (reviewer)

Estrin reassociates f32 FMAs, so it is **not** bit-identical to Horner
by
construction. The squad reviewer measured the actual error with a
two-build
dump-and-diff (base Horner build vs this Estrin build, identical input
grid of
404,045 points spanning `[-6,6]` densely, the `0.921875` split boundary,
large `|x|`,
denormals, and NaN/±inf), on **Intel Core i7-13800H, Windows 11**:

- **Estrin vs Horner: max difference = exactly 1 ULP.** 4,401 of 404,043
finite
points (1.09%) differ, every one of them by a single ULP; max absolute
diff
  5.96e-8. Nothing moves by more than one ULP anywhere in the domain.
- **vs f64 `erf` truth:** max |Horner − erf| = 5.77e-8, max |Estrin −
erf| = 6.55e-8
(worst at x≈0.933, just past the split). Both within ~1 ULP of the
correctly
rounded result; Estrin is at most ~0.8e-8 looser than Horner and never
worse than
  1 ULP from it.
- **Special values are all correct and identical between the two:**
large `|x|`
(7,8,10,20,50,100,1000,1e30,f32::MAX) saturate to ±1.0; denormals
produce
  bit-identical tiny outputs; NaN → NaN; +inf → +1.0; −inf → −1.0.

This is a conscious, documented 1-ULP change, not a hand-wave. The
accuracy gate
`erf_reassociation_costs_no_accuracy` (worst scaled error pinned at
5.97e-8 = 2^-24
over `[-6,6]`) **passes** on this box, so a future regrouping that
widens the
envelope fails instead of being absorbed by the looser `ERF_BOUND`.

## Performance — measured on this box, ORT-independent

The original body quoted "1.66× → 1.36× vs ORT" (≈ −17.7% kernel time).
I could not
reproduce that specific figure: no pinned ORT reference on this Windows
box, and it is
different hardware. Instead I timed `erf_avx2` directly, same machine,
two builds,
single-thread, min-of-200 × 9 rounds over a 1,048,576-element
mixed-range buffer:

| build | per-round min (ms) | best |
| --- | --- | ---: |
| Horner (base) | 0.6346–0.6348 (8/9 rounds) | 0.6346 |
| **Estrin (this PR)** | **0.6042–0.6043 (8/9 rounds)** | **0.6042** |

The ranges do not overlap, so the direction is solid: **Estrin is ~5.0%
faster
(1.05×)** on the erf kernel here. That is a real win but smaller than
the −17.7%
originally claimed — reported honestly as the measured kernel-internal
delta rather
than an unverified ORT ratio. The mechanism (a latency-bound chain freed
by shorter
dependency depth) is consistent with a positive but hardware-dependent
magnitude.

## Validation (reviewer)

- `cargo test -p onnx-runtime-ep-cpu --lib` — **1353 passed, 0 failed,
17 ignored**
  (base 1345 + #1183/#1149 + this test), including the accuracy gate.
- `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` —
clean.
- No MLAS feature involved; `map_ps`'s target_feature (#1227) verified
intact in the
  merged result.

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Co-authored-by: resch <resch@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 18, 2026
…-ORT headline) (#1259)

## What

Records the correction of record for **#1218** (erf's Estrin-scheme
reassociation), appended as **§32** to the CPU-EP-vs-ORT ledger — the
same document and style where the #1226 (§30.2) and #1230 withdrawals
live.

#1218 landed as *"1.66x -> 1.36x vs ORT"* (~17.7% kernel-time
reduction). On the review host that ratio is **unreproducible** — there
is no pinned ORT reference build here, and the original number came from
different hardware — so it is **withdrawn**. Stated plainly: it was
**neither reproduced nor refuted**, because it was measured against a
reference we do not run.

## What the doc now records

- **The withdrawal**, with the reason (no ORT arm on this box), and a
note that the overstated figure reached **no benchmark table, README, or
source comment** — only the (unrewritable) squash-commit subject.
- **The measured number:** ORT-independent same-binary A/B on
`erf_avx2`, single-thread, min-of-200 × 9 rounds, 1M elements, on
**Intel Core i7-13800H** — Horner 0.6346 ms → Estrin 0.6042 ms,
non-overlapping ranges → **~5.0% (1.05x)**. Real, but a fraction of the
implied ~17.7%.
- **The durable correctness fact, stated in as many words:** this
reassociation is **NOT bit-identical**. Estrin differs from Horner by
**at most exactly 1 ULP** (1.09% of 404k points, max abs 5.96e-8); vs
f64 truth Horner 5.77e-8 / Estrin 6.55e-8 (both within ~1 ULP of
correctly-rounded erf); large |x| → ±1.0, denormals bit-identical, NaN →
NaN, ±inf → ±1.0. Names the gate test
**`erf_reassociation_costs_no_accuracy`** as the tripwire.

## Validation

Documentation only — no source, no benchmarks re-run. No markdown
linter/build exists in-tree, so there is nothing to build; the change is
prose appended to an existing report.

Refs #1218. Companion to the separate threshold-tuning issue filed for
#1149.

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant