Skip to content

fix(cpu): undo the 2.3x unary regression #1130 shipped - #1136

Merged
justinchuby merged 3 commits into
mainfrom
deckard/fix-cgu-regression
Aug 17, 2026
Merged

justinchuby merged 3 commits into
mainfrom
deckard/fix-cgu-regression

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 17, 2026 •

Copy link
Copy Markdown
Owner

What happened

#1130 (mine) wrapped the Clip MLAS call in run_chunked so it would use the
thread pool. run_chunked is generic, so doing that from selection.rs created a new
instantiation in a module that had never had one.

The runtime path of every other unary op was untouched — not one instruction — and all
1332 tests still passed. But the crate's codegen units repartitioned and the AVX2 unary
kernels in simd_activations.rs stopped being vectorised.

This shipped. It was found while collecting numbers for the final report: Sqrt had
gone from beating ORT by 1.9× to losing at 0.7×, and the ORT-relative position of five
ops had collapsed in a way no code change explained.

Evidence

Bisected by rebuilding main with one file at a time reverted to 1f1ce4b74
(the commit before #1130). n = 65536, 1 thread, taskset -c 8-23, 4 interleaved
rounds, µs p50. Only selection.rs matters — reverting relu.rs or conv.rs
alone restored nothing:

op main (#1130) revert relu.rs revert selection.rs revert conv.rs pre-#1130
Sqrt 48.5 48.4 21.3 44.6 21.5
Tanh 55.2 57.0 31.6 60.0 31.8
Sigmoid 56.9 55.6 31.0 57.1 31.1
QuickGelu 64.5 64.5 42.5 64.2 42.5
FastGelu 79.6 79.3 55.5 79.3 55.8
Erf 62.2 62.4 62.2 62.2 62.5
Relu 21.3 26.5 22.6 23.0 20.4

Reproduced independently by building the same commit in a second worktree, so it is not
a build-directory artefact. Adding #[inline] to run_chunked did not help, which
rules out a plain inlining decision and points at codegen-unit partitioning.

The diagnosis, confirmed mechanically

If the cause really is codegen-unit partitioning, then forcing the crate into a single
codegen unit should erase the regression with no source change at all. It does.
CARGO_PROFILE_RELEASE_CODEGEN_UNITS=1 on the unfixed commit 34095af0f, same
machine, 4 interleaved rounds, µs p50:

op n main, default CGUs main, codegen-units=1 this PR, default CGUs
Sqrt 64 Ki 48.3 21.6 21.2
Sqrt 1 Mi 667.4 261.4 274.4
Sigmoid 1 Mi 813.3 504.8 441.6
QuickGelu 1 Mi 929.4 575.8 573.0
FastGelu 1 Mi 1171.2 788.2 808.5

So the diagnosis is not inferred from a bisect alone — the proposed mechanism, applied
directly, reproduces the cure.

codegen-units = 1 (or LTO) in the release profile would remove this whole fragility
class permanently, and is the more durable answer. It is deliberately not in this PR:
it is a workspace-wide build-policy change that affects every crate and every
contributor's build time (the plugin alone went 14 s to 63 s here), it would need its own
measurement across the whole EP rather than the activation kernels, and it does not
belong in a regression fix. Filed as the follow-up this PR's limitation section points
at. The source fix costs nothing and is independent of it.

The fix

run_chunked is private to simd_activations.rs again — the compiler now enforces the
rule, not a convention. Callers elsewhere go through one of two entry points that are
instantiated in that module:

  • run_chunked_fn(input, output, body: fn(&[f32], &mut [f32])) — deliberately a fn
    pointer, not impl Fn, so every caller shares one instantiation. Used by Relu and
    SiLU.
  • clip_chunked(input, output, min, max) — Clip needs captured bounds. It takes the
    serial decision itself so the short case is a direct call rather than one through a
    closure the optimiser can no longer see into; without that, Clip itself paid 12%.

Result

n = 65536 and 1 Mi, 1 thread, 5 interleaved rounds, µs p50:

op n main (#1130) this PR pre-#1130
Sqrt 64 Ki 48.5 21.4 21.7
Sqrt 1 Mi 667.4 259.2 275.1
Tanh 1 Mi 776.1 437.8 436.4
Sigmoid 1 Mi 807.6 440.1 502.2
QuickGelu 1 Mi 923.1 574.7 607.0
FastGelu 1 Mi 1166.9 785.3 788.1
Clip 1 Mi 268.7 267.5 267.4
Swish 1 Mi 728.8 720.8 1385.6

Everything is back to its pre-#1130 level and #1130's own win is kept: Swish is
still 1.92× faster than before #1130, and Clip/Relu still reach the pool at
≥ PAR_MIN_LEN.

Regression guard

chunking_instantiation_is_local::no_module_outside_this_one_instantiates_run_chunked
walks the crate source and fails if any module other than simd_activations.rs
instantiates run_chunked, naming the offending file and line.

Verified to falsify: re-widening the visibility and pointing relu.rs:144 back at
run_chunked makes it fail with
Offending call sites: ["…/kernels/relu.rs:144"].

This class of bug produces no wrong answers, no test failures and no diff in the file
that slows down, so a mechanical guard is the only thing that catches it.

Tests

  • cargo test -p onnx-runtime-ep-cpu --features mlas --lib → 1332 passed, 0 failed
  • cargo test -p onnx-runtime-ep-cpu --lib → 1305 passed, 0 failed
  • cargo fmt --all clean.
  • No numerical change: clip_chunked's serial branch calls exactly the function the
    closure called, and run_chunked_fn forwards unchanged.

Limitations

  • The mechanism is codegen-unit partitioning, which the compiler makes no promises
    about. The guard encodes the rule that was measured to work on this toolchain; it
    cannot prove the next refactor is safe. That is why the guard names the symptom and
    the measured cost in its failure message.
  • Measured on one machine (AMD EPYC 9V74, AVX2/FMA/F16C, taskset-pinned, 5 rounds).
    The direction is unambiguous — up to 2.3× on Sqrt and roughly 1.5–1.8× on the
    other four — but the exact figures are not portable.

#1130 wrapped `Clip` in `run_chunked` from `selection.rs`. `run_chunked`
is generic, so that added an instantiation in a new module. The runtime
path of every other unary op was untouched, and every test still passed,
but the crate's codegen units repartitioned and the AVX2 unary kernels in
simd_activations.rs stopped being vectorised.

Measured at n = 65536, one thread, same commit, only that instantiation
moved:

  Sqrt       21.5 -> 48.5 us
  Tanh       30.4 -> 54.7 us
  Sigmoid    32.2 -> 57.1 us
  QuickGelu  42.6 -> 64.6 us
  FastGelu   55.7 -> 79.2 us

Bisected file by file: reverting selection.rs alone restored all five,
and reverting relu.rs or conv.rs alone restored none of them.

`run_chunked` is private to simd_activations.rs again. Callers elsewhere
go through `run_chunked_fn`, which is not generic, or `clip_chunked`,
both instantiated in that module. `clip_chunked` takes the serial
decision itself so the short case stays a direct call.

A test now walks the crate source and fails if any other module
instantiates `run_chunked`. Verified to falsify.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.90909% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 79.94%. Comparing base (34095af) to head (34a50b2).
⚠️ Report is 7 commits behind head on main.

Files with missing lines Patch % Lines
...nnx-runtime-ep-cpu/src/kernels/simd_activations.rs 90.90% 1 Missing and 2 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff            @@
##             main    #1136    +/-   ##
========================================
  Coverage   79.94%   79.94%            
========================================
  Files         368      368            
  Lines      160611   160826   +215     
  Branches   160611   160826   +215     
========================================
+ Hits       128395   128579   +184     
- Misses      27499    27525    +26     
- Partials     4717     4722     +5     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.40% <ø> (+0.09%) ⬆️
mlas 85.64% <ø> (-0.10%) ⬇️
offline 79.69% <90.90%> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...tes/onnx-runtime-ep-cpu/src/kernels/activations.rs 91.71% <ø> (-0.21%) ⬇️
crates/onnx-runtime-ep-cpu/src/kernels/relu.rs 90.74% <ø> (ø)
...rates/onnx-runtime-ep-cpu/src/kernels/selection.rs 82.64% <ø> (ø)
...nnx-runtime-ep-cpu/src/kernels/simd_activations.rs 97.11% <90.90%> (-0.12%) ⬇️

... and 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 58.43 µs 106.53 µs +82.3%
🔴 grammar_masking/llguidance_compute_mask/32 83.32 µs 117.59 µs +41.1%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 83.73 µs 117.19 µs +40.0%
🔴 gather/small_f16_threads=1-internal/4096 518.5 ns 697.3 ns +34.5%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 1.29 ms 1.71 ms +33.0%
⚠️ matmul/large_generic_f32_threads=8/32x1024x1024 5.05 ms 6.52 ms +29.0%
⚠️ sampling_latency/min_p_per_token 214.97 µs 268.08 µs +24.7%
⚠️ gather/large_f16_threads=1-internal/131072 16.49 µs 20.32 µs +23.3%
⚠️ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 777.68 µs 956.81 µs +23.0%
⚠️ matmul/large_generic_f32_threads=1/32x1024x1024 9.02 ms 11.02 ms +22.1%
⚠️ sampling_latency/top_p_per_token 392.85 µs 471.22 µs +20.0%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 585.79 µs 667.19 µs +13.9%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.06 µs 18.05 µs +12.4%
✅ gather/large_f32_threads=1-internal/131072 39.44 µs 44.26 µs +12.2%
✅ reduce_mean/medium_f32_threads=1-internal/65536 259.42 µs 289.21 µs +11.5%
✅ sampling_latency/top_k_per_token 52.83 µs 58.46 µs +10.7%
✅ matmul/medium_generic_f16_threads=8/32x512x512 49.85 µs 54.80 µs +9.9%
✅ add/large_bf16_threads=1-internal/4194304 1.78 ms 1.95 ms +9.8%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.82 ms 4.17 ms +9.2%
✅ qwen3_sampling_processors/top_k_partial_selection 154.71 µs 162.11 µs +4.8%
✅ logit_processing/seven_processor_chain_per_step 331.32 µs 344.83 µs +4.1%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 88.22 µs 90.80 µs +2.9%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.20 ms 2.24 ms +2.0%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 522.81 µs 533.33 µs +2.0%
✅ matmul/small_generic_f16_threads=1/1x256x256 36.61 µs 37.23 µs +1.7%
✅ matmul/small_generic_f32_threads=8/1x256x256 55.59 µs 56.27 µs +1.2%
✅ gather/small_bf16_threads=1-internal/4096 520.2 ns 517.4 ns -0.5%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 187.20 µs 186.17 µs -0.6%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 508.18 µs 501.66 µs -1.3%
✅ add/medium_f16_threads=1-internal/262144 111.99 µs 109.75 µs -2.0%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 606.50 µs 594.30 µs -2.0%
✅ add/small_bf16_threads=1-internal/1024 479.9 ns 469.1 ns -2.2%
✅ gather/small_f32_threads=1-internal/4096 761.4 ns 743.3 ns -2.4%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.33 ms 6.11 ms -3.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 724.37 µs 690.59 µs -4.7%
✅ matmul/small_generic_f16_threads=8/1x256x256 35.48 µs 33.54 µs -5.5%
✅ gather/medium_f32_threads=1-internal/32768 4.19 µs 3.95 µs -5.9%
✅ tokenization/encode_tokens_per_second 416.90 µs 390.31 µs -6.4%
✅ sampling_latency/greedy_per_token 3.44 µs 3.21 µs -6.6%
✅ matmul/medium_generic_f16_threads=1/32x512x512 39.18 µs 36.32 µs -7.3%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.29 ms 2.12 ms -7.7%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.13 ms 1.03 ms -8.2%
✅ gather/medium_bf16_threads=1-internal/32768 2.91 µs 2.66 µs -8.5%
✅ gather/large_bf16_threads=1-internal/131072 18.24 µs 16.58 µs -9.1%
✅ add/large_f32_threads=1-internal/4194304 675.98 µs 614.22 µs -9.1%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 106.64 µs 96.69 µs -9.3%
✅ add/medium_f32_threads=1-internal/262144 27.68 µs 25.03 µs -9.6%
✅ add/large_f16_threads=1-internal/4194304 1.77 ms 1.57 ms -11.5%
✅ add/small_f16_threads=1-internal/1024 530.5 ns 468.9 ns -11.6%
✅ matmul/small_generic_f32_threads=1/1x256x256 45.66 µs 40.25 µs -11.9%
✅ matmul/small_generic_bf16_threads=8/1x256x256 48.20 µs 42.47 µs -11.9%
✅ kv_cache/alloc_dealloc_pages 47.29 µs 40.82 µs -13.7%
✅ tokenization/decode_tokens_per_second 6.72 ms 5.76 ms -14.3%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 45.42 µs 37.56 µs -17.3%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.10 ms 1.67 ms -20.5%
🟢 add/medium_bf16_threads=1-internal/262144 137.35 µs 108.38 µs -21.1%
🟢 add/small_f32_threads=1-internal/1024 277.7 ns 189.4 ns -31.8%
🟢 matmul/medium_generic_f32_threads=1/32x512x512 3.54 ms 2.37 ms -32.9%
🟢 gather/medium_f16_threads=1-internal/32768 3.98 µs 2.47 µs -37.9%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 8.11 4.38 5.87 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

deckard and others added 2 commits August 17, 2026 18:25
Independent review findings on #1136.

MINOR: `run_chunked_fn` was ungated while both callers are mlas-gated, so
the non-mlas build warned it was never used. Gated to match.

MINOR: inserting the two new functions left `parallel_dispatches`' doc
comment attached to `run_chunked_fn`, where it contradicted the function
below it. Moved back.

MINOR: the source guard matched the bare substring `run_chunked(`, which
missed `run_chunked ::<T>(` and would trip on prose. It now matches the
identifier on a word boundary and requires the next token to open a call
or a turbofish, so it rejects `run_chunked_fn`, `run_chunked_rows`, test
names and string literals. Verified to falsify against the turbofish
spelling.

NIT: the failure message said 2.3x for all three of Sqrt, Tanh and
Sigmoid; only Sqrt is 2.3x. Corrected to "up to 2.3x".

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Two NITs from the re-review.

The guard's comment claimed it rejects prose inside string literals. It
does not: it is a substring heuristic, so a call reached through an
aliased import, split across two lines or generated by a macro slips
past, and the literal text inside a string would trip it. The comment now
says so, and says what it does catch, which is the accidental case.

The PR body said "2.3x on five ops". Only Sqrt is 2.3x; the other four
are 1.5-1.8x. Same imprecision that was already corrected in the guard's
failure message.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 17, 2026 19:25
@justinchuby
justinchuby merged commit fdeaf77 into main Aug 17, 2026
13 of 18 checks passed
@justinchuby
justinchuby deleted the deckard/fix-cgu-regression branch August 17, 2026 19:25
justinchuby added a commit that referenced this pull request Aug 18, 2026
…oss the activation family)

The AVX2 elementwise kernels only vectorise when the compiler can see the
dispatcher, the chunk loop and the per-vector body together. The release
profile's default sixteen codegen units splits that chain and the loops come
out scalar. #1136 found one *cause* of the repartition -- instantiating the
generic `run_chunked` from another module -- and fixed it by making
`run_chunked` private. That closed one door; the compiler was still free to
split the crate on its own, and it does.

Pin `codegen-units = 1` for `onnx-runtime-ep-cpu` alone. The rest of the
workspace keeps parallel codegen; this crate goes from ~19 s to ~65 s.

Kernel level (`activation_bench`, 5 interleaved rounds, medians, `taskset -c
8-15`): 77 of 105 cases improve by more than 1.15x, worst case `Sqrt` f32 at
4096 elements at 2.65x. The 16-element shapes are flat at 0.99-1.00x, which is
the control: they are dispatch-bound, so a de-vectorized loop cannot show up in
them, and nothing else should.

Session level, against ORT's own CPU EP through the ORT API, `intra_op = 1` on
both sides, 31 interleaved iterations, p50 of whole-`Run`, as `ours / ORT`
where above 1.00 means we are slower:

| case, 1 Mi f32 | before | after |
|---|---|---|
| `Tanh`         | 2.45 | 1.43 |
| `Sigmoid`      | 2.34 | 1.38 |
| `Erf`          | 2.06 | 1.57 |
| `Gelu` (tanh)  | 2.02 | 1.40 |
| `Gelu` (exact) | 1.95 | 1.51 |
| `Exp`          | 2.41 | 1.29 |
| `FastGelu`     | 2.00 | 1.40 |
| `QuickGelu`    | 1.45 | 0.95 |
| `Sqrt`         | 1.53 | 0.72 |
| `Relu`         | 1.04 | 1.03 |

`Sqrt` and `QuickGelu` cross from loss to win; `Relu` is the control and does
not move, being memory-bound at this size. The f16 rows move the same way
(`Tanh` 2.06 -> 1.36, `Exp` 2.02 -> 1.33), as does the 4 Ki grid (`Tanh` 2.06 ->
1.52).

This also explains a discrepancy: `main` had been measurably slower than the
ratios published in `CPU_ACTIVATION_GAPS.md` -- `Tanh` at 1 Mi was 0.41 against
a published 0.82 -- with no source change to account for it. The published
numbers were taken from a build whose partition happened to be favourable, and
were not reachable on `main` until this pin.

Two supporting pieces:

* `unary_bench_cases()` extends the session A/B harness, which until now only
  covered the matmul family, to the elementwise grid the activation doc is
  written against. Same harness, same pinning refusal, so a number from it is
  directly comparable to a matmul number.
* `codegen_units_are_pinned` reads the setting back out of the workspace
  manifest. A build setting is exactly the kind of thing a rebase drops
  silently -- no test fails, no path changes, everything just gets slower --
  which is how this class of regression reached `main` in the first place.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 18, 2026
…oss the activation family) (#1174)

## What

Pin `codegen-units = 1` for `onnx-runtime-ep-cpu` in the workspace
release
profile.

The AVX2 elementwise kernels only vectorise when the compiler can see
the
dispatcher, the chunk loop and the per-vector body together. The default
sixteen codegen units split that chain and the loops come out
**scalar**.

#1136 found one *cause* of the repartition — instantiating the generic
`run_chunked` from another module — and fixed it by making `run_chunked`
private. That closed one door. The compiler was still free to split the
crate
on its own, and it does.

## Why this is a regression, not a tuning knob

`main` has been measurably slower than the ratios this repo publishes in
`docs/performance/CPU_ACTIVATION_GAPS.md`, with no source change to
account for
it. `Tanh` at 1 Mi is published at 0.82 of ORT; measured on `main` today
it is
**0.41**. With the pin it is 0.70. The published numbers came from a
build whose
partition happened to be favourable and were not reachable on `main` at
all.

## Evidence — kernel level

`activation_bench`, two binaries built from **identical source**
(default vs
`CARGO_PROFILE_RELEASE_CODEGEN_UNITS=1`), run interleaved 5 rounds,
medians,
`taskset -c 8-15`. 105 cases (7 ops × 5 shapes × 3 dtypes).

| case | 16 CGUs | 1 CGU | ratio |
|---|---|---|---|
| `Sqrt` f32 / 4096 | 0.639 ns/elem | 0.241 | **2.65x** |
| `Sqrt` f32 / 3072 | 0.642 | 0.244 | 2.63x |
| `Tanh` f32 / 3072 | 0.731 | 0.381 | 1.92x |
| `Tanh` f32 / 2 Mi | 0.189 | 0.099 | 1.90x |
| `Sigmoid` f32 / 3072 | 0.735 | 0.389 | 1.89x |
| `Sqrt` f16 / 4096 | 0.915 | 0.499 | 1.83x |
| `Sigmoid` f32 / 2 Mi | 0.196 | 0.108 | 1.81x |
| `QuickGelu` f32 / **16** (control) | 8.556 | 8.657 | 0.99x |
| `Tanh` bf16 / **16** (control) | 21.149 | 21.305 | 0.99x |

**77 of 105 cases move by more than 1.15x.** The ones that do not are
the
16-element shapes — dispatch-bound, so a de-vectorized loop cannot show
up in
them, and nothing else should. That split is the signature of
de-vectorization,
not of noise.

Reproduced independently through the per-package override actually being
landed
here (`[profile.release.package.onnx-runtime-ep-cpu]`), same 77/105 and
same
2.65x worst case.

## Evidence — session level, against ORT

Through the ORT session API, our EP vs ORT's own CPU EP in one process,
interleaved iteration by iteration, `intra_op = 1` on **both** sides
(the harness refuses a half-pinned comparison), 31 iterations after 3
warmups,
p50/p90 of whole-`Run`, `taskset -c 8-15`.

Columns are `ours_ms / ort_ms`, so **lower is better and below 1.00
means we
win**.

### float32, 1 Mi, one thread

| case | before p50 | after p50 | before p90 | after p90 |
|---|---|---|---|---|
| `Tanh` | 2.449 | **1.433** | 3.135 | 1.381 |
| `Sigmoid` | 2.342 | **1.383** | 2.286 | 1.372 |
| `Erf` | 2.058 | **1.565** | 2.049 | 1.573 |
| `Gelu` (tanh) | 2.020 | **1.397** | 2.017 | 1.379 |
| `Gelu` (exact) | 1.951 | **1.507** | 1.934 | 1.497 |
| `Exp` | 2.408 | **1.285** | 2.333 | 1.271 |
| `FastGelu` | 2.003 | **1.402** | 1.996 | 1.394 |
| `QuickGelu` | 1.450 | **0.953** | 1.428 | 0.945 |
| `Sqrt` | 1.530 | **0.718** | 1.508 | 0.734 |
| `Relu` (control) | 1.038 | 1.033 | 1.060 | 1.058 |

`Sqrt` and `QuickGelu` cross from loss to **win**. `Relu` is the
control: it is
memory-bound at 1 Mi, so the codegen partition cannot move it, and it
does not.

### float32 4 Ki and float16 1 Mi, one thread

| case | before p50 | after p50 |
|---|---|---|
| `Tanh` f32 4 Ki | 2.058 | **1.523** |
| `Sigmoid` f32 4 Ki | 2.051 | **1.503** |
| `Erf` f32 4 Ki | 1.905 | **1.598** |
| `Sqrt` f32 4 Ki | 1.616 | **1.150** |
| `Tanh` f16 1 Mi | 2.060 | **1.359** |
| `Exp` f16 1 Mi | 2.015 | **1.332** |

Reproduce:

```sh
NXRT_MM_BENCH=1 NXRT_MM_BENCH_THREADS=1 ONNX_GENAI_MLAS_THREADPOOL_THREADS=1 \
NXRT_MM_BENCH_CASE=f32_1m NXRT_MM_BENCH_ITERS=31 taskset -c 8-15 \
cargo test --release -p onnx-runtime-ep-cpu-plugin --test plugin_ort_e2e \
  plugin_path_ab -- --nocapture --ignored
```

## Also in this PR

- **`unary_bench_cases()`** — extends the session A/B harness, which
until now
covered only the matmul family, to the elementwise grid the activation
doc is
written against. Same harness, same refusal to report a half-pinned
ratio, so
a number from it is directly comparable to a matmul number. `#[ignore]`d
like
  the rest of the harness, so no CI cost.
- **`codegen_units_are_pinned`** — reads the setting back out of the
workspace
manifest. A build setting is exactly what a rebase drops silently: no
test
fails, no path changes, everything just gets slower. That is how this
class of
  regression reached `main` in the first place, so it gets a test.
- Doc update recording the before/after grid and the reproduce line.

## Numerics

Unchanged — this is a compiler partitioning setting, not a code change.
Full
suites green: `onnx-runtime-ep-cpu` 1324 passed / 0 failed, and every
`onnx-runtime-ep-cpu-plugin` suite passed with
`NXRT_REQUIRE_ORT_TESTS=1`
(so ORT-dependent tests are hard failures rather than skips).

## Cost

`onnx-runtime-ep-cpu` builds in ~65 s instead of ~19 s. Scoped to this
one
package, so nothing else in the workspace changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants