Skip to content

perf(cpu-ep): build the CPU kernels as one codegen unit (1.2-2.1x across the activation family) - #1174

Merged
justinchuby merged 1 commit into
mainfrom
squad/resch-cpu-codegen-units
Aug 18, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/resch-cpu-codegen-units

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What

Pin codegen-units = 1 for onnx-runtime-ep-cpu in the workspace release
profile.

The AVX2 elementwise kernels only vectorise when the compiler can see the
dispatcher, the chunk loop and the per-vector body together. The default
sixteen codegen units split that chain and the loops come out scalar.

#1136 found one cause of the repartition — instantiating the generic
run_chunked from another module — and fixed it by making run_chunked
private. That closed one door. The compiler was still free to split the crate
on its own, and it does.

Why this is a regression, not a tuning knob

main has been measurably slower than the ratios this repo publishes in
docs/performance/CPU_ACTIVATION_GAPS.md, with no source change to account for
it. Tanh at 1 Mi is published at 0.82 of ORT; measured on main today it is
0.41. With the pin it is 0.70. The published numbers came from a build whose
partition happened to be favourable and were not reachable on main at all.

Evidence — kernel level

activation_bench, two binaries built from identical source (default vs
CARGO_PROFILE_RELEASE_CODEGEN_UNITS=1), run interleaved 5 rounds, medians,
taskset -c 8-15. 105 cases (7 ops × 5 shapes × 3 dtypes).

case 16 CGUs 1 CGU ratio
Sqrt f32 / 4096 0.639 ns/elem 0.241 2.65x
Sqrt f32 / 3072 0.642 0.244 2.63x
Tanh f32 / 3072 0.731 0.381 1.92x
Tanh f32 / 2 Mi 0.189 0.099 1.90x
Sigmoid f32 / 3072 0.735 0.389 1.89x
Sqrt f16 / 4096 0.915 0.499 1.83x
Sigmoid f32 / 2 Mi 0.196 0.108 1.81x
QuickGelu f32 / 16 (control) 8.556 8.657 0.99x
Tanh bf16 / 16 (control) 21.149 21.305 0.99x

77 of 105 cases move by more than 1.15x. The ones that do not are the
16-element shapes — dispatch-bound, so a de-vectorized loop cannot show up in
them, and nothing else should. That split is the signature of de-vectorization,
not of noise.

Reproduced independently through the per-package override actually being landed
here ([profile.release.package.onnx-runtime-ep-cpu]), same 77/105 and same
2.65x worst case.

Evidence — session level, against ORT

Through the ORT session API, our EP vs ORT's own CPU EP in one process,
interleaved iteration by iteration, intra_op = 1 on both sides
(the harness refuses a half-pinned comparison), 31 iterations after 3 warmups,
p50/p90 of whole-Run, taskset -c 8-15.

Columns are ours_ms / ort_ms, so lower is better and below 1.00 means we
win
.

float32, 1 Mi, one thread

case before p50 after p50 before p90 after p90
Tanh 2.449 1.433 3.135 1.381
Sigmoid 2.342 1.383 2.286 1.372
Erf 2.058 1.565 2.049 1.573
Gelu (tanh) 2.020 1.397 2.017 1.379
Gelu (exact) 1.951 1.507 1.934 1.497
Exp 2.408 1.285 2.333 1.271
FastGelu 2.003 1.402 1.996 1.394
QuickGelu 1.450 0.953 1.428 0.945
Sqrt 1.530 0.718 1.508 0.734
Relu (control) 1.038 1.033 1.060 1.058

Sqrt and QuickGelu cross from loss to win. Relu is the control: it is
memory-bound at 1 Mi, so the codegen partition cannot move it, and it does not.

float32 4 Ki and float16 1 Mi, one thread

case before p50 after p50
Tanh f32 4 Ki 2.058 1.523
Sigmoid f32 4 Ki 2.051 1.503
Erf f32 4 Ki 1.905 1.598
Sqrt f32 4 Ki 1.616 1.150
Tanh f16 1 Mi 2.060 1.359
Exp f16 1 Mi 2.015 1.332

Reproduce:

NXRT_MM_BENCH=1 NXRT_MM_BENCH_THREADS=1 ONNX_GENAI_MLAS_THREADPOOL_THREADS=1 \
NXRT_MM_BENCH_CASE=f32_1m NXRT_MM_BENCH_ITERS=31 taskset -c 8-15 \
cargo test --release -p onnx-runtime-ep-cpu-plugin --test plugin_ort_e2e \
  plugin_path_ab -- --nocapture --ignored

Also in this PR

  • unary_bench_cases() — extends the session A/B harness, which until now
    covered only the matmul family, to the elementwise grid the activation doc is
    written against. Same harness, same refusal to report a half-pinned ratio, so
    a number from it is directly comparable to a matmul number. #[ignore]d like
    the rest of the harness, so no CI cost.
  • codegen_units_are_pinned — reads the setting back out of the workspace
    manifest. A build setting is exactly what a rebase drops silently: no test
    fails, no path changes, everything just gets slower. That is how this class of
    regression reached main in the first place, so it gets a test.
  • Doc update recording the before/after grid and the reproduce line.

Numerics

Unchanged — this is a compiler partitioning setting, not a code change. Full
suites green: onnx-runtime-ep-cpu 1324 passed / 0 failed, and every
onnx-runtime-ep-cpu-plugin suite passed with NXRT_REQUIRE_ORT_TESTS=1
(so ORT-dependent tests are hard failures rather than skips).

Cost

onnx-runtime-ep-cpu builds in ~65 s instead of ~19 s. Scoped to this one
package, so nothing else in the workspace changes.

…oss the activation family)

The AVX2 elementwise kernels only vectorise when the compiler can see the
dispatcher, the chunk loop and the per-vector body together. The release
profile's default sixteen codegen units splits that chain and the loops come
out scalar. #1136 found one *cause* of the repartition -- instantiating the
generic `run_chunked` from another module -- and fixed it by making
`run_chunked` private. That closed one door; the compiler was still free to
split the crate on its own, and it does.

Pin `codegen-units = 1` for `onnx-runtime-ep-cpu` alone. The rest of the
workspace keeps parallel codegen; this crate goes from ~19 s to ~65 s.

Kernel level (`activation_bench`, 5 interleaved rounds, medians, `taskset -c
8-15`): 77 of 105 cases improve by more than 1.15x, worst case `Sqrt` f32 at
4096 elements at 2.65x. The 16-element shapes are flat at 0.99-1.00x, which is
the control: they are dispatch-bound, so a de-vectorized loop cannot show up in
them, and nothing else should.

Session level, against ORT's own CPU EP through the ORT API, `intra_op = 1` on
both sides, 31 interleaved iterations, p50 of whole-`Run`, as `ours / ORT`
where above 1.00 means we are slower:

| case, 1 Mi f32 | before | after |
|---|---|---|
| `Tanh`         | 2.45 | 1.43 |
| `Sigmoid`      | 2.34 | 1.38 |
| `Erf`          | 2.06 | 1.57 |
| `Gelu` (tanh)  | 2.02 | 1.40 |
| `Gelu` (exact) | 1.95 | 1.51 |
| `Exp`          | 2.41 | 1.29 |
| `FastGelu`     | 2.00 | 1.40 |
| `QuickGelu`    | 1.45 | 0.95 |
| `Sqrt`         | 1.53 | 0.72 |
| `Relu`         | 1.04 | 1.03 |

`Sqrt` and `QuickGelu` cross from loss to win; `Relu` is the control and does
not move, being memory-bound at this size. The f16 rows move the same way
(`Tanh` 2.06 -> 1.36, `Exp` 2.02 -> 1.33), as does the 4 Ki grid (`Tanh` 2.06 ->
1.52).

This also explains a discrepancy: `main` had been measurably slower than the
ratios published in `CPU_ACTIVATION_GAPS.md` -- `Tanh` at 1 Mi was 0.41 against
a published 0.82 -- with no source change to account for it. The published
numbers were taken from a build whose partition happened to be favourable, and
were not reachable on `main` until this pin.

Two supporting pieces:

* `unary_bench_cases()` extends the session A/B harness, which until now only
  covered the matmul family, to the elementwise grid the activation doc is
  written against. Same harness, same pinning refusal, so a number from it is
  directly comparable to a matmul number.
* `codegen_units_are_pinned` reads the setting back out of the workspace
  manifest. A build setting is exactly the kind of thing a rebase drops
  silently -- no test fails, no path changes, everything just gets slower --
  which is how this class of regression reached `main` in the first place.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Reviewed by Opus (code-review agent, read-only, full diff): no blockers, one should-fix on the doc and three nits. All acted on before marking ready.

should-fix — the doc overstated the evidence. The original text said the published ratios "were unreachable until the pin landed", but the pinned column is still below the published column on seven rows (Tanh 0.82 published vs 0.70 pinned), which a strictly-improved build on one harness cannot produce — meaning published came from a different harness and does not belong in a progression narrative beside two session-level columns. Rewritten: the controlled A/B is 16 CGUs vs 1 CGU on the same commit and same harness, the claim is only the 1.4–1.9x recovery between them, and published is explicitly labelled a different measurement kept for continuity, to be replaced row by row as each is re-measured.

nits, all fixed:

  • The guard test would have failed on codegen-units = 1 # pinned — a test that fails on its own documentation teaches people to delete it. It now strips an inline comment before comparing.
  • A per-package override does not inherit from release into bench, so cargo bench would have rebuilt this crate at sixteen units and measured the de-vectorized code. [profile.bench.package.onnx-runtime-ep-cpu] added, and the guard now checks both profiles.
  • The Cargo.toml comment cited 2.1x as the extreme where the module table and this PR body both cite 2.65x; corrected.

Confirmed by the review, worth recording:

  • Scoping to one package has no downstream-monomorphization hole. The hot loops are concrete pub(crate) non-generic functions, the one generic helper (run_chunked) is deliberately confined to simd_activations.rs by fix(cpu): undo the 2.3x unary regression #1130 shipped #1136's guard, and release carries no LTO — so all of their codegen lands in this crate's units. The through-cdylib session measurements corroborate it.
  • The generated TextFormat is valid: default-domain opset_import alongside { domain: "com.microsoft" version: 1 }, domain: "" on the node, Gelu at opset 20 with type: STRING s: "tanh", shapes [1, 4096] and [1, 1048576] matching the doc's 4 K / 1 M.
  • tolerance is inert for the unary cases (the A/B measures, it does not compare); documented on unary_case so nobody mistakes these for numerics coverage.

Re-verified after the fixes: onnx-runtime-ep-cpu 1324 passed / 0 failed, guard test green, and the full onnx-runtime-ep-cpu-plugin suite green under NXRT_REQUIRE_ORT_TESTS=1.

@justinchuby
justinchuby enabled auto-merge (squash) August 18, 2026 02:46
justinchuby added a commit that referenced this pull request Aug 18, 2026
…d build actually shows

`codegen-units = 1` for `onnx-runtime-ep-cpu` (#1174) makes the *serial*
route materially faster, so the point where splitting starts to repay the
fork moves up. Re-ran `bench_half_gemm_parallel_threshold` at RAYON 2/4/8/16
with that pin in place, two runs per thread count:

|     m*k*n |  T=2 |       T=4 | T=8  |      T=16 |
|-----------|------|-----------|------|-----------|
|   262_144 | 1.20 | 0.92/0.91 | 1.52 | 0.64/0.95 |
|   393_216 | 1.28 | 0.99/0.96 | 1.66 | 0.80/1.10 |
|   524_288 | 1.32 | 0.99/0.98 | 1.81 | 1.03/1.19 |
|   786_432 | 1.33 | 1.03/0.92 | 1.35 | 0.97/1.25 |
| 1_048_576 | 1.37 | 1.05/1.05 | 1.46 | 1.31/1.34 |

The rule is unchanged -- the smallest size that wins at *every* measured
thread count in *every* run -- but the answer is now `1_048_576`, not
`524_288`: `524_288` is a wash at T=4 (0.99/0.98) and `786_432` regresses
there (0.92) and at T=16 (0.97). Below the threshold the loss is still the
0.32-0.37x the guard exists to stop, so the guard itself is unaffected.

Boundary tests are moved with the constant so they keep testing the
boundary: the "must split" shape becomes 8x512x384, the exact-threshold
shape becomes 8x512x256, and the one-block shape becomes 1x1024x2048.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/medium_generic_f32_threads=8/32x512x512 879.45 µs 2.08 ms +136.9%
🔴 gather/large_f16_threads=1-internal/131072 11.33 µs 22.82 µs +101.3%
🔴 gather/large_f32_threads=1-internal/131072 26.15 µs 49.12 µs +87.8%
🔴 gather/large_bf16_threads=1-internal/131072 12.14 µs 22.49 µs +85.3%
🔴 gather/small_bf16_threads=1-internal/4096 438.2 ns 687.3 ns +56.8%
🔴 matmul/small_generic_f32_threads=1/1x256x256 33.92 µs 49.43 µs +45.7%
🔴 gather/medium_f32_threads=1-internal/32768 3.41 µs 4.66 µs +36.6%
🔴 matmul/small_generic_f32_threads=8/1x256x256 32.44 µs 43.98 µs +35.6%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 28.72 µs 38.38 µs +33.6%
⚠️ matmul/small_generic_bf16_threads=1/1x256x256 29.56 µs 36.81 µs +24.5%
⚠️ matmul/small_generic_f16_threads=1/1x256x256 30.33 µs 37.68 µs +24.2%
⚠️ matmul/small_generic_bf16_threads=8/1x256x256 29.17 µs 35.47 µs +21.6%
⚠️ gather/medium_bf16_threads=1-internal/32768 2.28 µs 2.71 µs +18.9%
⚠️ matmul/medium_generic_f32_threads=1/32x512x512 2.19 ms 2.59 ms +18.3%
⚠️ matmul/medium_generic_bf16_threads=1/32x512x512 487.00 µs 575.35 µs +18.1%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 29.16 µs 34.23 µs +17.4%
✅ add/large_bf16_threads=1-internal/4194304 1.74 ms 1.96 ms +12.8%
✅ add/large_f16_threads=1-internal/4194304 1.70 ms 1.91 ms +12.3%
✅ add/small_f16_threads=1-internal/1024 437.3 ns 480.5 ns +9.9%
✅ matmul/small_generic_f16_threads=8/1x256x256 34.41 µs 37.74 µs +9.7%
✅ add/small_bf16_threads=1-internal/1024 429.8 ns 469.6 ns +9.2%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 364.99 µs 394.96 µs +8.2%
✅ sampling_latency/top_k_per_token 49.59 µs 53.07 µs +7.0%
✅ sampling_latency/min_p_per_token 194.27 µs 205.71 µs +5.9%
✅ add/medium_bf16_threads=1-internal/262144 104.40 µs 107.81 µs +3.3%
✅ gather/small_f32_threads=1-internal/4096 668.5 ns 689.0 ns +3.1%
✅ gather/medium_f16_threads=1-internal/32768 2.26 µs 2.33 µs +3.0%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.23 ms 3.32 ms +2.7%
✅ sampling_latency/greedy_per_token 2.98 µs 3.03 µs +1.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 1.97 ms 1.99 ms +1.1%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 483.42 µs 487.45 µs +0.8%
✅ kv_cache/alloc_dealloc_pages 36.32 µs 36.62 µs +0.8%
✅ tokenization/decode_tokens_per_second 5.74 ms 5.79 ms +0.8%
✅ sampling_latency/top_p_per_token 411.79 µs 413.43 µs +0.4%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 8.86 ms 8.85 ms -0.1%
✅ add/medium_f16_threads=1-internal/262144 110.76 µs 110.64 µs -0.1%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.33 ms 5.30 ms -0.5%
✅ add/medium_f32_threads=1-internal/262144 26.28 µs 25.82 µs -1.8%
✅ qwen3_sampling_processors/top_k_top_p_fast 617.59 µs 606.10 µs -1.9%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 80.00 µs 78.07 µs -2.4%
✅ grammar_masking/llguidance_compute_mask/32 72.26 µs 70.27 µs -2.7%
✅ logit_processing/seven_processor_chain_per_step 304.37 µs 295.57 µs -2.9%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.39 µs 13.90 µs -3.4%
✅ gather/small_f16_threads=1-internal/4096 502.8 ns 480.4 ns -4.4%
✅ tokenization/encode_tokens_per_second 378.70 µs 353.61 µs -6.6%
✅ reduce_mean/large_f32_threads=1-internal/262144 990.01 µs 914.85 µs -7.6%
✅ reduce_mean/medium_f32_threads=1-internal/65536 248.75 µs 227.44 µs -8.6%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 174.31 µs 157.19 µs -9.8%
✅ qwen3_sampling_processors/top_k_partial_selection 144.34 µs 128.60 µs -10.9%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 4.10 ms 3.54 ms -13.6%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 52.07 µs 44.83 µs -13.9%
✅ add/large_f32_threads=1-internal/4194304 751.70 µs 640.35 µs -14.8%
🟢 matmul/large_generic_bf16_threads=1/32x1024x1024 2.23 ms 1.88 ms -15.5%
🟢 add/small_f32_threads=1-internal/1024 250.3 ns 202.9 ns -19.0%
🟢 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 577.70 µs 464.64 µs -19.6%
🟢 matmul/large_generic_f16_threads=1/32x1024x1024 108.30 µs 76.11 µs -29.7%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.12 ms 1.25 ms -41.1%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 104.19 µs 41.33 µs -60.3%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.63 ms 514.16 µs -68.5%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.83 3.49 5.01 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@codecov

codecov Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.00000% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.45%. Comparing base (b160543) to head (2187e7c).
⚠️ Report is 19 commits behind head on main.

Files with missing lines Patch % Lines
...nnx-runtime-ep-cpu/src/kernels/simd_activations.rs 90.00% 2 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1174      +/-   ##
==========================================
+ Coverage   79.60%   80.45%   +0.85%     
==========================================
  Files         357      359       +2     
  Lines      153790   157353    +3563     
  Branches   153790   157353    +3563     
==========================================
+ Hits       122420   126594    +4174     
+ Misses      26850    26196     -654     
- Partials     4520     4563      +43     
Flag Coverage Δ
mlas 85.46% <ø> (?)
offline 80.35% <90.00%> (+0.75%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...nnx-runtime-ep-cpu/src/kernels/simd_activations.rs 88.59% <90.00%> (+0.01%) ⬆️

... and 23 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@justinchuby
justinchuby merged commit 314ccab into main Aug 18, 2026
14 of 19 checks passed
@justinchuby
justinchuby deleted the squad/resch-cpu-codegen-units branch August 18, 2026 04:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant