Repository navigation
perf(cpu-ep): build the CPU kernels as one codegen unit (1.2-2.1x across the activation family) - #1174
Conversation
…oss the activation family) The AVX2 elementwise kernels only vectorise when the compiler can see the dispatcher, the chunk loop and the per-vector body together. The release profile's default sixteen codegen units splits that chain and the loops come out scalar. #1136 found one *cause* of the repartition -- instantiating the generic `run_chunked` from another module -- and fixed it by making `run_chunked` private. That closed one door; the compiler was still free to split the crate on its own, and it does. Pin `codegen-units = 1` for `onnx-runtime-ep-cpu` alone. The rest of the workspace keeps parallel codegen; this crate goes from ~19 s to ~65 s. Kernel level (`activation_bench`, 5 interleaved rounds, medians, `taskset -c 8-15`): 77 of 105 cases improve by more than 1.15x, worst case `Sqrt` f32 at 4096 elements at 2.65x. The 16-element shapes are flat at 0.99-1.00x, which is the control: they are dispatch-bound, so a de-vectorized loop cannot show up in them, and nothing else should. Session level, against ORT's own CPU EP through the ORT API, `intra_op = 1` on both sides, 31 interleaved iterations, p50 of whole-`Run`, as `ours / ORT` where above 1.00 means we are slower: | case, 1 Mi f32 | before | after | |---|---|---| | `Tanh` | 2.45 | 1.43 | | `Sigmoid` | 2.34 | 1.38 | | `Erf` | 2.06 | 1.57 | | `Gelu` (tanh) | 2.02 | 1.40 | | `Gelu` (exact) | 1.95 | 1.51 | | `Exp` | 2.41 | 1.29 | | `FastGelu` | 2.00 | 1.40 | | `QuickGelu` | 1.45 | 0.95 | | `Sqrt` | 1.53 | 0.72 | | `Relu` | 1.04 | 1.03 | `Sqrt` and `QuickGelu` cross from loss to win; `Relu` is the control and does not move, being memory-bound at this size. The f16 rows move the same way (`Tanh` 2.06 -> 1.36, `Exp` 2.02 -> 1.33), as does the 4 Ki grid (`Tanh` 2.06 -> 1.52). This also explains a discrepancy: `main` had been measurably slower than the ratios published in `CPU_ACTIVATION_GAPS.md` -- `Tanh` at 1 Mi was 0.41 against a published 0.82 -- with no source change to account for it. The published numbers were taken from a build whose partition happened to be favourable, and were not reachable on `main` until this pin. Two supporting pieces: * `unary_bench_cases()` extends the session A/B harness, which until now only covered the matmul family, to the elementwise grid the activation doc is written against. Same harness, same pinning refusal, so a number from it is directly comparable to a matmul number. * `codegen_units_are_pinned` reads the setting back out of the workspace manifest. A build setting is exactly the kind of thing a rebase drops silently -- no test fails, no path changes, everything just gets slower -- which is how this class of regression reached `main` in the first place. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
699dd91 to
2187e7c
Compare
|
Reviewed by Opus (code-review agent, read-only, full diff): no blockers, one should-fix on the doc and three nits. All acted on before marking ready. should-fix — the doc overstated the evidence. The original text said the published ratios "were unreachable until the pin landed", but the pinned column is still below the published column on seven rows ( nits, all fixed:
Confirmed by the review, worth recording:
Re-verified after the fixes: |
…d build actually shows `codegen-units = 1` for `onnx-runtime-ep-cpu` (#1174) makes the *serial* route materially faster, so the point where splitting starts to repay the fork moves up. Re-ran `bench_half_gemm_parallel_threshold` at RAYON 2/4/8/16 with that pin in place, two runs per thread count: | m*k*n | T=2 | T=4 | T=8 | T=16 | |-----------|------|-----------|------|-----------| | 262_144 | 1.20 | 0.92/0.91 | 1.52 | 0.64/0.95 | | 393_216 | 1.28 | 0.99/0.96 | 1.66 | 0.80/1.10 | | 524_288 | 1.32 | 0.99/0.98 | 1.81 | 1.03/1.19 | | 786_432 | 1.33 | 1.03/0.92 | 1.35 | 0.97/1.25 | | 1_048_576 | 1.37 | 1.05/1.05 | 1.46 | 1.31/1.34 | The rule is unchanged -- the smallest size that wins at *every* measured thread count in *every* run -- but the answer is now `1_048_576`, not `524_288`: `524_288` is a wash at T=4 (0.99/0.98) and `786_432` regresses there (0.92) and at T=16 (0.97). Below the threshold the loss is still the 0.32-0.37x the guard exists to stop, so the guard itself is unaffected. Boundary tests are moved with the constant so they keep testing the boundary: the "must split" shape becomes 8x512x384, the exact-threshold shape becomes 8x512x256, and the one-block shape becomes 1x1024x2048. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1174 +/- ##
==========================================
+ Coverage 79.60% 80.45% +0.85%
==========================================
Files 357 359 +2
Lines 153790 157353 +3563
Branches 153790 157353 +3563
==========================================
+ Hits 122420 126594 +4174
+ Misses 26850 26196 -654
- Partials 4520 4563 +43
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
What
Pin
codegen-units = 1foronnx-runtime-ep-cpuin the workspace releaseprofile.
The AVX2 elementwise kernels only vectorise when the compiler can see the
dispatcher, the chunk loop and the per-vector body together. The default
sixteen codegen units split that chain and the loops come out scalar.
#1136 found one cause of the repartition — instantiating the generic
run_chunkedfrom another module — and fixed it by makingrun_chunkedprivate. That closed one door. The compiler was still free to split the crate
on its own, and it does.
Why this is a regression, not a tuning knob
mainhas been measurably slower than the ratios this repo publishes indocs/performance/CPU_ACTIVATION_GAPS.md, with no source change to account forit.
Tanhat 1 Mi is published at 0.82 of ORT; measured onmaintoday it is0.41. With the pin it is 0.70. The published numbers came from a build whose
partition happened to be favourable and were not reachable on
mainat all.Evidence — kernel level
activation_bench, two binaries built from identical source (default vsCARGO_PROFILE_RELEASE_CODEGEN_UNITS=1), run interleaved 5 rounds, medians,taskset -c 8-15. 105 cases (7 ops × 5 shapes × 3 dtypes).Sqrtf32 / 4096Sqrtf32 / 3072Tanhf32 / 3072Tanhf32 / 2 MiSigmoidf32 / 3072Sqrtf16 / 4096Sigmoidf32 / 2 MiQuickGeluf32 / 16 (control)Tanhbf16 / 16 (control)77 of 105 cases move by more than 1.15x. The ones that do not are the
16-element shapes — dispatch-bound, so a de-vectorized loop cannot show up in
them, and nothing else should. That split is the signature of de-vectorization,
not of noise.
Reproduced independently through the per-package override actually being landed
here (
[profile.release.package.onnx-runtime-ep-cpu]), same 77/105 and same2.65x worst case.
Evidence — session level, against ORT
Through the ORT session API, our EP vs ORT's own CPU EP in one process,
interleaved iteration by iteration,
intra_op = 1on both sides(the harness refuses a half-pinned comparison), 31 iterations after 3 warmups,
p50/p90 of whole-
Run,taskset -c 8-15.Columns are
ours_ms / ort_ms, so lower is better and below 1.00 means wewin.
float32, 1 Mi, one thread
TanhSigmoidErfGelu(tanh)Gelu(exact)ExpFastGeluQuickGeluSqrtRelu(control)SqrtandQuickGelucross from loss to win.Reluis the control: it ismemory-bound at 1 Mi, so the codegen partition cannot move it, and it does not.
float32 4 Ki and float16 1 Mi, one thread
Tanhf32 4 KiSigmoidf32 4 KiErff32 4 KiSqrtf32 4 KiTanhf16 1 MiExpf16 1 MiReproduce:
NXRT_MM_BENCH=1 NXRT_MM_BENCH_THREADS=1 ONNX_GENAI_MLAS_THREADPOOL_THREADS=1 \ NXRT_MM_BENCH_CASE=f32_1m NXRT_MM_BENCH_ITERS=31 taskset -c 8-15 \ cargo test --release -p onnx-runtime-ep-cpu-plugin --test plugin_ort_e2e \ plugin_path_ab -- --nocapture --ignoredAlso in this PR
unary_bench_cases()— extends the session A/B harness, which until nowcovered only the matmul family, to the elementwise grid the activation doc is
written against. Same harness, same refusal to report a half-pinned ratio, so
a number from it is directly comparable to a matmul number.
#[ignore]d likethe rest of the harness, so no CI cost.
codegen_units_are_pinned— reads the setting back out of the workspacemanifest. A build setting is exactly what a rebase drops silently: no test
fails, no path changes, everything just gets slower. That is how this class of
regression reached
mainin the first place, so it gets a test.Numerics
Unchanged — this is a compiler partitioning setting, not a code change. Full
suites green:
onnx-runtime-ep-cpu1324 passed / 0 failed, and everyonnx-runtime-ep-cpu-pluginsuite passed withNXRT_REQUIRE_ORT_TESTS=1(so ORT-dependent tests are hard failures rather than skips).
Cost
onnx-runtime-ep-cpubuilds in ~65 s instead of ~19 s. Scoped to this onepackage, so nothing else in the workspace changes.