Repository navigation
perf(cpu-ep): AVX2 kernels for LeakyRelu, HardSigmoid, ThresholdedRelu and Selu - #1243
Conversation
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
…u and Selu Give the last four activations on the generic path a slice kernel and route them through the zero-copy contiguous_f32 fast path, eliminating the widen / map / collect / copy the generic path incurred. Selu also gains a genuinely vectorised exponential; the other three are compare-and-blend. NaN and signed-zero semantics settled against ORT and pinned: HardSigmoid's clamp operand order keeps NaN (falsifier: flipping it makes NaN -> 1 -> 0), and Selu's -0 goes down the exp branch returning +0 to match ORT (falsifier: putting selu_scalar back on x > 0 makes scalar and vector paths disagree). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
7f2ee11 to
2c5e84d
Compare
Reviewer verification — validated and mergingReproduced on Intel Core i7-13800H (14C/20T, AVX2+FMA, no AVX-512), Windows 11; ORT-independent same-binary A/Bs (no ORT build on this box). Where the win comes from (attributed honestly):
So for LeakyRelu/HardSigmoid/ThresholdedRelu the speedup is the plumbing, not the intrinsics; for Selu it is both. The PR's original AMD EPYC 9V74 Correctness falsifiers confirmed load-bearing: HardSigmoid clamp operand order pins NaN→NaN; Selu Suite: lib activation tests 1415 passed / 0 failed / 17 ignored; clippy Merging via squash. |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1243 +/- ##
=========================================
+ Coverage 0 80.70% +80.70%
=========================================
Files 0 363 +363
Lines 0 160682 +160682
Branches 0 160682 +160682
=========================================
+ Hits 0 129680 +129680
- Misses 0 26345 +26345
- Partials 0 4657 +4657
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
AVX2 kernels for LeakyRelu, HardSigmoid, ThresholdedRelu and Selu
LeakyRelu,HardSigmoid,ThresholdedReluandSeluhad no slice kernel, sothey took
ActivationKernel's generic path: widen into aVec,map()aper-element
match, collect into a secondVec, copy that into the outputtensor. This PR gives each a slice kernel and routes them through the zero-copy
contiguous_f32fast path (established in #1240), and adds a genuinelyvectorised exponential for Selu.
Stacked on #1240 (which is stacked on #1235).
Reviewer measurements (Intel Core i7-13800H) — where the win actually comes from
Measured by the reviewer on an Intel Core i7-13800H (14C/20T, AVX2+FMA, no
AVX-512), Windows 11. No ORT build is available on this box, so these are
ORT-independent, same-binary A/Bs, 1M f32 elements, single-thread, min-of-200 ×
9 rounds.
Two different comparisons, because they answer different questions:
Fast-path routing vs the generic path (what this PR changes at the call
site). Emulating the old generic path (
to_vec→map→copy_from_slice)against the new
*_f32_slicefast path, for LeakyRelu:This — eliminating two allocations and two extra passes — is the dominant win
for the three cheap ops.
The AVX2 kernel vs its scalar reference (isolating the hand-written SIMD
arithmetic alone, both in the same release binary):
For the three cheap ops the hand-written AVX2 kernel is no faster than the
scalar loop, because LLVM already auto-vectorises a compare-and-blend in
release mode. Selu is the exception: its exponential does not auto-vectorise,
so the vector kernel is a real ~4.2× on the arithmetic itself.
Honest attribution: for LeakyRelu/HardSigmoid/ThresholdedRelu the speedup is
the fast-path plumbing, not the SIMD; for Selu it is both. That is worth stating
plainly so nobody credits the compare-and-blend intrinsics with a win the
allocator elimination actually produced.
Peer measurement (original PR, AMD EPYC 9V74)
Originally measured on an AMD EPYC 9V74, 16 physical cores, no AVX-512 as a
ratio
ours / ORTin a shared process, at 1M/1-thread: ThresholdedRelu11.30 → 0.93, HardSigmoid 8.62 → 0.71, LeakyRelu 7.58 → 0.65, Selu 2.56 → 0.27.
These are peers, not corrections to the reviewer's numbers: different
microarchitecture, core count, and reference (one is
ours/ORT, the other is aninternal same-binary A/B). The AMD
ours/ORTratios and the i7 fast-path-vs-generic ratio are consistent — both attribute the cheap-op gains to removing the
generic path's allocations and passes.
The PR's own note that the three cheap ops remain ~1.2–1.5× ORT at 4k (a
per-call plugin-dispatch floor, not the kernel) stands and is a good separate
follow-up.
Correctness — falsifiers verified by the reviewer
The PR ships two load-bearing invariants, both confirmed RED-on-break:
minps/maxpsreturn their secondoperand for an unordered input, so the value must be second or a NaN silently
becomes
1then0.hard_sigmoid_special_valuespins NaN→NaN.x > 0sends-0down the exp branch; ORTreturns
+0, so both scalar and vector spell italpha*(exp(x)-1)(there isno AVX2
expm1).scalar_paths_agree_with_the_vector_pathspins scalar andvector to the same answer, and
selu_zero_sign_does_not_depend_on_dtypepinsit across dtypes.
Bit-identity where the op is exact (LeakyRelu, ThresholdedRelu asserted
bit-identical to their scalars over 20 003 points); HardSigmoid allows one FMA
rounding; Selu uses an alpha-scaled bound.
Verification (reviewer, Intel Core i7-13800H, Windows 11)
cargo test -p onnx-runtime-ep-cpu --lib— 1415 passed / 0 failed / 17ignored for the activation family (all four ops' special-value,
scalar-vs-vector bit-parity, dense-sweep and signed-zero tests green).
task_runtime::pool::tests::slot_exhaustion_declines_instead_of_blocking(from perf(cpu): add a CPU task runtime with an adaptive-spin native pool #1201's CPU task runtime, untouched by this PR) failed once under
contention and passes deterministically in isolation. Not introduced here.
cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings— clean.#[target_feature(enable = "avx2,fma")]confirmed intact onmap_psand
map_bias_ps.dispatch!and touch no MLASpath.