Skip to content

CPU EP: native Celu and Mish kernels, AVX2 Log (removes the last silent ORT activation deferrals) - #1235

Merged
justinchuby merged 3 commits into
mainfrom
squad/resch-celu-mish-log
Aug 18, 2026
Merged

justinchuby merged 3 commits into
mainfrom
squad/resch-celu-mish-log

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 18, 2026 •

Copy link
Copy Markdown
Owner

What this changes

Adds native CPU-EP kernels for Celu (opset 12) and Mish (opset 18), which had no kernel at all and were therefore silently handed to ORT's CPU EP, and vectorises Log (f32), which was still taking a scalar libm call per element. Celu/Mish are wired through all four capability filters (registry, PHASE1_OPS, dtype table, shape-inference table, plugin descriptors) and added to the two activation-coverage tests.

Behavioural claim — verified independently (reviewer)

The headline "removes the last silent ORT activation deferrals" is a behavioural claim, so I falsified it rather than trusting it, on Intel Core i7-13800H, Windows 11:

  • The deferral was real and silent. GetCapability runs three fail-closed filters; an op absent from the shape-inference table (Filter 2) or the dtype descriptor list (Filter 3) is handed to ORT with no diagnostic. Celu/Mish had no kernel, so they never reached supports_op.
  • Falsification (RED/GREEN). With the PR applied, activation_and_norm_ops_clear_every_capability_filter passes. Removing just the Celu/Mish entries from the shape-inference table (compute.rs, Filter 2) turns it RED with ::Celu: no shape rule (shape filter declines it) … would be handed to ORT's CPU EP for both ops; restoring returns it to GREEN. The test genuinely guards the deferral.
  • "The last" is defended structurally, not by spot-check. The coverage test's OWNED list now enumerates the full activation/normalisation family and holds every member to both the shape and dtype filters; the companion every_registered_op_has_a_shape_rule_or_is_a_known_gap enumerates all registered ops, and its remaining shape-table gap (data-dependent, internal-fusion, and inferrable-but-unwritten ops) contains no activations. Log was never a deferral — it already had a scalar kernel — so this PR is a perf change for it, not a coverage change.

Correctness

The kernels' own tests (dense sweeps vs the scalar reference, subnormals, and the full IEEE special set) pass; the scalar reference, the f64 path, and the AVX2 path are held to agree, including NaN propagation (f32::max/min drop NaN, so each kernel guards it to match ORT). Full cargo test -p onnx-runtime-ep-cpu --lib = 1362 passed, 0 failed, 17 ignored (base 1353 + 9 new). cargo clippy --all-targets -D warnings clean. map_ps's #[target_feature] (#1227) verified intact.

Performance — measured on this box, ORT-independent

The gaps doc's ratios were taken on an AMD EPYC 9V74 against ORT. I have no ORT reference on this host, so I measured the ORT-independent internal win instead: the AVX2 kernel vs the scalar-native fallback (which is what runs on non-AVX2 hosts and small tensors), on Intel Core i7-13800H (14C/20T), Windows 11, single thread, 1,048,576 elements, min-of-200 × 9 rounds.

kernel scalar-native (ms) AVX2 (ms) speedup
Log (positive domain) 2.89 1.14 ~2.54×
Celu (alpha=1, mixed) 2.77 0.68 ~4.07×
Mish (mixed) 17.82 2.50 ~7.13×

These are hardware-dependent (see docs/performance/CPU_ACTIVATION_GAPS.md §32 lesson) and are not the same quantity as the doc's ORT-relative ratios — mine compare AVX2 against the native scalar path on this machine; the AMD EPYC figures compare native against ORT. Presented as peers, each with its hardware named. Both agree on direction: the native kernels are a clear win.

Resch and others added 2 commits August 18, 2026 08:21
…ivation deferrals

`Celu` and `Mish` had no kernel in this EP at all. That is not a gap the
existing tests could report: the coverage list in
`no_activation_or_norm_op_is_left_to_ort` is written from the ops that
exist, so an op nobody implemented is absent from the kernel registry
*and* from the list that is supposed to catch its absence. The plugin's
fail-closed shape filter then dropped the claim and ORT's CPU EP ran
both -- exactly the deferral the architecture forbids.

`Log` was registered, but fell into `unary_math`'s scalar catch-all: one
`libm` call per element, which measured 1.76x ORT at 4k and 1.66x at 1M.

- `simd_activations`: AVX2 `log_ps` (Cephes/Eigen `plog` coefficients, so
  it evaluates what ORT evaluates), plus `celu_ps`, `softplus_ps` and
  `mish_ps`, with scalar references and slice entry points. Pure native --
  routed through `dispatch!`, never `dispatch_mlas!`.
- `unary_math`: `MathOp::Log` now takes the vector path.
- `activations`: `Celu { alpha }` and `Mish` variants and factories.
- `kernels/mod.rs`: capability arm, op-name list and registry entries.
- `ep-plugin/compute.rs`: shape rules for both, without which
  `GetCapability` fails closed and the claim is silently dropped.
- Tests: special values, denormals and dense sweeps against the scalar
  references, plus a new end-to-end parity test that runs all three on
  this EP with `disable_cpu_ep_fallback=1` and compares elementwise
  against a plain-ORT session.

Log, one thread, interleaved A/B (ratio = ours/ORT, lower is better):
4k 1.763 -> 0.782, 1M 1.651 -> 0.566.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…two docs

Opus review found one real defect and three smaller ones.

MUST-FIX -- `celu_scalar` destroyed NaN. Rust's `f32::max`/`min` are IEEE
`maxNum`/`minNum`, so they return the *other* operand for NaN: transcribing
ONNX's `max(0,x) + min(0, alpha*(exp(x/alpha)-1))` literally collapses both
terms to zero and answers `0.0`. The AVX2 path was deliberately written to
propagate NaN, and ORT propagates it too, so the result depended on whether
the slice was long enough for the vector path and on whether the host had
AVX2 -- which is the one thing `dispatch!` exists to prevent. Reachable by
any tensor under `SIMD_MIN_LEN`, by every `Float64` tensor (which never
touches SIMD), and by any non-AVX2 host. Guarded in `celu_scalar`, and
`Activation::apply` now delegates to it instead of keeping a second copy of
the formula; `apply_f64` carries the same guard.

Two new tests pin it, both verified to fail without the fix:
`celu_scalar_path_agrees_with_the_vector_path` runs the same specials through
a slice too short for SIMD and a padded one, and
`celu_and_mish_propagate_nan_on_every_path` covers f32 and f64 `apply`. The
existing special-value tests could not catch this: `special_inputs()` pads to
`SIMD_MIN_LEN`, so all of them take the vector path.

SHOULD-FIX -- `mod log_c` landed between `mod exp_c`'s doc comment and
`mod exp_c`, so `exp_c` lost its `#[cfg(target_arch = "x86_64")]` and its doc
described the wrong module. Same shape as the `map_ps` defect in #1227, in
the same file, found the same way. Each module now carries its own doc and
its own gate.

SHOULD-FIX -- added `Celu`/`Mish` to the `OWNED` list in
`activation_and_norm_ops_clear_every_capability_filter`, which checks the
dtype filter as well as the shape filter. Its comment said to add them the
moment a kernel landed; it has.

NIT -- the four new kernels had been inserted into the middle of
`exp_full_ps`'s doc comment, splitting it mid-sentence. Moved below it.

Also refreshed `CPU_ACTIVATION_GAPS.md`: "activations with no kernel at all"
is now empty, and `Celu`/`Mish`/`Log` moved to the closed table.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 18, 2026 08:59
@justinchuby
justinchuby enabled auto-merge (squash) August 18, 2026 08:59
`cargo clippy --all-targets` on the plugin test crate flagged
`manual_contains` and `neg_cmp_op_on_partial_ord` in the parity test
added by this PR. Both are trivial, but CI runs clippy with
`-D warnings`, so they are build failures rather than notes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown

⚠️ Benchmark Change Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
⚠️ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.57 ms 7.58 ms +15.4%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.14 ms 11.46 ms +13.0%
✅ qwen3_sampling_processors/top_k_partial_selection 142.79 µs 156.24 µs +9.4%
✅ qwen3_sampling_processors/top_k_top_p_fast 696.33 µs 761.06 µs +9.3%
✅ gather/large_f16_threads=1-internal/131072 12.34 µs 13.34 µs +8.1%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.65 ms 2.83 ms +6.7%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.66 ms 3.89 ms +6.1%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 398.11 µs 421.98 µs +6.0%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 539.00 µs 563.41 µs +4.5%
✅ tokenization/decode_tokens_per_second 7.02 ms 7.23 ms +2.9%
✅ matmul/small_generic_bf16_threads=8/1x256x256 48.58 µs 49.72 µs +2.3%
✅ grammar_masking/llguidance_compute_mask/32 72.53 µs 74.18 µs +2.3%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 570.48 µs 580.60 µs +1.8%
✅ tokenization/encode_tokens_per_second 390.62 µs 395.11 µs +1.2%
✅ sampling_latency/greedy_per_token 3.32 µs 3.36 µs +1.1%
✅ matmul/small_generic_f32_threads=8/1x256x256 49.23 µs 48.82 µs -0.8%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.40 ms 2.38 ms -1.0%
✅ gather/medium_bf16_threads=1-internal/32768 2.61 µs 2.58 µs -1.0%
✅ logit_processing/seven_processor_chain_per_step 325.40 µs 316.26 µs -2.8%
✅ gather/small_bf16_threads=1-internal/4096 506.6 ns 490.4 ns -3.2%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 676.86 µs 648.93 µs -4.1%
✅ add/large_f32_threads=1-internal/4194304 983.45 µs 941.61 µs -4.3%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 180.17 µs 172.23 µs -4.4%
✅ matmul/medium_generic_f16_threads=1/32x512x512 37.19 µs 35.19 µs -5.4%
✅ matmul/medium_generic_f16_threads=8/32x512x512 37.71 µs 35.64 µs -5.5%
✅ sampling_latency/top_k_per_token 57.36 µs 53.61 µs -6.5%
✅ matmul/small_generic_f16_threads=8/1x256x256 46.29 µs 42.64 µs -7.9%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 4.73 ms 4.29 ms -9.2%
✅ matmul/small_generic_f32_threads=1/1x256x256 44.44 µs 39.83 µs -10.4%
✅ gather/medium_f16_threads=1-internal/32768 2.70 µs 2.42 µs -10.5%
✅ kv_cache/alloc_dealloc_pages 42.25 µs 37.61 µs -11.0%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.51 ms 2.23 ms -11.4%
✅ add/large_f16_threads=1-internal/4194304 2.85 ms 2.52 ms -11.4%
✅ add/small_f32_threads=1-internal/1024 222.3 ns 196.9 ns -11.4%
✅ sampling_latency/min_p_per_token 252.32 µs 221.77 µs -12.1%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.35 ms 1.17 ms -13.3%
🟢 add/medium_bf16_threads=1-internal/262144 156.37 µs 132.10 µs -15.5%
🟢 gather/small_f16_threads=1-internal/4096 620.4 ns 518.1 ns -16.5%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 317.96 µs 261.43 µs -17.8%
🟢 sampling_latency/top_p_per_token 481.08 µs 391.56 µs -18.6%
🟢 matmul/large_generic_f16_threads=1/32x1024x1024 111.60 µs 90.16 µs -19.2%
🟢 gather/large_f32_threads=1-internal/131072 32.88 µs 26.18 µs -20.4%
🟢 reduce_mean/small_f32_threads=1-internal/4096 20.53 µs 16.09 µs -21.6%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.34 ms 1.82 ms -22.1%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.29 ms 986.77 µs -23.3%
🟢 gather/medium_f32_threads=1-internal/32768 5.48 µs 4.11 µs -24.9%
🟢 matmul/large_generic_f16_threads=8/32x1024x1024 120.32 µs 89.59 µs -25.5%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 55.38 µs 40.39 µs -27.1%
🟢 add/small_f16_threads=1-internal/1024 616.0 ns 440.8 ns -28.4%
🟢 matmul/small_generic_f16_threads=1/1x256x256 43.72 µs 31.10 µs -28.9%
🟢 add/large_bf16_threads=1-internal/4194304 2.51 ms 1.76 ms -30.0%
🟢 gather/small_f32_threads=1-internal/4096 948.3 ns 658.5 ns -30.6%
🟢 add/small_bf16_threads=1-internal/1024 634.5 ns 435.0 ns -31.4%
🟢 add/medium_f16_threads=1-internal/262144 183.58 µs 117.48 µs -36.0%
🟢 add/medium_f32_threads=1-internal/262144 38.04 µs 23.61 µs -37.9%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 97.04 µs 58.95 µs -39.3%
🟢 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 134.31 µs 81.50 µs -39.3%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.28 ms 749.62 µs -41.4%
🟢 gather/large_bf16_threads=1-internal/131072 17.15 µs 10.04 µs -41.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.53 3.75 6.77 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

Reviewer verification — reproduced, merging

Validated on Intel Core i7-13800H (14C/20T), Windows 11, off current main (1353-test baseline confirmed first).

Behavioural claim falsified and confirmed. The "silent ORT deferral" is real: an activation absent from the shape-inference table or the dtype descriptor list is handed to ORT with no diagnostic. With the PR, activation_and_norm_ops_clear_every_capability_filter passes; deleting just the Celu/Mish shape-inference entries turns it RED (::Celu: no shape rule … would be handed to ORT's CPU EP), restoring returns GREEN. "The last" holds structurally — the coverage test enumerates the whole activation/norm family, and the registered-op shape-gap list contains no activations. Log was already kernelled, so it is a perf change, not a deferral fix.

Correctness: kernels' dense sweeps, subnormals, and IEEE specials pass; scalar/f64/AVX2 paths agree incl. NaN propagation. cargo test --lib = 1362 passed / 0 failed / 17 ignored (1353 + 9); clippy -D warnings clean; #1227's map_ps target_feature intact.

Performance (ORT-independent, this box): AVX2 vs native-scalar, 1M elems, single-thread, min-of-200 × 9 rounds — Log 2.89→1.14 ms (~2.54×), Celu 2.77→0.68 ms (~4.07×), Mish 17.82→2.50 ms (~7.13×). The doc's AMD-EPYC vs-ORT ratios are a peer, not reproduced here (no ORT reference on this host).

Merging by squash.

@justinchuby
justinchuby merged commit e0d61a1 into main Aug 18, 2026
6 of 17 checks passed
@justinchuby
justinchuby deleted the squad/resch-celu-mish-log branch August 18, 2026 15:56
justinchuby added a commit that referenced this pull request Aug 18, 2026
…1240)

# Zero-copy activation fast path, and an AVX2 Elu kernel

Two changes in the same seam of `ActivationKernel::execute`:

1. **Generalise the contiguous-f32 fast path.** `silu_contiguous_f32`
becomes
`contiguous_f32(input, output, kernel)` — the kernel is now a parameter
— so
Celu, Mish and Elu write their result straight into the output buffer
instead
   of taking the general path's two allocations and three passes
   (`to_dense_f32_widen` → `map` → `write_dense_f32_narrow`).
2. **An AVX2 Elu kernel** (`elu_ps` / `elu_f32_slice` / `elu_avx2`)
built on the
`exp_full_ps` that Celu and Mish already use, replacing a scalar `expf`
loop.

Base: `origin/main` with #1235 (Celu/Mish/Log) landed first.

---

## Reviewer measurements (Intel Core i7-13800H)

All numbers below were taken by the reviewer on an **Intel Core
i7-13800H
(14C/20T, AVX2+FMA, no AVX-512), Windows 11**. No ORT reference build is
available on this box, so the perf figure is an **ORT-independent,
same-binary
A/B**: the scalar Elu reference (`elu_scalar` loop) versus the AVX2
slice kernel
(`elu_f32_slice`), 1M f32 elements, single-thread, min-of-200 iterations
× 9
rounds, `black_box`ed on both arms of the same release binary.

| Elu, 1M elems, 1 thread | scalar-native | AVX2 slice | speedup |
|---|---|---|---|
| min-of-200 × 9 rounds | 2.5842 ms | 0.5107 ms | **~5.06×** |

This measures the vectorisation win itself (scalar vs AVX2) rather than
a ratio
against ORT, which is what the same-binary A/B can establish without an
ORT
build present.

### Peer measurement (original PR, AMD EPYC 9V74)

The PR was originally measured on an **AMD EPYC 9V74, 16 physical cores,
AVX2/FMA/F16C, no AVX-512**, as a ratio `ours / ORT` in a shared
process. There,
Elu went from **2.06× slower than ORT to ~4.2× faster** at 1M (4.96 ms →
0.57 ms
against ORT's 2.39 ms), with Celu/Mish also improving from the fast path
alone
(their kernels are unchanged in this PR).

These two results are **peers, not corrections**: different
microarchitecture,
different core count, and different reference (one is `ours/ORT`, the
other is an
internal scalar-vs-AVX2 A/B). Both show the AVX2 Elu kernel is a
substantial
win; neither refutes the other.

---

## The fast-path fallback is proven, not assumed

`contiguous_f32` forms `&[f32]` and `&mut [f32]` over raw tensor
pointers, which
is only sound when both tensors are dense f32, agree on shape and
strides, and do
not overlap in memory. A fast path that silently produced wrong results
on an
unusual layout would be worse than a slow one, so the reviewer added a
direct
test — `contiguous_f32_fast_path_declines_every_unsafe_layout` — that
pins the
guard:

- **dtype mismatch** (f64 input) → declines;
- **non-dense strides** (a broadcast stride of 0) → declines;
- **aliasing** (input and output over the same buffer) → declines;
- **control**: two distinct dense f32 buffers of equal shape → the fast
path
  runs *and* is the thing that produced the output.

The closure passed on the decline arms panics if invoked, so the test
fails if
the fast path is ever wrongly taken. Verified by falsification:
neutering the
overlap check makes the aliasing arm run the kernel and the test goes
**RED**
(`fast path ran on a layout it must decline`); restoring it returns
**GREEN**.
The overlap check refuses even exact in-place aliasing, because forming
`&[f32]` and `&mut [f32]` over the same bytes is UB in Rust regardless
of what
the elementwise machine code would compute.

---

## Numerics

- `elu_special_values` — `-Inf → -alpha`, `-0 → -0` (bit-checked), `+0 →
+0`,
`+Inf`, `NaN`, `±MAX`, `±1e30`, and ordinary values against the scalar.
- `elu_scalar_path_agrees_with_the_vector_path` — nine specials run once
through
a sub-`SIMD_MIN_LEN` slice (scalar) and once padded past it (vector),
compared
  **bit for bit**, at alpha ∈ {0.25, 1, 2, 7.5}.
- `elu_matches_scalar_across_alphas` — 20 003 points over `[-90, 90]` at
four
  alphas with an alpha-scaled bound.

`-0.0`: Elu is a select so the sign survives (`Elu(-0) = -0`), where
Celu's
trailing addition gives `Celu(-0) = +0`; both are pinned. The scalar
reference
moves from `alpha * x.exp_m1()` to `alpha * (x.exp() - 1.0)` so the
scalar path,
the vector path and ORT are the same function (there is no AVX2
`expm1`); the
two spellings differ by at most ~6e-8 absolute, inside this module's
4e-7 bound.

---

## Verification (reviewer, Intel Core i7-13800H, Windows 11)

- `cargo test -p onnx-runtime-ep-cpu --lib` — **1366 passed / 0 failed /
17
ignored** (+4 over #1235: three Elu tests plus the fallback falsifier).
- `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` —
clean.
- #1227's `#[target_feature(enable = "avx2,fma")]` confirmed intact on
**both**
  `map_ps` and `map_bias_ps` after the merge.
- MLAS remains non-default; `elu_f32_slice` uses `dispatch!` like its
neighbours
  and touches no MLAS path.

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 18, 2026
…u and Selu (#1243)

# AVX2 kernels for LeakyRelu, HardSigmoid, ThresholdedRelu and Selu

`LeakyRelu`, `HardSigmoid`, `ThresholdedRelu` and `Selu` had no slice
kernel, so
they took `ActivationKernel`'s generic path: widen into a `Vec`, `map()`
a
per-element `match`, collect into a second `Vec`, copy that into the
output
tensor. This PR gives each a slice kernel and routes them through the
zero-copy
`contiguous_f32` fast path (established in #1240), and adds a genuinely
vectorised exponential for Selu.

Stacked on #1240 (which is stacked on #1235).

---

## Reviewer measurements (Intel Core i7-13800H) — where the win actually
comes from

Measured by the reviewer on an **Intel Core i7-13800H (14C/20T,
AVX2+FMA, no
AVX-512), Windows 11**. No ORT build is available on this box, so these
are
ORT-independent, same-binary A/Bs, 1M f32 elements, single-thread,
min-of-200 ×
9 rounds.

**Two different comparisons, because they answer different questions:**

1. **Fast-path routing vs the generic path** (what this PR changes at
the call
site). Emulating the old generic path (`to_vec` → `map` →
`copy_from_slice`)
   against the new `*_f32_slice` fast path, for LeakyRelu:

   | LeakyRelu, 1M, 1 thread | generic path | fast path | speedup |
   |---|---|---|---|
   | min-of-200 × 9 | 1.8784 ms | 0.1712 ms | **~10.97×** |

This — eliminating two allocations and two extra passes — is the
dominant win
   for the three cheap ops.

2. **The AVX2 kernel vs its scalar reference** (isolating the
hand-written SIMD
   arithmetic alone, both in the same release binary):

   | op, 1M, 1 thread | scalar | AVX2 | speedup |
   |---|---|---|---|
   | LeakyRelu | 0.1533 ms | 0.1524 ms | ~1.01× |
   | HardSigmoid | 0.1599 ms | 0.1583 ms | ~1.01× |
   | ThresholdedRelu | 0.1644 ms | 0.1619 ms | ~1.02× |
   | **Selu** | 1.4301 ms | 0.3391 ms | **~4.22×** |

For the three cheap ops the hand-written AVX2 kernel is **no faster than
the
scalar loop**, because LLVM already auto-vectorises a compare-and-blend
in
release mode. Selu is the exception: its exponential does not
auto-vectorise,
   so the vector kernel is a real ~4.2× on the arithmetic itself.

**Honest attribution:** for LeakyRelu/HardSigmoid/ThresholdedRelu the
speedup is
the fast-path plumbing, not the SIMD; for Selu it is both. That is worth
stating
plainly so nobody credits the compare-and-blend intrinsics with a win
the
allocator elimination actually produced.

### Peer measurement (original PR, AMD EPYC 9V74)

Originally measured on an **AMD EPYC 9V74, 16 physical cores, no
AVX-512** as a
ratio `ours / ORT` in a shared process, at 1M/1-thread: ThresholdedRelu
11.30 → 0.93, HardSigmoid 8.62 → 0.71, LeakyRelu 7.58 → 0.65, Selu 2.56
→ 0.27.
These are **peers, not corrections** to the reviewer's numbers:
different
microarchitecture, core count, and reference (one is `ours/ORT`, the
other is an
internal same-binary A/B). The AMD `ours/ORT` ratios and the i7
fast-path-vs-
generic ratio are consistent — both attribute the cheap-op gains to
removing the
generic path's allocations and passes.

The PR's own note that the three cheap ops remain **~1.2–1.5× ORT at
4k** (a
per-call plugin-dispatch floor, not the kernel) stands and is a good
separate
follow-up.

---

## Correctness — falsifiers verified by the reviewer

The PR ships two load-bearing invariants, both confirmed RED-on-break:

- **HardSigmoid clamp operand order.** `minps`/`maxps` return their
*second*
operand for an unordered input, so the value must be second or a NaN
silently
  becomes `1` then `0`. `hard_sigmoid_special_values` pins NaN→NaN.
- **Selu signed-zero.** ONNX's `x > 0` sends `-0` down the exp branch;
ORT
returns `+0`, so both scalar and vector spell it `alpha*(exp(x)-1)`
(there is
no AVX2 `expm1`). `scalar_paths_agree_with_the_vector_paths` pins scalar
and
vector to the same answer, and `selu_zero_sign_does_not_depend_on_dtype`
pins
  it across dtypes.

Bit-identity where the op is exact (LeakyRelu, ThresholdedRelu asserted
bit-identical to their scalars over 20 003 points); HardSigmoid allows
one FMA
rounding; Selu uses an alpha-scaled bound.

---

## Verification (reviewer, Intel Core i7-13800H, Windows 11)

- `cargo test -p onnx-runtime-ep-cpu --lib` — **1415 passed / 0 failed /
17
  ignored** for the activation family (all four ops' special-value,
  scalar-vs-vector bit-parity, dense-sweep and signed-zero tests green).
- One unrelated flake observed under full-suite load:

`task_runtime::pool::tests::slot_exhaustion_declines_instead_of_blocking`
(from #1201's CPU task runtime, untouched by this PR) failed once under
contention and passes deterministically in isolation. Not introduced
here.
- `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` —
clean.
- #1227's `#[target_feature(enable = "avx2,fma")]` confirmed intact on
`map_ps`
  and `map_bias_ps`.
- MLAS remains non-default; all four kernels use `dispatch!` and touch
no MLAS
  path.

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant