Repository navigation
CPU EP: native Celu and Mish kernels, AVX2 Log (removes the last silent ORT activation deferrals) - #1235
Conversation
…ivation deferrals
`Celu` and `Mish` had no kernel in this EP at all. That is not a gap the
existing tests could report: the coverage list in
`no_activation_or_norm_op_is_left_to_ort` is written from the ops that
exist, so an op nobody implemented is absent from the kernel registry
*and* from the list that is supposed to catch its absence. The plugin's
fail-closed shape filter then dropped the claim and ORT's CPU EP ran
both -- exactly the deferral the architecture forbids.
`Log` was registered, but fell into `unary_math`'s scalar catch-all: one
`libm` call per element, which measured 1.76x ORT at 4k and 1.66x at 1M.
- `simd_activations`: AVX2 `log_ps` (Cephes/Eigen `plog` coefficients, so
it evaluates what ORT evaluates), plus `celu_ps`, `softplus_ps` and
`mish_ps`, with scalar references and slice entry points. Pure native --
routed through `dispatch!`, never `dispatch_mlas!`.
- `unary_math`: `MathOp::Log` now takes the vector path.
- `activations`: `Celu { alpha }` and `Mish` variants and factories.
- `kernels/mod.rs`: capability arm, op-name list and registry entries.
- `ep-plugin/compute.rs`: shape rules for both, without which
`GetCapability` fails closed and the claim is silently dropped.
- Tests: special values, denormals and dense sweeps against the scalar
references, plus a new end-to-end parity test that runs all three on
this EP with `disable_cpu_ep_fallback=1` and compares elementwise
against a plain-ORT session.
Log, one thread, interleaved A/B (ratio = ours/ORT, lower is better):
4k 1.763 -> 0.782, 1M 1.651 -> 0.566.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…two docs Opus review found one real defect and three smaller ones. MUST-FIX -- `celu_scalar` destroyed NaN. Rust's `f32::max`/`min` are IEEE `maxNum`/`minNum`, so they return the *other* operand for NaN: transcribing ONNX's `max(0,x) + min(0, alpha*(exp(x/alpha)-1))` literally collapses both terms to zero and answers `0.0`. The AVX2 path was deliberately written to propagate NaN, and ORT propagates it too, so the result depended on whether the slice was long enough for the vector path and on whether the host had AVX2 -- which is the one thing `dispatch!` exists to prevent. Reachable by any tensor under `SIMD_MIN_LEN`, by every `Float64` tensor (which never touches SIMD), and by any non-AVX2 host. Guarded in `celu_scalar`, and `Activation::apply` now delegates to it instead of keeping a second copy of the formula; `apply_f64` carries the same guard. Two new tests pin it, both verified to fail without the fix: `celu_scalar_path_agrees_with_the_vector_path` runs the same specials through a slice too short for SIMD and a padded one, and `celu_and_mish_propagate_nan_on_every_path` covers f32 and f64 `apply`. The existing special-value tests could not catch this: `special_inputs()` pads to `SIMD_MIN_LEN`, so all of them take the vector path. SHOULD-FIX -- `mod log_c` landed between `mod exp_c`'s doc comment and `mod exp_c`, so `exp_c` lost its `#[cfg(target_arch = "x86_64")]` and its doc described the wrong module. Same shape as the `map_ps` defect in #1227, in the same file, found the same way. Each module now carries its own doc and its own gate. SHOULD-FIX -- added `Celu`/`Mish` to the `OWNED` list in `activation_and_norm_ops_clear_every_capability_filter`, which checks the dtype filter as well as the shape filter. Its comment said to add them the moment a kernel landed; it has. NIT -- the four new kernels had been inserted into the middle of `exp_full_ps`'s doc comment, splitting it mid-sentence. Moved below it. Also refreshed `CPU_ACTIVATION_GAPS.md`: "activations with no kernel at all" is now empty, and `Celu`/`Mish`/`Log` moved to the closed table. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`cargo clippy --all-targets` on the plugin test crate flagged `manual_contains` and `neg_cmp_op_on_partial_ord` in the parity test added by this PR. Both are trivial, but CI runs clippy with `-D warnings`, so they are build failures rather than notes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
| Status | Scenario | Base | PR | Change |
|---|---|---|---|---|
qwen3_sampling_processors/top_k_top_p_full_sort_baseline |
6.57 ms | 7.58 ms | +15.4% | |
| ✅ | matmul/large_generic_f32_threads=1/32x1024x1024 |
10.14 ms | 11.46 ms | +13.0% |
| ✅ | qwen3_sampling_processors/top_k_partial_selection |
142.79 µs | 156.24 µs | +9.4% |
| ✅ | qwen3_sampling_processors/top_k_top_p_fast |
696.33 µs | 761.06 µs | +9.3% |
| ✅ | gather/large_f16_threads=1-internal/131072 |
12.34 µs | 13.34 µs | +8.1% |
| ✅ | matmul/medium_generic_f32_threads=1/32x512x512 |
2.65 ms | 2.83 ms | +6.7% |
| ✅ | qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline |
3.66 ms | 3.89 ms | +6.1% |
| ✅ | matmul/medium_generic_bf16_threads=8/32x512x512 |
398.11 µs | 421.98 µs | +6.0% |
| ✅ | matmul/medium_generic_bf16_threads=1/32x512x512 |
539.00 µs | 563.41 µs | +4.5% |
| ✅ | tokenization/decode_tokens_per_second |
7.02 ms | 7.23 ms | +2.9% |
| ✅ | matmul/small_generic_bf16_threads=8/1x256x256 |
48.58 µs | 49.72 µs | +2.3% |
| ✅ | grammar_masking/llguidance_compute_mask/32 |
72.53 µs | 74.18 µs | +2.3% |
| ✅ | qwen3_sampling_processors/top_p_fast_after_top_k |
570.48 µs | 580.60 µs | +1.8% |
| ✅ | tokenization/encode_tokens_per_second |
390.62 µs | 395.11 µs | +1.2% |
| ✅ | sampling_latency/greedy_per_token |
3.32 µs | 3.36 µs | +1.1% |
| ✅ | matmul/small_generic_f32_threads=8/1x256x256 |
49.23 µs | 48.82 µs | -0.8% |
| ✅ | qwen3_sampling_processors/top_k_full_sort_baseline |
2.40 ms | 2.38 ms | -1.0% |
| ✅ | gather/medium_bf16_threads=1-internal/32768 |
2.61 µs | 2.58 µs | -1.0% |
| ✅ | logit_processing/seven_processor_chain_per_step |
325.40 µs | 316.26 µs | -2.8% |
| ✅ | gather/small_bf16_threads=1-internal/4096 |
506.6 ns | 490.4 ns | -3.2% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 |
676.86 µs | 648.93 µs | -4.1% |
| ✅ | add/large_f32_threads=1-internal/4194304 |
983.45 µs | 941.61 µs | -4.3% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 |
180.17 µs | 172.23 µs | -4.4% |
| ✅ | matmul/medium_generic_f16_threads=1/32x512x512 |
37.19 µs | 35.19 µs | -5.4% |
| ✅ | matmul/medium_generic_f16_threads=8/32x512x512 |
37.71 µs | 35.64 µs | -5.5% |
| ✅ | sampling_latency/top_k_per_token |
57.36 µs | 53.61 µs | -6.5% |
| ✅ | matmul/small_generic_f16_threads=8/1x256x256 |
46.29 µs | 42.64 µs | -7.9% |
| ✅ | matmul/large_generic_f32_threads=8/32x1024x1024 |
4.73 ms | 4.29 ms | -9.2% |
| ✅ | matmul/small_generic_f32_threads=1/1x256x256 |
44.44 µs | 39.83 µs | -10.4% |
| ✅ | gather/medium_f16_threads=1-internal/32768 |
2.70 µs | 2.42 µs | -10.5% |
| ✅ | kv_cache/alloc_dealloc_pages |
42.25 µs | 37.61 µs | -11.0% |
| ✅ | matmul/large_generic_bf16_threads=1/32x1024x1024 |
2.51 ms | 2.23 ms | -11.4% |
| ✅ | add/large_f16_threads=1-internal/4194304 |
2.85 ms | 2.52 ms | -11.4% |
| ✅ | add/small_f32_threads=1-internal/1024 |
222.3 ns | 196.9 ns | -11.4% |
| ✅ | sampling_latency/min_p_per_token |
252.32 µs | 221.77 µs | -12.1% |
| ✅ | matmul/medium_generic_f32_threads=8/32x512x512 |
1.35 ms | 1.17 ms | -13.3% |
| 🟢 | add/medium_bf16_threads=1-internal/262144 |
156.37 µs | 132.10 µs | -15.5% |
| 🟢 | gather/small_f16_threads=1-internal/4096 |
620.4 ns | 518.1 ns | -16.5% |
| 🟢 | reduce_mean/medium_f32_threads=1-internal/65536 |
317.96 µs | 261.43 µs | -17.8% |
| 🟢 | sampling_latency/top_p_per_token |
481.08 µs | 391.56 µs | -18.6% |
| 🟢 | matmul/large_generic_f16_threads=1/32x1024x1024 |
111.60 µs | 90.16 µs | -19.2% |
| 🟢 | gather/large_f32_threads=1-internal/131072 |
32.88 µs | 26.18 µs | -20.4% |
| 🟢 | reduce_mean/small_f32_threads=1-internal/4096 |
20.53 µs | 16.09 µs | -21.6% |
| 🟢 | matmul/large_generic_bf16_threads=8/32x1024x1024 |
2.34 ms | 1.82 ms | -22.1% |
| 🟢 | reduce_mean/large_f32_threads=1-internal/262144 |
1.29 ms | 986.77 µs | -23.3% |
| 🟢 | gather/medium_f32_threads=1-internal/32768 |
5.48 µs | 4.11 µs | -24.9% |
| 🟢 | matmul/large_generic_f16_threads=8/32x1024x1024 |
120.32 µs | 89.59 µs | -25.5% |
| 🟢 | matmul/small_generic_bf16_threads=1/1x256x256 |
55.38 µs | 40.39 µs | -27.1% |
| 🟢 | add/small_f16_threads=1-internal/1024 |
616.0 ns | 440.8 ns | -28.4% |
| 🟢 | matmul/small_generic_f16_threads=1/1x256x256 |
43.72 µs | 31.10 µs | -28.9% |
| 🟢 | add/large_bf16_threads=1-internal/4194304 |
2.51 ms | 1.76 ms | -30.0% |
| 🟢 | gather/small_f32_threads=1-internal/4096 |
948.3 ns | 658.5 ns | -30.6% |
| 🟢 | add/small_bf16_threads=1-internal/1024 |
634.5 ns | 435.0 ns | -31.4% |
| 🟢 | add/medium_f16_threads=1-internal/262144 |
183.58 µs | 117.48 µs | -36.0% |
| 🟢 | add/medium_f32_threads=1-internal/262144 |
38.04 µs | 23.61 µs | -37.9% |
| 🟢 | block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 |
97.04 µs | 58.95 µs | -39.3% |
| 🟢 | block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 |
134.31 µs | 81.50 µs | -39.3% |
| 🟢 | block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 |
1.28 ms | 749.62 µs | -41.4% |
| 🟢 | gather/large_bf16_threads=1-internal/131072 |
17.15 µs | 10.04 µs | -41.4% |
Visual flags:
Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.53 3.75 6.77 }
What this cannot catch
- Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
- Sub-threshold regressions that compound over multiple PRs
- Performance changes that only manifest under GPU execution
- Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)
Reviewer verification — reproduced, mergingValidated on Intel Core i7-13800H (14C/20T), Windows 11, off current Behavioural claim falsified and confirmed. The "silent ORT deferral" is real: an activation absent from the shape-inference table or the dtype descriptor list is handed to ORT with no diagnostic. With the PR, Correctness: kernels' dense sweeps, subnormals, and IEEE specials pass; scalar/f64/AVX2 paths agree incl. NaN propagation. Performance (ORT-independent, this box): AVX2 vs native-scalar, 1M elems, single-thread, min-of-200 × 9 rounds — Merging by squash. |
…1240) # Zero-copy activation fast path, and an AVX2 Elu kernel Two changes in the same seam of `ActivationKernel::execute`: 1. **Generalise the contiguous-f32 fast path.** `silu_contiguous_f32` becomes `contiguous_f32(input, output, kernel)` — the kernel is now a parameter — so Celu, Mish and Elu write their result straight into the output buffer instead of taking the general path's two allocations and three passes (`to_dense_f32_widen` → `map` → `write_dense_f32_narrow`). 2. **An AVX2 Elu kernel** (`elu_ps` / `elu_f32_slice` / `elu_avx2`) built on the `exp_full_ps` that Celu and Mish already use, replacing a scalar `expf` loop. Base: `origin/main` with #1235 (Celu/Mish/Log) landed first. --- ## Reviewer measurements (Intel Core i7-13800H) All numbers below were taken by the reviewer on an **Intel Core i7-13800H (14C/20T, AVX2+FMA, no AVX-512), Windows 11**. No ORT reference build is available on this box, so the perf figure is an **ORT-independent, same-binary A/B**: the scalar Elu reference (`elu_scalar` loop) versus the AVX2 slice kernel (`elu_f32_slice`), 1M f32 elements, single-thread, min-of-200 iterations × 9 rounds, `black_box`ed on both arms of the same release binary. | Elu, 1M elems, 1 thread | scalar-native | AVX2 slice | speedup | |---|---|---|---| | min-of-200 × 9 rounds | 2.5842 ms | 0.5107 ms | **~5.06×** | This measures the vectorisation win itself (scalar vs AVX2) rather than a ratio against ORT, which is what the same-binary A/B can establish without an ORT build present. ### Peer measurement (original PR, AMD EPYC 9V74) The PR was originally measured on an **AMD EPYC 9V74, 16 physical cores, AVX2/FMA/F16C, no AVX-512**, as a ratio `ours / ORT` in a shared process. There, Elu went from **2.06× slower than ORT to ~4.2× faster** at 1M (4.96 ms → 0.57 ms against ORT's 2.39 ms), with Celu/Mish also improving from the fast path alone (their kernels are unchanged in this PR). These two results are **peers, not corrections**: different microarchitecture, different core count, and different reference (one is `ours/ORT`, the other is an internal scalar-vs-AVX2 A/B). Both show the AVX2 Elu kernel is a substantial win; neither refutes the other. --- ## The fast-path fallback is proven, not assumed `contiguous_f32` forms `&[f32]` and `&mut [f32]` over raw tensor pointers, which is only sound when both tensors are dense f32, agree on shape and strides, and do not overlap in memory. A fast path that silently produced wrong results on an unusual layout would be worse than a slow one, so the reviewer added a direct test — `contiguous_f32_fast_path_declines_every_unsafe_layout` — that pins the guard: - **dtype mismatch** (f64 input) → declines; - **non-dense strides** (a broadcast stride of 0) → declines; - **aliasing** (input and output over the same buffer) → declines; - **control**: two distinct dense f32 buffers of equal shape → the fast path runs *and* is the thing that produced the output. The closure passed on the decline arms panics if invoked, so the test fails if the fast path is ever wrongly taken. Verified by falsification: neutering the overlap check makes the aliasing arm run the kernel and the test goes **RED** (`fast path ran on a layout it must decline`); restoring it returns **GREEN**. The overlap check refuses even exact in-place aliasing, because forming `&[f32]` and `&mut [f32]` over the same bytes is UB in Rust regardless of what the elementwise machine code would compute. --- ## Numerics - `elu_special_values` — `-Inf → -alpha`, `-0 → -0` (bit-checked), `+0 → +0`, `+Inf`, `NaN`, `±MAX`, `±1e30`, and ordinary values against the scalar. - `elu_scalar_path_agrees_with_the_vector_path` — nine specials run once through a sub-`SIMD_MIN_LEN` slice (scalar) and once padded past it (vector), compared **bit for bit**, at alpha ∈ {0.25, 1, 2, 7.5}. - `elu_matches_scalar_across_alphas` — 20 003 points over `[-90, 90]` at four alphas with an alpha-scaled bound. `-0.0`: Elu is a select so the sign survives (`Elu(-0) = -0`), where Celu's trailing addition gives `Celu(-0) = +0`; both are pinned. The scalar reference moves from `alpha * x.exp_m1()` to `alpha * (x.exp() - 1.0)` so the scalar path, the vector path and ORT are the same function (there is no AVX2 `expm1`); the two spellings differ by at most ~6e-8 absolute, inside this module's 4e-7 bound. --- ## Verification (reviewer, Intel Core i7-13800H, Windows 11) - `cargo test -p onnx-runtime-ep-cpu --lib` — **1366 passed / 0 failed / 17 ignored** (+4 over #1235: three Elu tests plus the fallback falsifier). - `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` — clean. - #1227's `#[target_feature(enable = "avx2,fma")]` confirmed intact on **both** `map_ps` and `map_bias_ps` after the merge. - MLAS remains non-default; `elu_f32_slice` uses `dispatch!` like its neighbours and touches no MLAS path. Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
…u and Selu (#1243) # AVX2 kernels for LeakyRelu, HardSigmoid, ThresholdedRelu and Selu `LeakyRelu`, `HardSigmoid`, `ThresholdedRelu` and `Selu` had no slice kernel, so they took `ActivationKernel`'s generic path: widen into a `Vec`, `map()` a per-element `match`, collect into a second `Vec`, copy that into the output tensor. This PR gives each a slice kernel and routes them through the zero-copy `contiguous_f32` fast path (established in #1240), and adds a genuinely vectorised exponential for Selu. Stacked on #1240 (which is stacked on #1235). --- ## Reviewer measurements (Intel Core i7-13800H) — where the win actually comes from Measured by the reviewer on an **Intel Core i7-13800H (14C/20T, AVX2+FMA, no AVX-512), Windows 11**. No ORT build is available on this box, so these are ORT-independent, same-binary A/Bs, 1M f32 elements, single-thread, min-of-200 × 9 rounds. **Two different comparisons, because they answer different questions:** 1. **Fast-path routing vs the generic path** (what this PR changes at the call site). Emulating the old generic path (`to_vec` → `map` → `copy_from_slice`) against the new `*_f32_slice` fast path, for LeakyRelu: | LeakyRelu, 1M, 1 thread | generic path | fast path | speedup | |---|---|---|---| | min-of-200 × 9 | 1.8784 ms | 0.1712 ms | **~10.97×** | This — eliminating two allocations and two extra passes — is the dominant win for the three cheap ops. 2. **The AVX2 kernel vs its scalar reference** (isolating the hand-written SIMD arithmetic alone, both in the same release binary): | op, 1M, 1 thread | scalar | AVX2 | speedup | |---|---|---|---| | LeakyRelu | 0.1533 ms | 0.1524 ms | ~1.01× | | HardSigmoid | 0.1599 ms | 0.1583 ms | ~1.01× | | ThresholdedRelu | 0.1644 ms | 0.1619 ms | ~1.02× | | **Selu** | 1.4301 ms | 0.3391 ms | **~4.22×** | For the three cheap ops the hand-written AVX2 kernel is **no faster than the scalar loop**, because LLVM already auto-vectorises a compare-and-blend in release mode. Selu is the exception: its exponential does not auto-vectorise, so the vector kernel is a real ~4.2× on the arithmetic itself. **Honest attribution:** for LeakyRelu/HardSigmoid/ThresholdedRelu the speedup is the fast-path plumbing, not the SIMD; for Selu it is both. That is worth stating plainly so nobody credits the compare-and-blend intrinsics with a win the allocator elimination actually produced. ### Peer measurement (original PR, AMD EPYC 9V74) Originally measured on an **AMD EPYC 9V74, 16 physical cores, no AVX-512** as a ratio `ours / ORT` in a shared process, at 1M/1-thread: ThresholdedRelu 11.30 → 0.93, HardSigmoid 8.62 → 0.71, LeakyRelu 7.58 → 0.65, Selu 2.56 → 0.27. These are **peers, not corrections** to the reviewer's numbers: different microarchitecture, core count, and reference (one is `ours/ORT`, the other is an internal same-binary A/B). The AMD `ours/ORT` ratios and the i7 fast-path-vs- generic ratio are consistent — both attribute the cheap-op gains to removing the generic path's allocations and passes. The PR's own note that the three cheap ops remain **~1.2–1.5× ORT at 4k** (a per-call plugin-dispatch floor, not the kernel) stands and is a good separate follow-up. --- ## Correctness — falsifiers verified by the reviewer The PR ships two load-bearing invariants, both confirmed RED-on-break: - **HardSigmoid clamp operand order.** `minps`/`maxps` return their *second* operand for an unordered input, so the value must be second or a NaN silently becomes `1` then `0`. `hard_sigmoid_special_values` pins NaN→NaN. - **Selu signed-zero.** ONNX's `x > 0` sends `-0` down the exp branch; ORT returns `+0`, so both scalar and vector spell it `alpha*(exp(x)-1)` (there is no AVX2 `expm1`). `scalar_paths_agree_with_the_vector_paths` pins scalar and vector to the same answer, and `selu_zero_sign_does_not_depend_on_dtype` pins it across dtypes. Bit-identity where the op is exact (LeakyRelu, ThresholdedRelu asserted bit-identical to their scalars over 20 003 points); HardSigmoid allows one FMA rounding; Selu uses an alpha-scaled bound. --- ## Verification (reviewer, Intel Core i7-13800H, Windows 11) - `cargo test -p onnx-runtime-ep-cpu --lib` — **1415 passed / 0 failed / 17 ignored** for the activation family (all four ops' special-value, scalar-vs-vector bit-parity, dense-sweep and signed-zero tests green). - One unrelated flake observed under full-suite load: `task_runtime::pool::tests::slot_exhaustion_declines_instead_of_blocking` (from #1201's CPU task runtime, untouched by this PR) failed once under contention and passes deterministically in isolation. Not introduced here. - `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` — clean. - #1227's `#[target_feature(enable = "avx2,fma")]` confirmed intact on `map_ps` and `map_bias_ps`. - MLAS remains non-default; all four kernels use `dispatch!` and touch no MLAS path. Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
What this changes
Adds native CPU-EP kernels for
Celu(opset 12) andMish(opset 18), which had no kernel at all and were therefore silently handed to ORT's CPU EP, and vectorisesLog(f32), which was still taking a scalarlibmcall per element. Celu/Mish are wired through all four capability filters (registry,PHASE1_OPS, dtype table, shape-inference table, plugin descriptors) and added to the two activation-coverage tests.Behavioural claim — verified independently (reviewer)
The headline "removes the last silent ORT activation deferrals" is a behavioural claim, so I falsified it rather than trusting it, on Intel Core i7-13800H, Windows 11:
GetCapabilityruns three fail-closed filters; an op absent from the shape-inference table (Filter 2) or the dtype descriptor list (Filter 3) is handed to ORT with no diagnostic. Celu/Mish had no kernel, so they never reachedsupports_op.activation_and_norm_ops_clear_every_capability_filterpasses. Removing just the Celu/Mish entries from the shape-inference table (compute.rs, Filter 2) turns it RED with::Celu: no shape rule (shape filter declines it) … would be handed to ORT's CPU EPfor both ops; restoring returns it to GREEN. The test genuinely guards the deferral.OWNEDlist now enumerates the full activation/normalisation family and holds every member to both the shape and dtype filters; the companionevery_registered_op_has_a_shape_rule_or_is_a_known_gapenumerates all registered ops, and its remaining shape-table gap (data-dependent, internal-fusion, and inferrable-but-unwritten ops) contains no activations.Logwas never a deferral — it already had a scalar kernel — so this PR is a perf change for it, not a coverage change.Correctness
The kernels' own tests (dense sweeps vs the scalar reference, subnormals, and the full IEEE special set) pass; the scalar reference, the f64 path, and the AVX2 path are held to agree, including NaN propagation (
f32::max/mindrop NaN, so each kernel guards it to match ORT). Fullcargo test -p onnx-runtime-ep-cpu --lib= 1362 passed, 0 failed, 17 ignored (base 1353 + 9 new).cargo clippy --all-targets -D warningsclean.map_ps's#[target_feature](#1227) verified intact.Performance — measured on this box, ORT-independent
The gaps doc's ratios were taken on an AMD EPYC 9V74 against ORT. I have no ORT reference on this host, so I measured the ORT-independent internal win instead: the AVX2 kernel vs the scalar-native fallback (which is what runs on non-AVX2 hosts and small tensors), on Intel Core i7-13800H (14C/20T), Windows 11, single thread, 1,048,576 elements, min-of-200 × 9 rounds.
Log(positive domain)Celu(alpha=1, mixed)Mish(mixed)These are hardware-dependent (see
docs/performance/CPU_ACTIVATION_GAPS.md§32 lesson) and are not the same quantity as the doc's ORT-relative ratios — mine compare AVX2 against the native scalar path on this machine; the AMD EPYC figures compare native against ORT. Presented as peers, each with its hardware named. Both agree on direction: the native kernels are a clear win.