Repository navigation
perf(ep-cpu): port MLAS's erf polynomial to AVX2, 3-28x on Erf and exact GELU - #1074
Merged
Merged
Conversation
…act GELU `Erf` and `Gelu(approximate="none")` were the two remaining activations that evaluated a `libm` transcendental per element, and they did it in `f64`. At 1048576 float32 elements the standalone `Erf` node took 39.8 ms against ORT's 0.89 ms — 0.022x. Every other activation in this crate had already been vectorised; these two had not, because an earlier comment asserted that no `f32` polynomial could meet the conformance suite's tolerance. That assertion was wrong twice over. The suite compares at `rtol=1e-4`, and ORT's own CPU `Erf` and `Gelu(none)` kernels call `MlasComputeErf`, which is a faithfully-rounded `f32` polynomial — so matching ORT means *using* that polynomial, not avoiding it. Ports `MlasErfKernel` (`onnxruntime/core/mlas/lib/erf.cpp`, already vendored in `crates/mlas-sys`) to `core::arch::x86_64`, adds the small `exp` it needs, and routes `Erf`, `Gelu(none)`, `com.microsoft::Gelu` and `BiasGelu` through it. The scalar `libm::erf` stays as the non-AVX2 fallback, exactly as `tanh` and `sigmoid` keep theirs. Session-level A/B against ORT 1.28.0, same host, interleaved, 1 thread, reps=21, base ab6cb01 (ratio = ORT/ours, >1 means we win): float32 Erf 0.203 -> 0.738 (512) ... 0.022 -> 0.576 (1048576) float32 Gelu(none) 0.251 -> 0.758 (512) ... 0.039 -> 0.699 (1048576) float16 Gelu(none) 0.414 -> 1.184 (512) ... 0.059 -> 0.898 (1048576) float16 `Gelu` is the range that matters for assignment honesty: the policy claims it unconditionally because ORT has no float16 `Gelu` kernel to defer to, so that 0.059x was a range this EP really was serving 17x slower than ORT. Accuracy, measured over 400 003-point sweeps including both branch boundaries and the saturation clamp: worst error 5.96e-8 (1 ulp below 1.0) for `erf`, 1.19e-7 scaled for exact GELU. Sign, signed zero, +/-Inf saturation to exactly +/-1, NaN propagation and exact oddness are pinned by tests. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1074 +/- ##
=======================================
Coverage 79.53% 79.54%
=======================================
Files 366 366
Lines 156827 156889 +62
Branches 156827 156889 +62
=======================================
+ Hits 124737 124793 +56
- Misses 27387 27396 +9
+ Partials 4703 4700 -3
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
|
| Status | Scenario | Base | PR | Change |
|---|---|---|---|---|
grammar_masking/llguidance_compute_mask/32 |
75.50 µs | 97.94 µs | +29.7% | |
block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 |
788.57 µs | 949.15 µs | +20.4% | |
logit_processing/seven_processor_chain_per_step |
320.52 µs | 383.99 µs | +19.8% | |
qwen3_sampling_processors/top_k_top_p_full_sort_baseline |
5.67 ms | 6.61 ms | +16.5% | |
| ✅ | kv_cache/alloc_dealloc_pages |
38.95 µs | 44.43 µs | +14.1% |
| ✅ | qwen3_sampling_processors/top_k_top_p_fast |
664.28 µs | 755.02 µs | +13.7% |
| ✅ | qwen3_sampling_processors/top_p_fast_after_top_k |
522.70 µs | 582.36 µs | +11.4% |
| ✅ | add/small_f16_threads=1-internal/1024 |
460.2 ns | 492.0 ns | +6.9% |
| ✅ | qwen3_sampling_processors/top_k_partial_selection |
139.53 µs | 146.89 µs | +5.3% |
| ✅ | sampling_latency/top_p_per_token |
383.93 µs | 397.23 µs | +3.5% |
| ✅ | tokenization/encode_tokens_per_second |
384.86 µs | 397.71 µs | +3.3% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 |
177.12 µs | 182.39 µs | +3.0% |
| ✅ | matmul/large_generic_bf16_threads=1/32x1024x1024 |
2.06 ms | 2.11 ms | +2.5% |
| ✅ | matmul/large_generic_bf16_threads=8/32x1024x1024 |
2.09 ms | 2.13 ms | +2.1% |
| ✅ | tokenization/decode_tokens_per_second |
6.44 ms | 6.55 ms | +1.7% |
| ✅ | matmul/medium_generic_bf16_threads=1/32x512x512 |
535.43 µs | 544.26 µs | +1.6% |
| ✅ | matmul/medium_generic_f32_threads=1/32x512x512 |
2.34 ms | 2.38 ms | +1.6% |
| ✅ | matmul/small_generic_bf16_threads=8/1x256x256 |
37.56 µs | 38.15 µs | +1.6% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 |
594.50 µs | 602.98 µs | +1.4% |
| ✅ | block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 |
85.85 µs | 86.87 µs | +1.2% |
| ✅ | qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline |
3.51 ms | 3.55 ms | +1.1% |
| ✅ | sampling_latency/min_p_per_token |
210.63 µs | 212.92 µs | +1.1% |
| ✅ | matmul/small_generic_bf16_threads=1/1x256x256 |
35.83 µs | 36.15 µs | +0.9% |
| ✅ | matmul/large_generic_f16_threads=1/32x1024x1024 |
89.85 µs | 90.66 µs | +0.9% |
| ✅ | sampling_latency/greedy_per_token |
3.20 µs | 3.22 µs | +0.8% |
| ✅ | matmul/small_generic_f16_threads=8/1x256x256 |
37.86 µs | 38.05 µs | +0.5% |
| ✅ | qwen3_sampling_processors/top_k_full_sort_baseline |
2.16 ms | 2.17 ms | +0.4% |
| ✅ | sampling_latency/top_k_per_token |
51.87 µs | 51.99 µs | +0.2% |
| ✅ | matmul/large_generic_f32_threads=8/32x1024x1024 |
6.16 ms | 6.17 ms | +0.1% |
| ✅ | add/large_bf16_threads=1-internal/4194304 |
1.66 ms | 1.66 ms | -0.0% |
| ✅ | matmul/medium_generic_f16_threads=8/32x512x512 |
42.48 µs | 42.16 µs | -0.8% |
| ✅ | matmul/large_generic_f16_threads=8/32x1024x1024 |
99.22 µs | 98.46 µs | -0.8% |
| ✅ | matmul/large_generic_f32_threads=1/32x1024x1024 |
9.54 ms | 9.44 ms | -1.1% |
| ✅ | reduce_mean/large_f32_threads=1-internal/262144 |
1.01 ms | 996.41 µs | -1.3% |
| ✅ | matmul/medium_generic_f16_threads=1/32x512x512 |
35.44 µs | 34.96 µs | -1.4% |
| ✅ | matmul/small_generic_f16_threads=1/1x256x256 |
36.39 µs | 35.83 µs | -1.5% |
| ✅ | reduce_mean/medium_f32_threads=1-internal/65536 |
251.93 µs | 247.96 µs | -1.6% |
| ✅ | reduce_mean/small_f32_threads=1-internal/4096 |
15.13 µs | 14.88 µs | -1.7% |
| ✅ | gather/medium_bf16_threads=1-internal/32768 |
2.40 µs | 2.35 µs | -2.0% |
| ✅ | gather/medium_f16_threads=1-internal/32768 |
2.40 µs | 2.34 µs | -2.6% |
| ✅ | gather/small_bf16_threads=1-internal/4096 |
494.3 ns | 475.5 ns | -3.8% |
| ✅ | gather/large_f16_threads=1-internal/131072 |
13.89 µs | 13.33 µs | -4.0% |
| ✅ | gather/medium_f32_threads=1-internal/32768 |
4.10 µs | 3.90 µs | -4.7% |
| ✅ | add/large_f16_threads=1-internal/4194304 |
1.79 ms | 1.70 ms | -4.8% |
| ✅ | matmul/small_generic_f32_threads=1/1x256x256 |
43.90 µs | 40.87 µs | -6.9% |
| ✅ | matmul/medium_generic_f32_threads=8/32x512x512 |
1.65 ms | 1.53 ms | -7.6% |
| ✅ | matmul/small_generic_f32_threads=8/1x256x256 |
50.42 µs | 46.45 µs | -7.9% |
| ✅ | matmul/medium_generic_bf16_threads=8/32x512x512 |
635.35 µs | 579.47 µs | -8.8% |
| ✅ | gather/large_bf16_threads=1-internal/131072 |
14.07 µs | 12.80 µs | -9.1% |
| ✅ | add/medium_bf16_threads=1-internal/262144 |
113.25 µs | 102.47 µs | -9.5% |
| ✅ | add/medium_f32_threads=1-internal/262144 |
27.52 µs | 24.50 µs | -10.9% |
| ✅ | add/medium_f16_threads=1-internal/262144 |
116.55 µs | 103.27 µs | -11.4% |
| ✅ | gather/large_f32_threads=1-internal/131072 |
39.88 µs | 34.84 µs | -12.6% |
| ✅ | gather/small_f32_threads=1-internal/4096 |
792.4 ns | 692.2 ns | -12.7% |
| ✅ | add/large_f32_threads=1-internal/4194304 |
949.24 µs | 828.48 µs | -12.7% |
| ✅ | gather/small_f16_threads=1-internal/4096 |
572.1 ns | 489.3 ns | -14.5% |
| 🟢 | add/small_bf16_threads=1-internal/1024 |
533.0 ns | 446.5 ns | -16.2% |
| 🟢 | block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 |
116.12 µs | 95.67 µs | -17.6% |
| 🟢 | add/small_f32_threads=1-internal/1024 |
275.1 ns | 193.3 ns | -29.7% |
Visual flags:
Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 6.43 4.14 5.63 }
What this cannot catch
- Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
- Sub-threshold regressions that compound over multiple PRs
- Performance changes that only manifest under GPU execution
- Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)
Independent review of #1074 returned four findings; this applies all four. * Restore the `#[inline]` attribute on `quick_gelu_ps`. An earlier edit in this branch joined it onto the end of the preceding doc-comment line, so it silently became comment text. Neither `fmt` nor `clippy` catches that. * Transcribe `EXP_LOWER_RANGE` exactly as MLAS spells it (`-88.3762626647949f`, bits `0xc2b0c0a5`). The previous `-88.376_264` rounded to `0xc2b0c0a6`, the only constant of the 26 that was not bit-identical. It is provably unreachable for `erf` -- the big branch's `R` maxes at 17.375652 at the `|x| = 3.925` clamp -- so no output changes, but the module's contract is verbatim MLAS. * Document the `SIMD_MIN_LEN` seam that `Erf` now inherits: below 32 elements the correctly-rounded scalar fallback runs, so the same value can differ by <=1 ulp with tensor length. `Tanh` and `Sigmoid` have had this since #1037 and ORT has it too. * Correct the float32 `Erf` 1048576 ratio in the policy rationale from a dispersion-limited 0.58 to a reps=31 re-measurement of 0.687. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
marked this pull request as ready for review
August 16, 2026 15:40
justinchuby
added a commit
that referenced
this pull request
Aug 18, 2026
…ap_bias_ps (#1227) ## What this is `map_ps` is the loop driver behind every unary activation on AVX2. A merge (the `#1037` / `#1074` layering) inserted `map_bias_ps` *between* `map_ps`'s doc comment and `map_ps` itself, so the comment and **both of its attributes** ended up on the wrong function: ```rust /// Apply an 8-lane kernel across a slice. ... <- map_ps's doc #[inline] #[target_feature(enable = "avx2,fma")] <- map_ps's attributes /// Like [`map_ps`], but adds a bias row ... pub(super) unsafe fn map_bias_ps(...) <- ...on map_bias_ps pub(super) unsafe fn map_ps(...) <- bare ``` `map_bias_ps` needs those attributes too, so nothing was broken outright — but `map_ps` was left with no `#[inline]` and no `avx2,fma`. ## Why the missing `target_feature` is a hazard LLVM will not inline a callee that requires a feature its caller lacks (a *strict* superset; equal feature sets inline fine). Every kernel closure `map_ps` takes (`erf_ps`, `tanh_ps`, `sigmoid_ps`, ...) is `avx2,fma`. With `map_ps` compiled at baseline features, those closures are not inlinable into its loop. It is harmless **today** only by luck of ordering: every caller (`tanh_avx2`, `erf_avx2`, ...) is itself `avx2,fma`, so `map_ps` gets inlined *upwards* into the caller first, and once there the closure folds in. That is an inlining decision, not a guarantee. Grow `map_ps` past the inline threshold — which is exactly what an unroll experiment does — and the closure becomes a real call per 8 elements, and a kernel like `erf_ps` has to re-materialise all of its constants on every one of those calls. I hit this while unrolling `map_ps`, which is how it surfaced. ## Evidence that this changes nothing today Disassembled `tanh_avx2` out of the built rlib before and after: ``` diff <(objdump -d ... base) <(objdump -d ... fixed) -> identical ``` Byte-identical, and `erf_avx2` / `sigmoid_avx2` keep their instruction counts (216 / 117). No `call` to any kernel remains in the lib. So this is a latent-hazard and documentation fix with **zero** codegen delta — not a perf change, and it needs no A/B. ## Also recorded here: a negative result on the divide While looking at these loops I measured one algorithmic change and it lost, so it is written down rather than left for someone to retry. `tanh_ps` and `sigmoid_ps` each end in `_mm256_div_ps`. `vdivps ymm` is the only instruction in either loop that is not fully pipelined (Zen 4 retires one about every 4.5 cycles), so replacing it with `vrcpps` + one Newton-Raphson step — `y1 = y0 + y0*(1 - q*y0)`, four pipelined uops for one blocking one — looked like the obvious win. It is not: | case | ORT p50 (control) | divide | rcp + NR | |---|---|---|---| | `bench_tanh_f32_4k` | 0.0031 ms, all 6 runs | 1.518 / 1.525 / 1.526 | 1.550 / 1.570 / 1.570 | Interleaved rebuild-and-alternate, 3 rounds, 1 thread, `taskset -c 8-15`. The 4k case is the only trustworthy one here — it is L1-resident and ORT's own p50 was identical to four decimal places across all six runs, whereas on the 1M cases ORT's p50 moved 4.6x mid-session and neither arm is quotable. Consistent ~3% **regression**. The `rcp` + 2 FMA dependency chain is about as long as the divide it replaces, so a latency-bound kernel gains nothing and just pays two extra uops. It also breaks four exactness proofs that #1121 landed: `tanh(±Inf)` comes back as `0.99999994` instead of exactly `±1.0`, because the input clamp makes `p` and `q` bit-equal at `|v| = 9` and only an exact divide turns that into exactly `1.0`. Buying that back would mean restoring the saturation blend #1121 proved redundant over 1.05 billion inputs — to fund a change that is already slower. Not pursued. No fallback involved; the kernel keeps its own divide. ## Validation - `cargo test --release -p onnx-runtime-ep-cpu --lib` — 1337 passed, 0 failed - `cargo clippy --release -p onnx-runtime-ep-cpu --all-targets` — clean - `cargo fmt --all --check` — clean - `tanh_avx2` disassembly identical to `main` ## Review Opus review: no blockers, no should-fixes. It independently confirmed the three things the PR rests on: every one of the 16 `map_ps` and 2 `map_bias_ps` call sites is already inside a `#[target_feature(enable = "avx2,fma")]` function (so nothing becomes unsound or fails to compile), `map_bias_ps` keeps exactly the attributes it had, and there are **no bare/function-pointer references** to either. It also did the scan I most wanted: all ~130 `fn` definitions in the file were checked for the same orphaned-doc-comment-after-attribute signature that caused this bug. **Zero other anomalies** — every `*_ps` helper and every `*_avx2` dispatcher carries the attributes it should. So this is the only instance. - **NIT (applied)** — the doc comment said LLVM won't inline a callee whose features are "a superset" of the caller's. Read non-strictly that would forbid inlining at *equal* feature sets, which is the opposite of what the fix relies on. Reworded to "requires a feature its caller lacks — a *strict* superset; equal feature sets inline fine". --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root cause
ErfandGelu(approximate="none")were the last two activations in this crate still evaluating alibmtranscendental per element, inf64. Every other activation family was vectorised in #1037. These two were explicitly left behind, with this comment inkernels/gelu.rs:That was wrong on both halves:
conformance/run_onnx_tests.pycompares atrtol=1e-4, atol=1e-5. A faithfully-roundedf32erfis ~1e-7 — three orders of magnitude inside it.f32polynomial.ErfandGelu(approximate="none")both callMlasComputeErf, whose FMA3 kernel evaluates a two-branch polynomial. So "matching ORT" meant adopting that polynomial, not avoiding it.The cost was the largest single gap in the whole activation matrix: at 1048576 float32 elements a standalone
Erfnode took 39.8 ms against ORT's 0.89 ms.Change
Ports
MlasErfKernelfromonnxruntime/core/mlas/lib/erf.cpp— already vendored atcrates/mlas-sys/vendor/mlas/— tocore::arch::x86_64inkernels/simd_activations.rs, alongside thetanh/logisticrationals ported in #1037. Themlascargo feature is off by default, so linkingMlasComputeErfdirectly was not an option for the default build; this is the same reason #1037 ported rather than linked.erf_ps— the two-branch kernel:|x| <= 0.921875usesx·(1 + P(x²)); above it,1 - exp(-R(|x|)). Both branches are evaluated for every lane and merged withor, which works because the inactive branch is forced to+0.0(the small branch byandnot, the big branch because zeroing its input collapses1 - exp(-0)to0). Branch-free is what buys the speedup — far more than the 8× the SIMD width alone gives.exp_ps— the small range-reducedexpthe big branch needs, usingMlasErfConstants' ownexpparameters, pluspower_of_2_ps(MlasPowerOf2Float32x4).erf_gelu_ps/erf_gelu_bias_f32_slice— exact GELU with thex/√2scale fused so the intermediate never reaches memory, and theBiasGelubias folded in-register exactly asFastGelu's already is.Wired into four call sites that previously ran the scalar
f64form:Erf(elementwise.rs),ai.onnx::Gelu(approximate="none")andcom.microsoft::Gelu(gelu.rs), andBiasGelu(contrib_fused.rs).kernels::gelu::exact_geluhad no remaining callers and is deleted; itsf64form survives assimd_activations::erf_gelu_scalar, the non-AVX2 fallback.libm::erfremains the fallback on every non-AVX2 target, so no existing platform changes accuracy or speed.Benchmarks
Session-level A/B (real ORT session latency, not kernel time, so ORT's per-node overhead and our plugin's are both included). Same host, same process, interleaved arms,
taskset -c 0-15, 1 intra-op thread,reps=21(the 1 048 576 row was re-measured atreps=31after review found it dispersion-limited). Ratio = ORT ns / ours; >1 means we are faster. Before =ab6cb0168, after = this branch. Both arms built identically; the float32 rows use a scratchNXRT_EXP_CLAIM_ALLoverride (never committed) to force assignment, because the policy defers float32 today.float32
Erf(measurement-only claim)float32
Gelu(approximate="none")(measurement-only claim)float16
Gelu(approximate="none")— the range this EP actually claimsThe assignment policy claims float16
Geluunconditionally, because ORT has no float16GeluCPU kernel: declining makes ORT inline the function body and this EP then picks up the ungoverned constituents, measured at 0.024–0.049x. So this is a range where the EP really was serving users a 17x-slower kernel.Control: float16
Erf, which the policy defers1.035 / 1.018 / 1.001 / 1.009 / 0.996 / 1.004— flat 1.0, i.e. the deferral still hands the node to ORT and this PR does not accidentally claim it.What this does not claim
ErfandGelu(none)still lose to ORT (0.66–0.76x) and this PR does not change their assignment — they stay deferred. What changed is that they now sit on the same ~0.70x plateau asTanh(0.70–0.75x) andSigmoid(0.70–0.73x) instead of being 30x worse than their neighbours. That residual plateau is not the transcendental: it is this crate's elementwise kernels being single-threaded while ORT spreads the same work over its intra-op pool.assignment_policy.rs's rationale is updated to say so, with the new numbers replacing the stale0.023-0.77x.Correctness
libm::erfis correctly rounded; MLAS's polynomial is faithfully rounded. Measured worst error over 400 003-point sweeps, scaled bymax(1, |x|):erf[-6, 6]+ both branch boundaries + clamp + subnormalserf[-1.5, 1.5](dense, whereerfis steep)[-25, 25]5.96e-8 is exactly one ulp below 1.0 — i.e. faithful, as advertised. The conformance suite's tolerance is
rtol=1e-4.New tests (11 in
simd_activations, all also exercised through the kernels):erf_dense_sweep_matches_f64_reference,erf_dense_sweep_near_origin,erf_gelu_dense_sweep_matches_f64_reference— the sweeps above, againstlibm::erfinf64, asserting a3e-7/4e-7bound.erf_special_values,erf_gelu_special_values—-Inf -> -1,+Inf -> 1,±0keeps its sign,NaNpropagates,±MAXand±1e30saturate. Exact GELU pins-Inf -> 0(the mathematical limit) as the tanh form already does.erf_saturates_to_exactly_one_past_the_clamp— the3.925clamp must not be observable: every input past it returns exactly±1.0, not1 - eps.erf_is_exactly_odd— the sign is applied by anorafter the polynomial rather than carried through it, so exact antisymmetry over 4096 points is a falsifier for the sign mask leaking into the arithmetic.erf_tail_lanes_match_the_vector_body— bit equality between every length in[32, 40)and the full run, so results cannot depend on tensor length mod 8.Existing coverage that had to keep passing:
erf_known_values,erf_odd_symmetry_and_limits,erf_bf16_reaches_dtype_without_touching_formula,gelu_*,bias_gelu_*, and the EP conformance lane.Validation
cargo test -p onnx-runtime-ep-cpu --lib— 1224 passed (debug) and 1224 passed (release).cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings— clean.cargo fmt --all --check— clean.Limitations
libm::erf, unchanged in both speed and accuracy. Results therefore differ by ~1 ulp between an AVX2 host and a non-AVX2 host — the same ISA dependence this module already documents fortanh/sigmoid, and the same one ORT has.Float64inputs are untouched; they keep the exactf64path.Independent review
Reviewed by an independent
claude-opus-4.8Rubber Duck against an adversarial brief (prove the branch merge, prove the constants, prove the special values, prove the benchmarks, find the regression). Verdict GO WITH FINDINGS, no blockers. The reviewer rebuilt both arms from scratch and re-derived rather than re-read:erf.cpp. 25 bit-identical. The 26th,EXP_LOWER_RANGE, was-88.376_264(0xc2b0c0a6) against MLAS's0xc2b0c0a5; the reviewer also proved it unreachable forerf(the big branch'sRmaxes at 17.375652 at the|x| = 3.925clamp, so the argument never approaches -88). Transcribed exactly anyway in866b95c, since the module's contract is "verbatim MLAS".+0.0, so theormerge cannot contaminate the small branch.Erf,Gelu(none)andBiasGeluat widths 1/8/37/40, against ORT 1.28.0 on the same host. NaN sign, NaN payload and sNaN quieting are bit-identical too.Erf1 048 576 cell (0.576) pessimistic against its own re-run; re-measured atreps=31and corrected the row to 0.687 (p90 0.604) with a matching ORT baseline, which also raises the kernel speedup to 31.1x.#[inline]attribute onquick_gelu_psinto the preceding doc comment (neitherfmtnorclippycatches this). Restored in866b95c; unrelated toerfbut a real defect this PR introduced.Erfnow inherits theSIMD_MIN_LEN = 32scalar/vector seam:erf(0.901059926)is0x3f4c2503in a 31-element tensor and0x3f4c2504in a 40-element one, and 286 of 2000 random values move by <=1 ulp purely with tensor length.TanhandSigmoidhave had this since CPU EP: vectorise the approximate activation family (Tanh/Sigmoid/FastGelu/QuickGelu) on AVX2+FMA #1037 and ORT has it too; documented in the module rather than paying 20x on short tensors to remove it.Out of scope, recorded here so it is not lost: ad-hoc benchmark scripts that let an
InferenceSessionoutlive the plugin registration segfault at interpreter teardown..work/sess_ab.pydisposes correctly and exits 0; the reviewer's scratch scripts did not. Nothing in the shipped code is implicated.