Skip to content

perf(cpu): run the MLAS activation routes through run_chunked - #1127

Merged
justinchuby merged 3 commits into
mainfrom
deckard/mlas-parallel
Aug 17, 2026
Merged

justinchuby merged 3 commits into
mainfrom
deckard/mlas-parallel

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 17, 2026 •

Copy link
Copy Markdown
Owner

What

Make the MLAS activation routes go through run_chunked, the shared chunking/parallel seam every pure-Rust route already uses.

One line:

-                f(input, output);
+                run_chunked(input, output, f);

Root cause

dispatch_mlas! called its kernel on the whole tensor and returned:

if input.len() >= SIMD_MIN_LEN {
    let f: fn(&[f32], &mut [f32]) = $mlas;
    f(input, output);   // <-- straight past run_chunked
    return;
}

run_chunked was the only thing that split work across the pool. So every MLAS-routed op ran single-threaded, regardless of how many threads the session had.

This was harmless when it was written — run_chunked was a plain loop then. #1105 taught it to parallelise above PAR_MIN_LEN, and from that moment the two builds diverged: an mlas-off build scaled with the pool, an mlas-on build did not. Since #1115 the wheel ships mlas on x86_64, so the shipped configuration was the non-scaling one.

Nothing caught it because the results were still correct, and correctness is all the existing tests checked. thread_invariance asserts a kernel gives the same answer serial and parallel — which a route that never goes parallel satisfies trivially.

Benchmark

Same machine (AMD EPYC 9V74, 32 vCPU, AVX2+FMA), ab_ort.py, --iters 30 --reps 5, medians of 3 interleaved rounds alternating the two builds, EP assignment asserted from ORT's own profiler (zero NOT-ASSIGNED). Both arms built --features mlas --release from the same base. p50 µs.

16 threads — the ops that take the MLAS route

op n before after speedup
Erf 1 Mi 897.3 393.9 2.28x
Erf 4 Mi 3558.7 561.6 6.34x
Gelu (exact) 1 Mi 1304.0 552.8 2.36x
Gelu (exact) 4 Mi 5458.7 1134.3 4.81x
Tanh 1 Mi 480.8 300.4 1.60x
Tanh 4 Mi 2219.6 425.9 5.21x
Sigmoid 1 Mi 503.5 298.6 1.69x
Sigmoid 4 Mi 1991.6 430.0 4.63x

(Tanh/Sigmoid measured on a base that predates #1124; they no longer take the MLAS route, but they exercised the same mechanism and are kept here as evidence of it.)

16 threads — ops with no MLAS route, as the control

Unchanged within noise, which is what says the win is the routing change and not drift:

op n before after
FastGelu 1 Mi 413.9 395.5
FastGelu 4 Mi 600.2 605.4
QuickGelu 1 Mi 329.7 347.1
QuickGelu 4 Mi 529.1 545.9
Sqrt 1 Mi 255.3 232.3
Sqrt 4 Mi 352.8 375.1

1 thread — no regression

run_chunked returns before touching rayon below PAR_MIN_LEN, and par_chunk_len declines to split when there is nothing to gain, so the single-threaded path is untouched:

op n before after
Erf 4 Mi 3628.0 3643.8
Gelu 4 Mi 5466.7 5479.0
Erf 65 Ki 62.2 63.5
Gelu 65 Ki 86.7 86.4

Every cell is inside run-to-run noise (<0.5%).

Correctness

run_chunked splits a slice into disjoint sub-slices and calls the same function on each. The MLAS activation entry points (MlasComputeErf, MlasComputeGeluErf) are elementwise and take no threadpool argument, so splitting cannot change a result: element i depends only on input i. The existing thread_invariance::unary_kernels_are_thread_count_invariant already covers erf and erf_gelu and asserts serial and parallel agree bit for bit — it now actually exercises the parallel branch for them, where before it compared serial against serial.

Regression test

parallel_reachability asserts the mechanism rather than the output: a tensor over PAR_MIN_LEN, submitted from outside the pool, must increment run_chunked's parallel-branch counter.

Proven to falsify — reverting just the one line makes it fail with:

erf: a 1052675-element call did not reach run_chunked's parallel branch, so it
runs single-threaded no matter how large the pool is. A kernel that calls its
backend directly and returns will fail here.

while pure_rust_kernels_go_through_run_chunked keeps passing, so it is not a blanket assertion that would pass on anything. It also guards the two preconditions that could make it vacuous (pool must have ≥2 threads; must not already be inside the pool).

The counter is #[cfg(test)] and thread-local, not a global atomic, so concurrently running tests cannot bump each other's count. Non-test builds get an empty #[inline(always)] function.

Tests

  • cargo test -p onnx-runtime-ep-cpu --features mlas --lib → 1312 passed, 0 failed
  • cargo test -p onnx-runtime-ep-cpu --lib (no mlas) → 1294 passed, 0 failed
  • cargo clippy -p onnx-runtime-ep-cpu --features mlas --lib --tests → clean
  • cargo fmt --all -- --check → clean

A pre-existing intermittent crash, ruled out as mine

While validating I hit an intermittent SIGSEGV in the full onnx-runtime-ep-cpu lib test binary. It is not from this change:

  • Activation tests alone, 40 consecutive runs on this branch: 0 failures.
  • Full suite, 14 consecutive runs on this branch: 0 segfaults.
  • Full suite, 14 consecutive runs on pristine origin/main: 1 segfault.
  • origin/main also fails kernels::sdpa::tests::sdpa_dispatch_matches_scalar_oracle_across_shapes intermittently (~1 in 8 full-suite runs), unrelated to activations.

Flagging for whoever owns sdpa/qmoe; not chased here as it is outside this PR's scope and reproduces without it.

Limitations

  • The thresholds are still PAR_MIN_LEN = 1 Mi and PAR_MIN_CHUNK = 256 Ki. Parallelise the f32 elementwise activation kernels #1105 recorded that one global constant cannot serve kernels with different per-element costs, and this PR does not change that — it only makes the MLAS routes obey whatever the constants say. Both are now on the same footing, which is the precondition for tuning them per kernel.
  • This closes a self-inflicted gap; it does not by itself make these ops beat ORT at 16 threads. At 4 Mi, Erf goes from 0.11x to 0.66x of ORT and Gelu from 0.09x to 0.42x. Real wins, still short of parity, and the remaining multi-thread gap stays open.
  • Measured on one 32-vCPU host.

Review

Independent Opus review: GO WITH FINDINGS.

The reviewer independently confirmed the parts that could have been wrong: that run_chunked splits into disjoint sub-slices at identical input/output offsets (so erf_gelu_mlas's internal 8192-element blocking and its (x, y)-paired NaN repair still line up inside a chunk, trailing partial block included); that MlasComputeErf/MlasComputeGeluErf are elementwise with no threadpool argument and no shared mutable state, so concurrent calls from rayon workers are sound; that the SIMD_MIN_LEN (32) / PAR_MIN_LEN (1 Mi) composition has no gap and short tensors still return before touching rayon; that the counter increments on the calling thread and cannot be cross-contaminated by concurrent tests; that note_parallel_dispatch compiles to nothing in a non-test build; and that the speedup arithmetic and the "serial against serial" claim are accurate.

Two findings:

  1. (MAJOR, pre-existing, outside this file) The same pathology exists in three sibling elementwise activations. silu_f32_slice (kernels/activations.rs:332), relu_contiguous_f32_mlas (kernels/relu.rs:144) and the two Clip paths (kernels/selection.rs:182, kernels/conv.rs:732) all call mlas_sys::compute_* on the whole tensor with no chunking seam — and mlas-sys documents those entry points as "Single threaded; callers shard across threads themselves". silu_f32_slice is worse than the ops fixed here: it follows the MLAS call with a second full-tensor serial correction loop. SiLU is the SwiGLU activation, so this is a hot path.

    Confirmed by grep: run_chunked/run_chunked_rows exist only in simd_activations.rs, and every entry point inside that file does route through the seam — so this PR is complete for its stated scope. Being a different subsystem with its own correctness wrinkle (the correction pass), it gets its own PR rather than being folded in here.

  2. (NIT) parallel_reachability hard-fails rather than skips on a single-threaded pool. This is deliberate and mirrors the pre-existing thread_invariance::assert_same guard, so it adds no environmental assumption the suite did not already make — a test that silently skipped would be worse, since the bug it guards is invisible in the output.

justinchuby and others added 3 commits August 17, 2026 12:56
The MLAS route called its kernel directly on the whole tensor and returned,
so it never went through `run_chunked` -- the shared chunking/parallel seam
every pure-Rust route uses. That was invisible until #1105 taught
`run_chunked` to parallelise: from then on an `mlas`-on build ran these ops
single-threaded no matter how large the pool, while an `mlas`-off build
scaled.

Since #1115 the wheel ships `mlas` on x86_64, so this was shipping. Measured
at 16 threads, p50 us, medians of 3 interleaved rounds with EP assignment
asserted:

           1 Mi                4 Mi
  Erf      897 ->  394 (2.3x)  3559 ->  562 (6.3x)
  Gelu    1304 ->  553 (2.4x)  5459 -> 1134 (4.8x)
  Tanh     481 ->  300 (1.6x)  2220 ->  426 (5.2x)
  Sigmoid  504 ->  299 (1.7x)  1992 ->  430 (4.6x)

Ops with no MLAS route are the control and are unchanged within noise
(FastGelu 414 -> 395 / 600 -> 605, QuickGelu 330 -> 347 / 529 -> 546,
Sqrt 255 -> 232 / 353 -> 375).

At 1 thread nothing moves: Erf 3628 -> 3644, Gelu 5467 -> 5479, every other
cell inside run-to-run noise.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The bypass produced correct results, so nothing caught it -- the
thread-invariance tests pass trivially for a route that never splits. This
asserts the mechanism instead: a tensor over PAR_MIN_LEN, submitted from
outside the pool, must increment run_chunked's parallel-branch counter.

Verified to fail on the pre-fix code (`erf: a 1052675-element call did not
reach run_chunked's parallel branch`) while the pure-Rust control kept
passing, so it is not a blanket assertion.

The counter is thread-local and `#[cfg(test)]`, so concurrent tests cannot
bump each other's count and non-test builds get an empty inlined function.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 17, 2026 14:37
@justinchuby
justinchuby merged commit b5309f7 into main Aug 17, 2026
11 of 16 checks passed
@justinchuby
justinchuby deleted the deckard/mlas-parallel branch August 17, 2026 14:37
@codecov

codecov Bot commented Aug 17, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.73%. Comparing base (233a636) to head (65b1f27).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1127      +/-   ##
==========================================
+ Coverage   80.71%   80.73%   +0.01%     
==========================================
  Files         368      368              
  Lines      159904   160093     +189     
  Branches   159904   160093     +189     
==========================================
+ Hits       129066   129244     +178     
- Misses      26120    26131      +11     
  Partials     4718     4718              
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.31% <ø> (ø)
mlas 84.83% <ø> (ø)
offline 80.54% <100.00%> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...nnx-runtime-ep-cpu/src/kernels/simd_activations.rs 97.45% <100.00%> (+0.04%) ⬆️

... and 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/medium_generic_bf16_threads=1/32x512x512 579.98 µs 1.23 ms +111.7%
🔴 gather/large_f16_threads=1-internal/131072 20.05 µs 34.50 µs +72.1%
🔴 gather/large_bf16_threads=1-internal/131072 16.60 µs 27.71 µs +67.0%
🔴 gather/medium_f32_threads=1-internal/32768 3.96 µs 5.72 µs +44.4%
🔴 matmul/medium_generic_f32_threads=1/32x512x512 2.44 ms 3.44 ms +41.3%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 5.10 ms 6.91 ms +35.3%
🔴 gather/small_f32_threads=1-internal/4096 684.6 ns 925.4 ns +35.2%
⚠️ matmul/medium_generic_f16_threads=8/32x512x512 38.71 µs 49.75 µs +28.5%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 34.46 µs 43.23 µs +25.5%
⚠️ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 66.32 µs 82.20 µs +24.0%
⚠️ matmul/large_generic_bf16_threads=1/32x1024x1024 2.08 ms 2.56 ms +23.2%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 89.65 µs 110.06 µs +22.8%
⚠️ gather/large_f32_threads=1-internal/131072 63.29 µs 77.47 µs +22.4%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 103.54 µs 124.38 µs +20.1%
⚠️ matmul/small_generic_f16_threads=8/1x256x256 39.30 µs 46.15 µs +17.4%
⚠️ gather/small_bf16_threads=1-internal/4096 509.5 ns 598.0 ns +17.4%
✅ matmul/small_generic_bf16_threads=8/1x256x256 42.81 µs 47.53 µs +11.0%
✅ matmul/small_generic_f32_threads=1/1x256x256 42.35 µs 46.99 µs +10.9%
✅ gather/small_f16_threads=1-internal/4096 494.2 ns 547.6 ns +10.8%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.35 ms 1.46 ms +8.3%
✅ gather/medium_f16_threads=1-internal/32768 2.85 µs 3.08 µs +8.1%
✅ add/large_bf16_threads=1-internal/4194304 2.09 ms 2.24 ms +7.2%
✅ matmul/small_generic_bf16_threads=1/1x256x256 38.26 µs 40.70 µs +6.4%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 91.30 µs 96.77 µs +6.0%
✅ reduce_mean/small_f32_threads=1-internal/4096 18.99 µs 19.92 µs +4.9%
✅ sampling_latency/top_p_per_token 429.29 µs 445.51 µs +3.8%
✅ matmul/small_generic_f16_threads=1/1x256x256 37.99 µs 39.38 µs +3.6%
✅ grammar_masking/llguidance_compute_mask/32 81.59 µs 84.13 µs +3.1%
✅ sampling_latency/top_k_per_token 57.93 µs 59.43 µs +2.6%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 554.20 µs 552.06 µs -0.4%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.13 ms 9.93 ms -2.0%
✅ add/large_f32_threads=1-internal/4194304 919.47 µs 897.83 µs -2.4%
✅ reduce_mean/medium_f32_threads=1-internal/65536 315.87 µs 307.20 µs -2.7%
✅ sampling_latency/min_p_per_token 219.41 µs 213.16 µs -2.8%
✅ matmul/small_generic_f32_threads=8/1x256x256 63.71 µs 61.69 µs -3.2%
✅ sampling_latency/greedy_per_token 4.19 µs 4.06 µs -3.2%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.78 ms 1.72 ms -3.7%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.26 ms 2.16 ms -4.4%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.04 ms 5.71 ms -5.6%
✅ add/large_f16_threads=1-internal/4194304 2.19 ms 2.06 ms -6.0%
✅ qwen3_sampling_processors/top_k_partial_selection 150.89 µs 140.12 µs -7.1%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 189.44 µs 175.17 µs -7.5%
✅ logit_processing/seven_processor_chain_per_step 349.96 µs 323.38 µs -7.6%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.83 ms 3.51 ms -8.3%
✅ tokenization/decode_tokens_per_second 7.42 ms 6.74 ms -9.1%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.33 ms 1.20 ms -9.4%
✅ kv_cache/alloc_dealloc_pages 42.92 µs 38.35 µs -10.6%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 526.31 µs 455.41 µs -13.5%
🟢 tokenization/encode_tokens_per_second 517.91 µs 439.36 µs -15.2%
🟢 qwen3_sampling_processors/top_k_top_p_fast 774.66 µs 655.05 µs -15.4%
🟢 add/medium_f16_threads=1-internal/262144 135.59 µs 113.83 µs -16.0%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.03 ms 860.03 µs -16.3%
🟢 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 615.66 µs 505.17 µs -17.9%
🟢 add/medium_f32_threads=1-internal/262144 31.93 µs 26.05 µs -18.4%
🟢 add/medium_bf16_threads=1-internal/262144 159.20 µs 127.53 µs -19.9%
🟢 add/small_f32_threads=1-internal/1024 287.8 ns 209.4 ns -27.2%
🟢 add/small_bf16_threads=1-internal/1024 637.1 ns 461.6 ns -27.5%
🟢 gather/medium_bf16_threads=1-internal/32768 4.50 µs 3.20 µs -28.9%
🟢 add/small_f16_threads=1-internal/1024 724.8 ns 492.3 ns -32.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 5.26 4.02 7.92 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 17, 2026
#1130)

## Root cause

`run_chunked` is the seam that #1105 taught to split activation work
across the rayon
pool. #1127 fixed `dispatch_mlas!`, which called its kernel directly and
returned,
bypassing that seam entirely — so every MLAS-routed op ran single
threaded no matter how
many threads were configured.

The independent review on #1127 reported as a **MAJOR** finding that
three more callers
have the identical pathology, in other files, so they were out of scope
there:

| caller | file |
|---|---|
| `silu_f32_slice` | `kernels/activations.rs` |
| `relu_contiguous_f32_mlas` | `kernels/relu.rs:144` |
| `Clip` (selection) | `kernels/selection.rs:182` |
| `Clip` (conv epilogue) | `kernels/conv.rs:732` |

`mlas-sys` documents `compute_silu`, `compute_relu` and `compute_clip`
as *"Single
threaded; callers shard across threads themselves."* Nobody sharded.
This PR wraps all
four call sites in `run_chunked`.

SiLU needed one extra step. Its MLAS route is followed by a correction
scan over the
whole tensor (MLAS's `compute_silu` is inaccurate outside
`±SILU_MLAS_SAFE_BOUND`).
Run whole-tensor, that scan streams the buffer a second time from DRAM.
It is now
blocked at `SILU_CORRECTION_BLOCK = 8192` so each block stays in L2, and
the scan is a
branch-free OR-reduction **over the input only** — the predicate
`!x.is_finite() || x.abs() > SILU_MLAS_SAFE_BOUND` depends solely on the
input, so the
common all-in-band case skips the write loop entirely.

## Benchmarks

Session level through the plugin `.so`, base = `origin/main` @
`b5309f799`, 16 threads,
3 interleaved rounds, randomised order, `# NOT-ASSIGNED: 0` on every run
(no node was
left to ORT's CPU EP). µs, p50.

| op | n | base | this PR | speedup | ORT | ORT-rel before | ORT-rel
after |
|---|---:|---:|---:|---:|---:|---:|---:|
| Clip | 1 Mi | 351.36 | 256.98 | **1.37×** | 34.73 | 0.099 | 0.135 |
| Clip | 4 Mi | 1284.90 | 516.92 | **2.49×** | 88.91 | 0.069 | 0.172 |
| Relu | 1 Mi | 334.00 | 252.63 | **1.32×** | 40.42 | 0.121 | 0.160 |
| Relu | 4 Mi | 1095.22 | 494.47 | **2.22×** | 78.41 | 0.072 | 0.159 |
| Swish | 1 Mi | 1415.31 | 404.02 | **3.50×** | 226.21 | 0.160 | 0.560 |
| Swish | 4 Mi | 5629.53 | 549.99 | **10.24×** | 474.99 | 0.084 |
**0.864** |

`Swish` (default domain, opset 24) is the ORT-visible spelling of SiLU
and is supported
by ORT 1.28, so SiLU does have a real single-node session-level A/B
after all — the
earlier note that it did not was wrong, and it is the op that gains the
most here.

Kernel-level SiLU, `serial_scope` vs parallel in-process, 32 threads, so
the MLAS route
is compared against itself with only the split changed:

| n | serial | parallel | speedup |
|---:|---:|---:|---:|
| 1 Mi | 2051.4 | 641.6 | 3.20× |
| 4 Mi | 8231.0 | 1266.0 | 6.50× |
| 16 Mi | 33589.2 | 3380.0 | 9.94× |

## Correctness

- `blocked_correction_matches_the_whole_tensor_loop_bit_for_bit` — the
blocked,
OR-reduced scan is compared bit for bit against the original
whole-tensor loop over
in-band values, out-of-band values, `±Inf`, NaN, `±0`, denormals and
values sitting
exactly on `SILU_MLAS_SAFE_BOUND`, at lengths that straddle the block
boundary.
- `silu_reaches_run_chunked_parallel_branch` — asserts the
**mechanism**, not the
output, using the `PARALLEL_DISPATCHES` counter added in #1127. Verified
to falsify:
  reverting the `run_chunked` wrapper makes it fail.
- `silu_is_thread_count_invariant` — identical results across pool
sizes.
- No tolerance was relaxed anywhere. No numerical behaviour changes:
this PR only
changes *who* runs the arithmetic, plus a blocking/reduction rewrite
that is proven
  bit identical.

`cargo test -p onnx-runtime-ep-cpu --features mlas --lib` → **1322
passed, 0 failed**.
Both feature configurations build. `cargo fmt` clean.

## Limitations

- **Clip, Relu and SiLU still lose to ORT** at these sizes
(0.135–0.864×). This PR is a
2.2–10.2× step toward the architectural requirement that our CPU EP beat
ORT on every
op it accepts; it does not finish the job, and no fallback was added.
The remaining
gap is the general 16-thread scaling gap tracked in
`docs/performance/CPU_ACTIVATION_GAPS.md`
— ORT scales these ops ~14× from 1→16 threads, we manage ~6×, because we
split over
our own rayon pool rather than ORT's intra-op pool. The `host_parallel`
seam over
  `KernelContext_ParallelFor` is the next step.
- 1 Mi is exactly `PAR_MIN_LEN`, so gains there are smaller and noisier
than at 4 Mi.
- 16-thread medians on this shared machine are noisy; untouched control
ops swung up to
36% across 3 rounds. The 4 Mi wins are far outside that band. The 1 Mi
Clip/Relu
  numbers are closer to it and should be read as directional.
- `run_chunked`, `PAR_MIN_LEN` and `parallel_dispatches` are widened to
`pub(crate)`
  because the three other callers live in sibling modules.

---------

Co-authored-by: Deckard <deckard@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant