Skip to content

perf(cpu): route Transpose and the elementwise fallback through the task runtime - #1207

Merged
justinchuby merged 5 commits into
mainfrom
squad/sebastian-cpu-task-runtime-transforms
Aug 18, 2026
Merged

justinchuby merged 5 commits into
mainfrom
squad/sebastian-cpu-task-runtime-transforms

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 18, 2026 •

Copy link
Copy Markdown
Owner

perf(cpu): route Transpose and the elementwise fallback through the task runtime

Moves the transpose blocked-copy paths (block_move, blocked_2d,
batched_blocked_2d) and the SIMD elementwise-activation fallback off
their ad-hoc rayon fan-out and onto the CPU task runtime landed in #1201,
so they share one pool with RoPE/Softmax instead of spinning up a second
scheduler.

What changed

  • Transpose partitions on a MIN_PARALLEL_TASK_BYTES (64 KiB) balance
    unit and fans out through task_runtime::chunk_runs_mut, with the
    task count derived from task_runtime::width().
  • The elementwise fallback routes through run_on_task_runtime, guarded
    by rayon::current_thread_index().is_some() || task_runtime::in_task()
    so a fan-out already inside a pool worker runs serially instead of
    nesting a second fan-out (this is the oversubscription guard).

Verification (this machine)

Hardware: Intel Core i7-13800H, 14 physical / 20 logical cores
(6 P-cores + 8 E-cores). Baseline: main at the rebase point.
Toolchain: cargo 1.97.1, release for any timing.

  • Correctness / bit-identity. These are pure scheduling changes:
    reordering parallel work must not change per-element math. The
    transpose and elementwise unit tests (which compare against a scalar
    reference) all pass, so output is bit-identical to base. cargo test -p onnx-runtime-ep-cpu --lib = 1405 passed, 0 failed, 17 ignored
    (baseline on the rebase point was the same count; no test regressed).
  • Routing is reachable. Production transpose call sites call
    task_runtime::chunk_runs_mut; the elementwise fallback reaches
    run_on_task_runtime. The only remaining rayon in transpose.rs is
    the #[cfg(test)] bench harness.
  • No deadlock/starvation under contention. The pool tests were run
    with 20 background CPU spinners saturating every logical core: 13
    passed, 0 failed. See the fixed slot-exhaustion test below.
  • Clippy --all-targets -- -D warnings: clean.

Flaky test fixed forward

task_runtime::pool::tests::slot_exhaustion_declines_instead_of_blocking
slept a fixed 50 ms and assumed the OS scheduled all SLOT_COUNT holder
threads into their dispatch bodies inside that window. Under the full lib
suite (cargo runs test binaries in parallel, saturating 20 logical CPUs)
some holders had not yet claimed a slot, the probe found a free slot, and
the "declines instead of blocking" assertion flaked. It now waits on a
per-holder held counter until every slot is provably occupied (10 s
deadline so a starved runner fails loudly rather than hanging). Green
both in isolation and under 20-spinner background load.

Perf

The transpose / activation / control tables in
docs/benchmarks/2026-08-15-cpu-ep-vs-ort-attention-moe.md §33.8 are the
author's native-ms A/B on this host class, with symmetric untouched
controls (GQA attention, GEMM) establishing the ±25% single-cell noise
floor. Transpose is bandwidth-bound and the gain is small and lives at
the wide end (t=16/32); the activation fallback shows up to ~3.6× at
t=16 and ~4× at t=32. I did not independently re-run the full 15-trial
ORT sweep; I verified the mechanism (width-based partitioning + the
nesting guard) and correctness. Doc subsections that were mislabelled
31.7/31.8/31.9 (stale from before the task-runtime section became 33)
are renumbered to 33.7/33.8/33.9 to stop colliding with the genuine §31.

Closes #1207. Working as sebastian (CPU perf).

@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from 3cf4f18 to e1ca38a Compare August 18, 2026 05:31
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from f9cb1aa to c6dc23c Compare August 18, 2026 05:31
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from e1ca38a to 53f362c Compare August 18, 2026 05:34
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from c6dc23c to 9356172 Compare August 18, 2026 05:34
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from f4d2634 to 6c4d69d Compare August 18, 2026 05:40
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from 3dc2157 to 1d12b3c Compare August 18, 2026 05:41
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from 6c4d69d to 0d88434 Compare August 18, 2026 05:47
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from 1d12b3c to 5e48ff1 Compare August 18, 2026 05:47
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch 2 times, most recently from d025c32 to abcbfd1 Compare August 18, 2026 06:05
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch 2 times, most recently from bc4646a to 29f7a67 Compare August 18, 2026 06:06
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from abcbfd1 to 7769607 Compare August 18, 2026 06:08
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from 29f7a67 to 132d15a Compare August 18, 2026 06:08
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from 7769607 to 3fedb98 Compare August 18, 2026 06:30
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from 132d15a to 42aaf72 Compare August 18, 2026 06:30
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from 3fedb98 to 0eb4888 Compare August 18, 2026 06:50
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch 2 times, most recently from 19907d2 to e52a6e7 Compare August 18, 2026 07:07
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from a4395c9 to bab6433 Compare August 18, 2026 07:13
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from e52a6e7 to 906f201 Compare August 18, 2026 07:14
@justinchuby
justinchuby marked this pull request as ready for review August 18, 2026 07:16
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from bab6433 to 8fe0695 Compare August 18, 2026 07:35
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from 906f201 to ebcc6cd Compare August 18, 2026 07:35
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from ebcc6cd to b4caf90 Compare August 18, 2026 08:05
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from 26a673c to cd6e3e1 Compare August 18, 2026 09:12
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from b4caf90 to 8fbb204 Compare August 18, 2026 09:12
@github-actions

github-actions Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/small_generic_f32_threads=8/1x256x256 35.16 µs 51.63 µs +46.9%
🔴 gather/large_f32_threads=1-internal/131072 32.30 µs 43.13 µs +33.5%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 1.41 ms 1.87 ms +32.4%
⚠️ matmul/small_generic_f16_threads=8/1x256x256 30.85 µs 40.09 µs +30.0%
⚠️ matmul/small_generic_f16_threads=1/1x256x256 30.92 µs 38.66 µs +25.1%
⚠️ matmul/small_generic_bf16_threads=8/1x256x256 32.84 µs 39.90 µs +21.5%
✅ qwen3_sampling_processors/top_k_partial_selection 138.65 µs 155.31 µs +12.0%
✅ matmul/small_generic_bf16_threads=1/1x256x256 32.17 µs 35.83 µs +11.4%
✅ sampling_latency/min_p_per_token 201.82 µs 218.19 µs +8.1%
✅ gather/large_f16_threads=1-internal/131072 15.96 µs 17.23 µs +7.9%
✅ sampling_latency/top_k_per_token 54.89 µs 58.77 µs +7.1%
✅ add/medium_f16_threads=1-internal/262144 101.70 µs 108.57 µs +6.8%
✅ add/large_f32_threads=1-internal/4194304 677.08 µs 720.89 µs +6.5%
✅ sampling_latency/top_p_per_token 374.73 µs 397.60 µs +6.1%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.37 ms 2.50 ms +5.6%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 134.84 µs 142.03 µs +5.3%
✅ matmul/small_generic_f32_threads=1/1x256x256 47.87 µs 50.38 µs +5.3%
✅ add/medium_bf16_threads=1-internal/262144 101.13 µs 104.90 µs +3.7%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.59 µs 16.17 µs +3.7%
✅ add/medium_f32_threads=1-internal/262144 24.99 µs 25.86 µs +3.5%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.51 ms 3.60 ms +2.6%
✅ add/small_f32_threads=1-internal/1024 201.9 ns 206.1 ns +2.1%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.79 ms 5.86 ms +1.2%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.00 ms 10.05 ms +0.5%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 7.21 ms 7.20 ms -0.1%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 677.79 µs 674.95 µs -0.4%
✅ matmul/medium_generic_f16_threads=8/32x512x512 44.08 µs 43.70 µs -0.9%
✅ reduce_mean/medium_f32_threads=1-internal/65536 255.33 µs 251.67 µs -1.4%
✅ add/large_bf16_threads=1-internal/4194304 1.73 ms 1.70 ms -1.5%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 545.06 µs 535.23 µs -1.8%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 97.08 µs 95.30 µs -1.8%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 130.67 µs 128.24 µs -1.9%
✅ sampling_latency/greedy_per_token 3.35 µs 3.29 µs -1.9%
✅ add/small_bf16_threads=1-internal/1024 452.3 ns 441.4 ns -2.4%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.04 ms 1.00 ms -3.3%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.21 ms 2.13 ms -3.7%
✅ add/small_f16_threads=1-internal/1024 471.8 ns 451.5 ns -4.3%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.30 ms 2.20 ms -4.4%
✅ logit_processing/seven_processor_chain_per_step 339.75 µs 323.11 µs -4.9%
✅ grammar_masking/llguidance_compute_mask/32 80.94 µs 76.88 µs -5.0%
✅ matmul/medium_generic_f16_threads=1/32x512x512 41.48 µs 39.27 µs -5.3%
✅ gather/small_f16_threads=1-internal/4096 514.7 ns 483.4 ns -6.1%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 2.48 ms 2.32 ms -6.6%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 630.34 µs 587.87 µs -6.7%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.32 ms 1.22 ms -7.5%
✅ gather/large_bf16_threads=1-internal/131072 19.12 µs 17.65 µs -7.7%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 102.28 µs 93.15 µs -8.9%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 111.23 µs 99.31 µs -10.7%
✅ kv_cache/alloc_dealloc_pages 45.08 µs 39.67 µs -12.0%
✅ qwen3_sampling_processors/top_k_top_p_fast 746.14 µs 655.93 µs -12.1%
✅ add/large_f16_threads=1-internal/4194304 2.09 ms 1.78 ms -14.6%
🟢 tokenization/decode_tokens_per_second 7.72 ms 6.55 ms -15.2%
🟢 gather/small_f32_threads=1-internal/4096 811.9 ns 682.0 ns -16.0%
🟢 tokenization/encode_tokens_per_second 479.75 µs 396.64 µs -17.3%
🟢 matmul/medium_generic_bf16_threads=1/32x512x512 710.49 µs 573.99 µs -19.2%
🟢 gather/small_bf16_threads=1-internal/4096 841.9 ns 506.9 ns -39.8%
🟢 gather/medium_f32_threads=1-internal/32768 7.65 µs 4.30 µs -43.8%
🟢 gather/medium_bf16_threads=1-internal/32768 5.56 µs 2.44 µs -56.1%
🟢 gather/medium_f16_threads=1-internal/32768 5.71 µs 2.48 µs -56.5%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.17 3.55 4.71 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 18, 2026
## What

`main` at 4d231ea fails two blocking CI steps on current stable
(rustc/rustfmt 1.9.0, 1.97.1):

```
$ cargo fmt --all -- --check
Diff in crates/onnx-runtime-ep-cuda/src/provider.rs
Diff in crates/onnx-runtime-memory-api/src/allocator.rs
Diff in crates/onnx-runtime-memory-api/src/capability.rs

$ cargo clippy --locked --all-targets -p onnx-runtime-session -- -D warnings
error: field `0` is never read
   --> crates/onnx-runtime-session/src/executor/mod.rs:175:45
```

Both landed while the runner queue was saturated (71 queued / 1 in
progress at the time of writing, nothing completed on `main` since
08:18Z), so no PR has seen a red check yet. Every open PR in the repo
currently inherits both failures.

## Why these fixes

**rustfmt** — mechanical normalisation, no semantic change.

**`ActivationPlanForTest`** (from #1226) is a tuple struct whose single
field is the `globals_lock()` `MutexGuard`. `dead_code` does not model
"this field's value is its `Drop`", so it fires. The guard must stay:
releasing it early is exactly the leaked-planner-gate race the struct
was added to prevent. So the lint is silenced with a comment explaining
the RAII intent, rather than the field removed.

## Verification

| check | result |
| --- | --- |
| `cargo fmt --all -- --check` | clean |
| `cargo clippy --locked --all-targets -p onnx-runtime-session -- -D
warnings` | clean |
| `cargo test --locked -p onnx-runtime-session --lib` | 181 passed, 0
failed |

Found while trying to land the CPU task-runtime stack (#1201 → #1202 →
#1207 → #1232 → #1238); this is unrelated to that work and is
deliberately kept out of it so it can go in on its own.

Working as sebastian (Performance Engineer)

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-kernels branch from cd6e3e1 to 28c75ea Compare August 18, 2026 16:14
Base automatically changed from squad/sebastian-cpu-task-runtime-kernels to main August 18, 2026 16:16
justinchuby and others added 5 commits August 18, 2026 09:16
…untime

Two more callers off Rayon and onto the CPU task runtime.

Transpose split its output into `units / workers` bands -- one static
share per worker, so the whole copy waited for whichever band was
unlucky. It now hands the runtime chunks of at least
MIN_PARALLEL_TASK_BYTES (64 KiB, a quarter of the whole-operation bar,
still comfortably bandwidth-bound) and lets the runtime claim them
dynamically and cap the task count. All three split sites --
`block_move`, `blocked_2d`, `batched_blocked_2d` -- use
`chunk_runs_mut`, so a task still gets one contiguous band in one call
and none of the three bodies changes shape.

`simd_activations` already borrowed ORT's pool inside the plugin EP.
Its *other* arm -- a native session, or an ORT session whose intra-op
pool is one thread, where `prefer_host` correctly declines -- still went
to Rayon and still paid the park latency this whole exercise is about:
67us for an isolated region back-to-back, 226us when the previous region
ended 20us earlier, which is exactly the spacing between two activations
in a decode step. That arm now goes to the native pool.

The new tail is outlined in `run_on_task_runtime` rather than written
inline, for the reason recorded on `try_host`: growing `run_chunked`
repartitions codegen units, and did cost `Relu` 34% at 1 Mi once already
with no path change at all.

The nesting guards now test `task_runtime::in_task()` as well as
`rayon::current_thread_index()`, because both pools can be the outer
region while the crate is mid-migration.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The elementwise ops are the ones most exposed to scheduling overhead --
a Gelu over an MLP intermediate is a few hundred microseconds of
arithmetic wrapped around a fan-out -- so they are the sharpest test of
a task runtime, and the harness had no cells for them.

The sweep is by token count at fixed hidden/intermediate widths, which
walks the tensor from L2-resident through L3 to memory-resident while
holding the fan-out cost fixed. That is the axis that decides whether
parallelising is worth it at all, so it is the axis a scheduler change
has to be measured along.

FastGelu rather than Gelu because Gelu only joins the default domain at
opset 20, and FastGelu is what a real transformer graph carries here.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ls, and the GEMM finding

Adds scripts/ort_ab/gen_gemm.py, which generates the MatMulNBits and dense
MatMul cells the control grid needed. That control was meant to be a null
result and instead found the largest remaining loss in the CPU EP: decode-shaped
quantised GEMM is 4.6x slower at 32 threads than at 16.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pure rustfmt output, no behaviour change.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…oc numbering

The slot-exhaustion pool test slept a fixed 50 ms and assumed the OS had
scheduled all SLOT_COUNT holder threads into their dispatch bodies inside
that window. Under a contended box (the rest of the lib suite running in
parallel) some holders had not yet claimed a slot, so the probe found a
free slot, succeeded, and the "declines instead of blocking" assertion
flaked. Replace the sleep with an explicit wait on a per-holder `held`
counter so the probe only runs once every slot is provably occupied,
bounded by a 10 s deadline so a starved runner fails loudly instead of
hanging. Verified green under 20 background CPU spinners.

Also renumber the task-runtime doc subsections that were still labelled
31.7/31.8/31.9 -- stale from before the section became 33 -- which
collided with the genuine section 31.7/31.8. They are now 33.7/33.8/33.9.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the squad/sebastian-cpu-task-runtime-transforms branch from 8fbb204 to 064afde Compare August 18, 2026 16:29
@justinchuby

Copy link
Copy Markdown
Owner Author

Validation (reviewer)

Reviewed, rebased onto current main, and validated on Intel Core i7-13800H (14 physical / 20 logical, 6 P + 8 E cores).

  • Rebase. Clean onto post-perf(cpu): route RoPE and Softmax through the CPU task runtime #1202 main — perf(cpu): route Transpose and the elementwise fallback through the task runtime #1207 touches the elementwise dispatch wrappers (run_on_task_runtime, the in_task() guard), not map_ps, so it did not collide with the concurrent activation-kernel cluster.
  • Bit-identity. Pure scheduling; transpose (56/56) and elementwise unit tests (scalar-reference comparisons) all pass → output bit-identical to base. Full lib suite: 1405 passed / 0 failed / 17 ignored, no regression vs the rebase-point baseline.
  • Oversubscription. The fallback is guarded by rayon::current_thread_index().is_some() || task_runtime::in_task(), so a fan-out already inside a pool worker runs serially rather than nesting a second one. It reuses perf(cpu): add a CPU task runtime with an adaptive-spin native pool #1201's single pool; no new steady-state threads.
  • Contended-CPU. Pool tests run under 20 background CPU spinners (all logical cores saturated): 13 passed / 0 failed. No deadlock or starvation.
  • Clippy --all-targets -- -D warnings: clean.
  • Fixed forward: the slot_exhaustion_declines_instead_of_blocking pool test was timing-flaky under suite-level contention; rewrote it to construct the "all slots held" precondition deterministically. Also renumbered stale 31.7/31.8/31.9 doc subsections to 33.x (they collided with the genuine §31).
  • Perf: §33.8 tables are the author's native-ms A/B on this host class with symmetric untouched controls; I verified the mechanism and correctness but did not independently re-run the full 15-trial ORT sweep. Transpose is bandwidth-bound (small win at the wide end); the activation fallback shows the larger wins at t=16/32.

Merging on local correctness + stress evidence.

@justinchuby
justinchuby merged commit 262136e into main Aug 18, 2026
8 of 17 checks passed
@justinchuby
justinchuby deleted the squad/sebastian-cpu-task-runtime-transforms branch August 18, 2026 16:31
@codecov

codecov Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.08%. Comparing base (508de85) to head (064afde).
⚠️ Report is 35 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1207      +/-   ##
==========================================
+ Coverage   79.82%   80.08%   +0.26%     
==========================================
  Files         361      363       +2     
  Lines      156651   160409    +3758     
  Branches   156651   160409    +3758     
==========================================
+ Hits       125046   128469    +3423     
- Misses      26994    27287     +293     
- Partials     4611     4653      +42     
Flag Coverage Δ
mlas 85.09% <ø> (?)
offline 79.99% <100.00%> (+0.16%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...nnx-runtime-ep-cpu/src/kernels/simd_activations.rs 90.05% <100.00%> (+0.35%) ⬆️
...rates/onnx-runtime-ep-cpu/src/kernels/transpose.rs 96.04% <100.00%> (-0.20%) ⬇️
...rates/onnx-runtime-ep-cpu/src/task_runtime/pool.rs 96.10% <100.00%> (+0.10%) ⬆️

... and 21 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

justinchuby added a commit that referenced this pull request Aug 19, 2026
…1232)

## The CPU budget was buying hyperthreads, not cores

`ONNX_GENAI_CPU_DECODE_THREADS=N` confines the process to `N` logical
CPUs.
`choose_budget_cpus` picked them as the `N` lowest indices of the chosen
NUMA
node. Every host we run on numbers SMT siblings adjacently — `0-1`,
`2-3`, … on
AMD EPYC and on Intel since Skylake-SP — so **a budget of `N` landed on
`N/2`
physical cores**.

The symptom was unmistakable once measured. Int4 `MatMulNBits`
4096×6144, 128
tokens, native-only, 20 runs, on the 16-core/32-thread EPYC 9V74:

| budget | wall per run | user CPU |
| ------ | ------------ | -------- |
| 1 | 79.4 ms | 2.04 s |
| 2 | 81.2 ms | **4.10 s** |
| 3 | 40.4 ms | 2.05 s |
| 4 | 53.9 ms | 3.35 s |

A budget of 2 was exactly as slow as a budget of 1 while burning two
cores. A
budget of 3 beat a budget of 4. `taskset` confirms the cause directly —
`0,2`
(two cores) ran in 50.1 ms against `0,1` (one core, two threads) at 81.8
ms,
using 36% less CPU.

## The fix

A **ranking, not a widening**:

* `scatter_across_cores` orders the candidate pool by
`CoreTopology::leaders_within` — one CPU per physical core first,
siblings
after — and truncates *after* ranking. The result stays a subset of the
same
pool, so the cpuset/cgroup guarantee is unchanged and a full-width
budget
  returns exactly what it did before.
* `smt_scaled_request` sizes the NUMA node search in cores rather than
logical
CPUs, falling back to the old logical sizing when no node is that large,
so a
  budget that fits on one core-rich node still stays on one node.
* the cross-node top-up now happens *before* ranking, so a second node's
fresh
cores can displace the first node's SMT siblings instead of being
appended to
  an already-full mask.
* `order_pin_targets` applies the same order to the flat decode Rayon
pool's pin
list (the builder pins worker `i` to `cpus[i % len]`, so the order
decides
whether an 8-worker pool occupies 8 cores or 4). The persistent SPMD
pool and
the `numa-split` sub-pools are deliberately left compact — their workers
spin,
and spreading spinning workers one per core is the experiment already
recorded
  in `core_topology`'s module docs, where it measured *worse*.

`choose_budget_cpus` gained a `cores: Option<&CoreTopology>` parameter
so the
policy stays pure and unit-testable; with `None` (no discoverable SMT
map) every
function here is the identity and the old behaviour stands exactly.

## Result

42 cells (20 GEMM, 22 transform) × 6 widths × 5 trials, paired arms in
one
driver invocation. Native-only view — the mask moves ORT's threads too,
so only
native-vs-native isolates the change. Geomean of `before/after`, >1 is
faster:

| width | GEMM | >1.05× | <0.95× | Transforms | >1.05× | <0.95× |
| ----- | ---- | ------ | ------ | ---------- | ------ | ------ |
| 1  | 0.996 | 0 | 1 | 0.997 | 1 | 3 |
| 2  | **1.769** | 19 | **0** | 1.191 | 12 | **0** |
| 4  | **1.642** | 20 | **0** | 1.095 | 11 | 2 |
| 8  | **1.441** | 16 | 1 | 0.913 | 4 | 9 |
| 16 | **1.245** | 15 | 1 | 0.961 | 6 | 9 |
| 32 | 0.976 | 8 | 6 | 1.062 | 9 | 4 |

The two ends are the control: a budget of 1 cannot be re-ranked and a
budget of
32 is the whole machine, so both must be flat — and both are. At 2 and 4
threads
**not one cell out of 42 regressed**.

The transform column at t=8/16 reads as a loss and is not one: those
cells are
20–500 µs at five trials on a shared box. Re-measured at 40 runs × 5
warmups
with the arms strictly interleaved, they are wins — `sm_bert_b8_s128` by
1.9× at
both widths, `rope_gptj_il_s512` by 1.35× at t=8. The five-trial grid is
left in
the docs unedited so the distinction stays visible; §34.5 has the raw
pairs.

## Validation

* `cargo test -p onnx-runtime-ep-cpu` — 1391 passed, 0 failed
* `cargo clippy -p onnx-runtime-ep-cpu --all-targets` — clean
* `cargo fmt --all --check` — clean
* 8 new unit tests, all of which fail if the ranking is reverted,
including the
cpuset-containment case (a leader that exists in the topology but not in
the
process's allowed set must not be invented), the asymmetric-node case
that
  exercises `smt_scaled_request` through `choose_budget_cpus`, the
top-up-before-ranking case, and the "a budget that fits one node does
not
  spill across nodes for cores" case.

Documented as **§34 Phase 14** in
`docs/benchmarks/2026-08-15-cpu-ep-vs-ort-attention-moe.md`.

## Rebase and convergence (2026-08-18)

#1201/#1202/#1207 were **squash-merged**, so this branch's ancestry no
longer
reached `main`. Its three commits were cherry-picked onto `main` at
`c55a3fab3`;
all applied cleanly, and the diff is now self-contained: 3 files,
+465/−29.
Review the whole PR, not "the last two commits".

Defect found and fixed during the rebase: the benchmark section was
numbered
`## 32. Phase 12`, which **collided with the existing §32 (Phase 12, erf
Estrin)** on `main`. Renumbered to `## 34. Phase 14` (subsections
34.1–34.6),
and the internal back-reference "Phase 11 built a task runtime"
corrected to
"Phase 13 (§33) built a task runtime" — §33 is where the task runtime
actually
landed. The squad decision note's `§32` cross-reference was updated to
match.

Re-validated on the rebased head (rustc 1.97.1, the same toolchain CI
resolves):

* `cargo test -p onnx-runtime-ep-cpu --lib` — **1426 passed, 0 failed,
17 ignored**
* `cargo clippy -p onnx-runtime-ep-cpu --all-targets` — clean
* `cargo fmt --all -- --check` — clean

CI note: the repository's Actions queue is saturated (every recent run
is
`queued`, nothing `in_progress`), so the two required checks — `Fast
(Linux
x86_64)` and `Rust quality` — cannot report. Both lanes were reproduced
locally,
step for step, from `.github/workflows/ci.yml`. `main` itself fails
**both** of
them today; #1346 is the fix for that and lands first.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
perf(cpu): run the MLAS prefill tiling on the CPU task runtime

The `m > 1` MLAS SQNBit prefill tiling (`run_mlas_shards`) issued a
single
`tiles.par_iter()` fan-out on global Rayon, sized by
`rayon::current_num_threads()`. With a co-resident ORT intra-op pool
spinning on the same cores, every parked-Rayon wake-up lands behind a
spinning thread, which is what made `gemm_nbits_*_t8` at t=32 the worst
cell in the benchmark ledger. This routes that fan-out through the CPU
task runtime from #1201 instead, with a work-size policy for the one
case
where the SMT-capped pool leaves hardware threads idle.

## MLAS remains opt-in and non-load-bearing

This does **not** enable MLAS anywhere. The changed code lives entirely
inside the pre-existing `m > 1 && active > 1 && !mlas_prefill_serial()`
path that already called `mlas_sys::sqnbit_gemm_into`. No Cargo feature,
`#[cfg(feature = "mlas")]` gate, or default-feature set is touched:
`mlas`
is still opt-in (`default = ["full"]`, and `full` does not include
`mlas`). The new policy helpers (`prefill_fan_out`,
`prefill_tile_grain`,
`PrefillFanOut`) are pure integer arithmetic marked
`#[cfg_attr(not(feature = "mlas"), allow(dead_code))]` so they compile
and
their unit tests run on the default (mlas-off) CI lane even though only
the MLAS path consults them. MLAS stays a labelled reference arm.

## What changed

- `prefill_fan_out(macs, lanes, wide)`: below `WIDE_PREFILL_MACS` (512
Mi
  MACs), or whenever global Rayon is not actually wider than the pool,
  fan out on the task runtime (cheap ~5 us dispatch, topology-aware,
  no fight with a co-resident ORT pool). Above it, use the wider global
  Rayon path -- a prefill tile is a multi-ms MLAS call whose dequantise
  step has enough load latency that SMT siblings pay off, so the SMT cap
  costs more than a 226 us park wake-up (0.25% of a 90 ms fan-out).
- `prefill_tile_grain`: a per-task tile floor so no task gets less than
  `MIN_PREFILL_TASK_MACS` (512 Ki) of arithmetic.

## Verification

Hardware: Intel Core i7-13800H, 14 physical / 20 logical (6 P + 8 E).
Baseline: `main` at the rebase point. Toolchain: cargo 1.97.1.

- **Bit-identity (the important one).** This is pure scheduling: the
`run_tile` closure and the `sqnbit_gemm_into` call are byte-for-byte the
  same, only the executor and grain differ, and every tile writes a
  disjoint `[row, row+rows) x [shard.start, shard.start+len)` window so
  order cannot matter. Confirmed empirically under `--features mlas`:
  `mlas_prefill_parallel_dispatch_matches_serial` and
  `mlas_prefill_dispatch_parity_subprocess` pass -- the routed parallel
  tiling matches the serial reference.
- **Policy tests (default features).** All six `prefill_*` unit tests
pass.
- **Falsified.** Flipping the threshold comparison in `prefill_fan_out`
  from `<` to `<=` turns `large_prefill_work_takes_the_wide_fan_out` RED
  (`left: TaskRuntime, right: Wide` at exactly `WIDE_PREFILL_MACS`) --
  the test is non-vacuous and guards the boundary. Restored to green.
- **`--features mlas` compiles clean;** clippy `--all-targets -D
warnings`
  clean on default features.

## Perf

The §35 (Phase 15) tables in the benchmark doc record up to 13.5x at
t=32
on the small int4 cells, dropping `gemm_nbits_*_t8` from 22-34x ORT to
2.6-2.8x. Those tables mix the author's EPYC 9V74 (16c/32t) and the
laptop measurements; per the repo's measurement rule they are peers,
named
by hardware. The mechanism (the work-size policy, the grain floor, the
disjoint-window safety) and correctness are verified, and the policy
decisions are reproduced in unit tests; no full ORT A/B sweep was re-run
on the laptop, so the headline speedup magnitudes are the author's, not
independently re-measured there.

## Rebase and convergence (2026-08-18)

#1201/#1202/#1207 and #1143 were **squash-merged**, so this branch's
ancestry
no longer reached `main`. Its three commits were cherry-picked onto
`main` at
`c55a3fab3` and applied cleanly. The diff is now self-contained: 2
files,
+355/-25 (`matmul_nbits.rs` and the benchmark doc). It is **not**
stacked on
#1232 any more.

Section numbering: #1232 lands first and takes §34/Phase 14, so this
PR's
section was renumbered to **§35/Phase 15**. (The `34.1×` figure in the
§35.4
matrix is a speed ratio, not a section reference, and is unchanged.)

Re-validated on the rebased head (rustc 1.97.1, the toolchain CI
resolves):

* `cargo test -p onnx-runtime-ep-cpu --lib` — **1433 passed, 0 failed,
17 ignored** (on top of #1232)
* `cargo clippy -p onnx-runtime-ep-cpu --all-targets` — clean
* `cargo fmt --all -- --check` — clean
* under `--features mlas`, the parity test that actually guards this
change,
  `mlas_prefill_parallel_dispatch_matches_serial`, **passes**

### The two `--features mlas` failures are pre-existing on `main`

Running with `--features mlas` fails two tests:
`feature_default_guard::mlas_is_not_a_default_feature` (it panics
*because*
`--features mlas` was passed explicitly) and

`kernels::simd_activations::mlas_ab::mlas_matches_rust_simd_on_special_values`
(a 1-ULP Erf disagreement between the MLAS and pure-Rust SIMD routes).
**Both
reproduce identically on `main` at `c55a3fab3` with this PR's changes
absent**,
so they are baseline-equivalent, not regressions, and they are out of
scope
here. Neither runs on any required lane: `mlas` is not a default
feature.

### The red-criterion comment is runner noise

The criterion report flags regressions including
`tokenization/decode_tokens_per_second` -47.6%. This diff cannot reach
the
tokenizer, and every changed line of `matmul_nbits.rs` is inside the
`#[cfg(feature = "mlas")]` `run_mlas_shards` path, which the benchmark
build
(default features) does not compile in. The default-feature build is
behaviourally identical to `main`; the deltas are shared-runner
variance.

### CI

The repository's Actions queue is saturated (every recent run is
`queued`,
nothing `in_progress`), so the two required checks — `Fast (Linux
x86_64)` and
`Rust quality` — cannot report. Both lanes were reproduced locally, step
for
step, from `.github/workflows/ci.yml`. `main` itself fails **both** of
them
today; #1346 is the fix for that and lands first.

Closes #1238. Working as sebastian (CPU perf).

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants