Repository navigation
perf(cpu): route Transpose and the elementwise fallback through the task runtime - #1207
Conversation
3cf4f18 to
e1ca38a
Compare
f9cb1aa to
c6dc23c
Compare
e1ca38a to
53f362c
Compare
c6dc23c to
9356172
Compare
f4d2634 to
6c4d69d
Compare
3dc2157 to
1d12b3c
Compare
6c4d69d to
0d88434
Compare
1d12b3c to
5e48ff1
Compare
d025c32 to
abcbfd1
Compare
bc4646a to
29f7a67
Compare
abcbfd1 to
7769607
Compare
29f7a67 to
132d15a
Compare
7769607 to
3fedb98
Compare
132d15a to
42aaf72
Compare
3fedb98 to
0eb4888
Compare
19907d2 to
e52a6e7
Compare
a4395c9 to
bab6433
Compare
e52a6e7 to
906f201
Compare
bab6433 to
8fe0695
Compare
906f201 to
ebcc6cd
Compare
ebcc6cd to
b4caf90
Compare
26a673c to
cd6e3e1
Compare
b4caf90 to
8fbb204
Compare
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
## What `main` at 4d231ea fails two blocking CI steps on current stable (rustc/rustfmt 1.9.0, 1.97.1): ``` $ cargo fmt --all -- --check Diff in crates/onnx-runtime-ep-cuda/src/provider.rs Diff in crates/onnx-runtime-memory-api/src/allocator.rs Diff in crates/onnx-runtime-memory-api/src/capability.rs $ cargo clippy --locked --all-targets -p onnx-runtime-session -- -D warnings error: field `0` is never read --> crates/onnx-runtime-session/src/executor/mod.rs:175:45 ``` Both landed while the runner queue was saturated (71 queued / 1 in progress at the time of writing, nothing completed on `main` since 08:18Z), so no PR has seen a red check yet. Every open PR in the repo currently inherits both failures. ## Why these fixes **rustfmt** — mechanical normalisation, no semantic change. **`ActivationPlanForTest`** (from #1226) is a tuple struct whose single field is the `globals_lock()` `MutexGuard`. `dead_code` does not model "this field's value is its `Drop`", so it fires. The guard must stay: releasing it early is exactly the leaked-planner-gate race the struct was added to prevent. So the lint is silenced with a comment explaining the RAII intent, rather than the field removed. ## Verification | check | result | | --- | --- | | `cargo fmt --all -- --check` | clean | | `cargo clippy --locked --all-targets -p onnx-runtime-session -- -D warnings` | clean | | `cargo test --locked -p onnx-runtime-session --lib` | 181 passed, 0 failed | Found while trying to land the CPU task-runtime stack (#1201 → #1202 → #1207 → #1232 → #1238); this is unrelated to that work and is deliberately kept out of it so it can go in on its own. Working as sebastian (Performance Engineer) --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
cd6e3e1 to
28c75ea
Compare
…untime Two more callers off Rayon and onto the CPU task runtime. Transpose split its output into `units / workers` bands -- one static share per worker, so the whole copy waited for whichever band was unlucky. It now hands the runtime chunks of at least MIN_PARALLEL_TASK_BYTES (64 KiB, a quarter of the whole-operation bar, still comfortably bandwidth-bound) and lets the runtime claim them dynamically and cap the task count. All three split sites -- `block_move`, `blocked_2d`, `batched_blocked_2d` -- use `chunk_runs_mut`, so a task still gets one contiguous band in one call and none of the three bodies changes shape. `simd_activations` already borrowed ORT's pool inside the plugin EP. Its *other* arm -- a native session, or an ORT session whose intra-op pool is one thread, where `prefer_host` correctly declines -- still went to Rayon and still paid the park latency this whole exercise is about: 67us for an isolated region back-to-back, 226us when the previous region ended 20us earlier, which is exactly the spacing between two activations in a decode step. That arm now goes to the native pool. The new tail is outlined in `run_on_task_runtime` rather than written inline, for the reason recorded on `try_host`: growing `run_chunked` repartitions codegen units, and did cost `Relu` 34% at 1 Mi once already with no path change at all. The nesting guards now test `task_runtime::in_task()` as well as `rayon::current_thread_index()`, because both pools can be the outer region while the crate is mid-migration. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The elementwise ops are the ones most exposed to scheduling overhead -- a Gelu over an MLP intermediate is a few hundred microseconds of arithmetic wrapped around a fan-out -- so they are the sharpest test of a task runtime, and the harness had no cells for them. The sweep is by token count at fixed hidden/intermediate widths, which walks the tensor from L2-resident through L3 to memory-resident while holding the fan-out cost fixed. That is the axis that decides whether parallelising is worth it at all, so it is the axis a scheduler change has to be measured along. FastGelu rather than Gelu because Gelu only joins the default domain at opset 20, and FastGelu is what a real transformer graph carries here. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ls, and the GEMM finding Adds scripts/ort_ab/gen_gemm.py, which generates the MatMulNBits and dense MatMul cells the control grid needed. That control was meant to be a null result and instead found the largest remaining loss in the CPU EP: decode-shaped quantised GEMM is 4.6x slower at 32 threads than at 16. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Pure rustfmt output, no behaviour change. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…oc numbering The slot-exhaustion pool test slept a fixed 50 ms and assumed the OS had scheduled all SLOT_COUNT holder threads into their dispatch bodies inside that window. Under a contended box (the rest of the lib suite running in parallel) some holders had not yet claimed a slot, so the probe found a free slot, succeeded, and the "declines instead of blocking" assertion flaked. Replace the sleep with an explicit wait on a per-holder `held` counter so the probe only runs once every slot is provably occupied, bounded by a 10 s deadline so a starved runner fails loudly instead of hanging. Verified green under 20 background CPU spinners. Also renumber the task-runtime doc subsections that were still labelled 31.7/31.8/31.9 -- stale from before the section became 33 -- which collided with the genuine section 31.7/31.8. They are now 33.7/33.8/33.9. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
8fbb204 to
064afde
Compare
Validation (reviewer)Reviewed, rebased onto current
Merging on local correctness + stress evidence. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1207 +/- ##
==========================================
+ Coverage 79.82% 80.08% +0.26%
==========================================
Files 361 363 +2
Lines 156651 160409 +3758
Branches 156651 160409 +3758
==========================================
+ Hits 125046 128469 +3423
- Misses 26994 27287 +293
- Partials 4611 4653 +42
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
…1232) ## The CPU budget was buying hyperthreads, not cores `ONNX_GENAI_CPU_DECODE_THREADS=N` confines the process to `N` logical CPUs. `choose_budget_cpus` picked them as the `N` lowest indices of the chosen NUMA node. Every host we run on numbers SMT siblings adjacently — `0-1`, `2-3`, … on AMD EPYC and on Intel since Skylake-SP — so **a budget of `N` landed on `N/2` physical cores**. The symptom was unmistakable once measured. Int4 `MatMulNBits` 4096×6144, 128 tokens, native-only, 20 runs, on the 16-core/32-thread EPYC 9V74: | budget | wall per run | user CPU | | ------ | ------------ | -------- | | 1 | 79.4 ms | 2.04 s | | 2 | 81.2 ms | **4.10 s** | | 3 | 40.4 ms | 2.05 s | | 4 | 53.9 ms | 3.35 s | A budget of 2 was exactly as slow as a budget of 1 while burning two cores. A budget of 3 beat a budget of 4. `taskset` confirms the cause directly — `0,2` (two cores) ran in 50.1 ms against `0,1` (one core, two threads) at 81.8 ms, using 36% less CPU. ## The fix A **ranking, not a widening**: * `scatter_across_cores` orders the candidate pool by `CoreTopology::leaders_within` — one CPU per physical core first, siblings after — and truncates *after* ranking. The result stays a subset of the same pool, so the cpuset/cgroup guarantee is unchanged and a full-width budget returns exactly what it did before. * `smt_scaled_request` sizes the NUMA node search in cores rather than logical CPUs, falling back to the old logical sizing when no node is that large, so a budget that fits on one core-rich node still stays on one node. * the cross-node top-up now happens *before* ranking, so a second node's fresh cores can displace the first node's SMT siblings instead of being appended to an already-full mask. * `order_pin_targets` applies the same order to the flat decode Rayon pool's pin list (the builder pins worker `i` to `cpus[i % len]`, so the order decides whether an 8-worker pool occupies 8 cores or 4). The persistent SPMD pool and the `numa-split` sub-pools are deliberately left compact — their workers spin, and spreading spinning workers one per core is the experiment already recorded in `core_topology`'s module docs, where it measured *worse*. `choose_budget_cpus` gained a `cores: Option<&CoreTopology>` parameter so the policy stays pure and unit-testable; with `None` (no discoverable SMT map) every function here is the identity and the old behaviour stands exactly. ## Result 42 cells (20 GEMM, 22 transform) × 6 widths × 5 trials, paired arms in one driver invocation. Native-only view — the mask moves ORT's threads too, so only native-vs-native isolates the change. Geomean of `before/after`, >1 is faster: | width | GEMM | >1.05× | <0.95× | Transforms | >1.05× | <0.95× | | ----- | ---- | ------ | ------ | ---------- | ------ | ------ | | 1 | 0.996 | 0 | 1 | 0.997 | 1 | 3 | | 2 | **1.769** | 19 | **0** | 1.191 | 12 | **0** | | 4 | **1.642** | 20 | **0** | 1.095 | 11 | 2 | | 8 | **1.441** | 16 | 1 | 0.913 | 4 | 9 | | 16 | **1.245** | 15 | 1 | 0.961 | 6 | 9 | | 32 | 0.976 | 8 | 6 | 1.062 | 9 | 4 | The two ends are the control: a budget of 1 cannot be re-ranked and a budget of 32 is the whole machine, so both must be flat — and both are. At 2 and 4 threads **not one cell out of 42 regressed**. The transform column at t=8/16 reads as a loss and is not one: those cells are 20–500 µs at five trials on a shared box. Re-measured at 40 runs × 5 warmups with the arms strictly interleaved, they are wins — `sm_bert_b8_s128` by 1.9× at both widths, `rope_gptj_il_s512` by 1.35× at t=8. The five-trial grid is left in the docs unedited so the distinction stays visible; §34.5 has the raw pairs. ## Validation * `cargo test -p onnx-runtime-ep-cpu` — 1391 passed, 0 failed * `cargo clippy -p onnx-runtime-ep-cpu --all-targets` — clean * `cargo fmt --all --check` — clean * 8 new unit tests, all of which fail if the ranking is reverted, including the cpuset-containment case (a leader that exists in the topology but not in the process's allowed set must not be invented), the asymmetric-node case that exercises `smt_scaled_request` through `choose_budget_cpus`, the top-up-before-ranking case, and the "a budget that fits one node does not spill across nodes for cores" case. Documented as **§34 Phase 14** in `docs/benchmarks/2026-08-15-cpu-ep-vs-ort-attention-moe.md`. ## Rebase and convergence (2026-08-18) #1201/#1202/#1207 were **squash-merged**, so this branch's ancestry no longer reached `main`. Its three commits were cherry-picked onto `main` at `c55a3fab3`; all applied cleanly, and the diff is now self-contained: 3 files, +465/−29. Review the whole PR, not "the last two commits". Defect found and fixed during the rebase: the benchmark section was numbered `## 32. Phase 12`, which **collided with the existing §32 (Phase 12, erf Estrin)** on `main`. Renumbered to `## 34. Phase 14` (subsections 34.1–34.6), and the internal back-reference "Phase 11 built a task runtime" corrected to "Phase 13 (§33) built a task runtime" — §33 is where the task runtime actually landed. The squad decision note's `§32` cross-reference was updated to match. Re-validated on the rebased head (rustc 1.97.1, the same toolchain CI resolves): * `cargo test -p onnx-runtime-ep-cpu --lib` — **1426 passed, 0 failed, 17 ignored** * `cargo clippy -p onnx-runtime-ep-cpu --all-targets` — clean * `cargo fmt --all -- --check` — clean CI note: the repository's Actions queue is saturated (every recent run is `queued`, nothing `in_progress`), so the two required checks — `Fast (Linux x86_64)` and `Rust quality` — cannot report. Both lanes were reproduced locally, step for step, from `.github/workflows/ci.yml`. `main` itself fails **both** of them today; #1346 is the fix for that and lands first. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
perf(cpu): run the MLAS prefill tiling on the CPU task runtime The `m > 1` MLAS SQNBit prefill tiling (`run_mlas_shards`) issued a single `tiles.par_iter()` fan-out on global Rayon, sized by `rayon::current_num_threads()`. With a co-resident ORT intra-op pool spinning on the same cores, every parked-Rayon wake-up lands behind a spinning thread, which is what made `gemm_nbits_*_t8` at t=32 the worst cell in the benchmark ledger. This routes that fan-out through the CPU task runtime from #1201 instead, with a work-size policy for the one case where the SMT-capped pool leaves hardware threads idle. ## MLAS remains opt-in and non-load-bearing This does **not** enable MLAS anywhere. The changed code lives entirely inside the pre-existing `m > 1 && active > 1 && !mlas_prefill_serial()` path that already called `mlas_sys::sqnbit_gemm_into`. No Cargo feature, `#[cfg(feature = "mlas")]` gate, or default-feature set is touched: `mlas` is still opt-in (`default = ["full"]`, and `full` does not include `mlas`). The new policy helpers (`prefill_fan_out`, `prefill_tile_grain`, `PrefillFanOut`) are pure integer arithmetic marked `#[cfg_attr(not(feature = "mlas"), allow(dead_code))]` so they compile and their unit tests run on the default (mlas-off) CI lane even though only the MLAS path consults them. MLAS stays a labelled reference arm. ## What changed - `prefill_fan_out(macs, lanes, wide)`: below `WIDE_PREFILL_MACS` (512 Mi MACs), or whenever global Rayon is not actually wider than the pool, fan out on the task runtime (cheap ~5 us dispatch, topology-aware, no fight with a co-resident ORT pool). Above it, use the wider global Rayon path -- a prefill tile is a multi-ms MLAS call whose dequantise step has enough load latency that SMT siblings pay off, so the SMT cap costs more than a 226 us park wake-up (0.25% of a 90 ms fan-out). - `prefill_tile_grain`: a per-task tile floor so no task gets less than `MIN_PREFILL_TASK_MACS` (512 Ki) of arithmetic. ## Verification Hardware: Intel Core i7-13800H, 14 physical / 20 logical (6 P + 8 E). Baseline: `main` at the rebase point. Toolchain: cargo 1.97.1. - **Bit-identity (the important one).** This is pure scheduling: the `run_tile` closure and the `sqnbit_gemm_into` call are byte-for-byte the same, only the executor and grain differ, and every tile writes a disjoint `[row, row+rows) x [shard.start, shard.start+len)` window so order cannot matter. Confirmed empirically under `--features mlas`: `mlas_prefill_parallel_dispatch_matches_serial` and `mlas_prefill_dispatch_parity_subprocess` pass -- the routed parallel tiling matches the serial reference. - **Policy tests (default features).** All six `prefill_*` unit tests pass. - **Falsified.** Flipping the threshold comparison in `prefill_fan_out` from `<` to `<=` turns `large_prefill_work_takes_the_wide_fan_out` RED (`left: TaskRuntime, right: Wide` at exactly `WIDE_PREFILL_MACS`) -- the test is non-vacuous and guards the boundary. Restored to green. - **`--features mlas` compiles clean;** clippy `--all-targets -D warnings` clean on default features. ## Perf The §35 (Phase 15) tables in the benchmark doc record up to 13.5x at t=32 on the small int4 cells, dropping `gemm_nbits_*_t8` from 22-34x ORT to 2.6-2.8x. Those tables mix the author's EPYC 9V74 (16c/32t) and the laptop measurements; per the repo's measurement rule they are peers, named by hardware. The mechanism (the work-size policy, the grain floor, the disjoint-window safety) and correctness are verified, and the policy decisions are reproduced in unit tests; no full ORT A/B sweep was re-run on the laptop, so the headline speedup magnitudes are the author's, not independently re-measured there. ## Rebase and convergence (2026-08-18) #1201/#1202/#1207 and #1143 were **squash-merged**, so this branch's ancestry no longer reached `main`. Its three commits were cherry-picked onto `main` at `c55a3fab3` and applied cleanly. The diff is now self-contained: 2 files, +355/-25 (`matmul_nbits.rs` and the benchmark doc). It is **not** stacked on #1232 any more. Section numbering: #1232 lands first and takes §34/Phase 14, so this PR's section was renumbered to **§35/Phase 15**. (The `34.1×` figure in the §35.4 matrix is a speed ratio, not a section reference, and is unchanged.) Re-validated on the rebased head (rustc 1.97.1, the toolchain CI resolves): * `cargo test -p onnx-runtime-ep-cpu --lib` — **1433 passed, 0 failed, 17 ignored** (on top of #1232) * `cargo clippy -p onnx-runtime-ep-cpu --all-targets` — clean * `cargo fmt --all -- --check` — clean * under `--features mlas`, the parity test that actually guards this change, `mlas_prefill_parallel_dispatch_matches_serial`, **passes** ### The two `--features mlas` failures are pre-existing on `main` Running with `--features mlas` fails two tests: `feature_default_guard::mlas_is_not_a_default_feature` (it panics *because* `--features mlas` was passed explicitly) and `kernels::simd_activations::mlas_ab::mlas_matches_rust_simd_on_special_values` (a 1-ULP Erf disagreement between the MLAS and pure-Rust SIMD routes). **Both reproduce identically on `main` at `c55a3fab3` with this PR's changes absent**, so they are baseline-equivalent, not regressions, and they are out of scope here. Neither runs on any required lane: `mlas` is not a default feature. ### The red-criterion comment is runner noise The criterion report flags regressions including `tokenization/decode_tokens_per_second` -47.6%. This diff cannot reach the tokenizer, and every changed line of `matmul_nbits.rs` is inside the `#[cfg(feature = "mlas")]` `run_mlas_shards` path, which the benchmark build (default features) does not compile in. The default-feature build is behaviourally identical to `main`; the deltas are shared-runner variance. ### CI The repository's Actions queue is saturated (every recent run is `queued`, nothing `in_progress`), so the two required checks — `Fast (Linux x86_64)` and `Rust quality` — cannot report. Both lanes were reproduced locally, step for step, from `.github/workflows/ci.yml`. `main` itself fails **both** of them today; #1346 is the fix for that and lands first. Closes #1238. Working as sebastian (CPU perf). --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
perf(cpu): route Transpose and the elementwise fallback through the task runtime
Moves the transpose blocked-copy paths (
block_move,blocked_2d,batched_blocked_2d) and the SIMD elementwise-activation fallback offtheir ad-hoc rayon fan-out and onto the CPU task runtime landed in #1201,
so they share one pool with RoPE/Softmax instead of spinning up a second
scheduler.
What changed
MIN_PARALLEL_TASK_BYTES(64 KiB) balanceunit and fans out through
task_runtime::chunk_runs_mut, with thetask count derived from
task_runtime::width().run_on_task_runtime, guardedby
rayon::current_thread_index().is_some() || task_runtime::in_task()so a fan-out already inside a pool worker runs serially instead of
nesting a second fan-out (this is the oversubscription guard).
Verification (this machine)
Hardware: Intel Core i7-13800H, 14 physical / 20 logical cores
(6 P-cores + 8 E-cores). Baseline:
mainat the rebase point.Toolchain: cargo 1.97.1, release for any timing.
reordering parallel work must not change per-element math. The
transpose and elementwise unit tests (which compare against a scalar
reference) all pass, so output is bit-identical to base.
cargo test -p onnx-runtime-ep-cpu --lib= 1405 passed, 0 failed, 17 ignored(baseline on the rebase point was the same count; no test regressed).
task_runtime::chunk_runs_mut; the elementwise fallback reachesrun_on_task_runtime. The only remaining rayon intranspose.rsisthe
#[cfg(test)]bench harness.with 20 background CPU spinners saturating every logical core: 13
passed, 0 failed. See the fixed slot-exhaustion test below.
--all-targets -- -D warnings: clean.Flaky test fixed forward
task_runtime::pool::tests::slot_exhaustion_declines_instead_of_blockingslept a fixed 50 ms and assumed the OS scheduled all
SLOT_COUNTholderthreads into their dispatch bodies inside that window. Under the full lib
suite (cargo runs test binaries in parallel, saturating 20 logical CPUs)
some holders had not yet claimed a slot, the probe found a free slot, and
the "declines instead of blocking" assertion flaked. It now waits on a
per-holder
heldcounter until every slot is provably occupied (10 sdeadline so a starved runner fails loudly rather than hanging). Green
both in isolation and under 20-spinner background load.
Perf
The transpose / activation / control tables in
docs/benchmarks/2026-08-15-cpu-ep-vs-ort-attention-moe.md§33.8 are theauthor's native-ms A/B on this host class, with symmetric untouched
controls (GQA attention, GEMM) establishing the ±25% single-cell noise
floor. Transpose is bandwidth-bound and the gain is small and lives at
the wide end (t=16/32); the activation fallback shows up to ~3.6× at
t=16 and ~4× at t=32. I did not independently re-run the full 15-trial
ORT sweep; I verified the mechanism (width-based partitioning + the
nesting guard) and correctness. Doc subsections that were mislabelled
31.7/31.8/31.9 (stale from before the task-runtime section became 33)
are renumbered to 33.7/33.8/33.9 to stop colliding with the genuine §31.
Closes #1207. Working as sebastian (CPU perf).