Skip to content

perf(cpu): stop building the decode pool for decode-shaped borrowed int4 - #1434

Merged
justinchuby merged 5 commits into
mainfrom
squad/sebastian-lazy-decode-pool
Aug 19, 2026
Merged

justinchuby merged 5 commits into
mainfrom
squad/sebastian-lazy-decode-pool

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 19, 2026 •

Copy link
Copy Markdown
Owner

Building the decode pool is unconditional today: every borrowed-int4 call installs a 16-worker Rayon pool before dispatching, including the decode-shaped calls that then route all of their work to the task runtime and never touch it. Those workers are constructed, parked, and torn down once per call for nothing.

This defers the construction for exactly one path — decode-shaped (m == 1) borrowed int4 — and hands the routing logic the width it would have installed, so the executor choice and the partition grain are both unchanged.

What is actually covered

The scope is deliberately narrow, and the narrowness is what makes it provable rather than measured:

  • Three call sites are converted: packed_nbits_gemv, gemv_nk, and the borrowed-int4 site gated on m == 1.
  • Those paths reach Rayon only through parallel_output_rows_repeated (parallel_output_rows delegates to it), which is the hookable dispatcher.
  • The m == 1 gate is what makes this safe by construction, not by inspection: borrowed_affine_int4_matmul_prefill drives Rayon directly, and it is unreachable at m == 1. The other seven with_decode_pool sites are untouched, so the kernels that use Rayon directly — packed_nbits_gemm, int8_matmul, int8_row, parallel_n16_output_rows, parallel_kai_output_rows — keep eager installation.

Routing is preserved rather than assumed. flat_fan_out branches on rayon::current_num_threads(), which equalled the decode width only because we were installed; effective_fan_out_width() reproduces exactly that width while deferred. The second commit extends the same helper to output_chunk_len, so the grain cannot drift from the installed grain either — without it a deferred call partitioned into 4096-row chunks where the installed one used 64.

Five tests, two of them verified as falsifiers by deliberately breaking the thing they check:

  • relaxing the gate to m >= 1 makes only_the_decode_shaped_borrowed_int4_path_defers_the_pool fail;
  • dropping the grain fix makes a_deferred_fan_out_partitions_for_the_pool_it_would_have_installed fail (chunk 64 vs 4096).

Measured

16-core budget, interleaved with alternating arm order in a single session.

before after
process threads 48 32
onnx-genai-decode-* workers 16, 340-490 ms CPU none
voluntary ctxsw / iter 24.61 16.12
total CPU 3.660 cpu-s 3.580 cpu-s
dispatches / iter 1.30 1.08

Latency is neutral, and that is the claim — not an improvement. Six A/B reps gave 0.938, 1.180, 1.130, 1.115, 1.031, 1.026 (median 1.073), and an A/A null control measured in the same window gave a 0.83-1.21 band. Every ratio is inside the band, so this PR does not demonstrate a latency change in either direction. The win is structural: 16 fewer threads and a third fewer voluntary context switches.

One number deserves an explicit caveat: total CPU is flat, not lower. The decode pool's CPU does not disappear, it reappears on the caller thread. That is attribution changing, not work being removed — the work was always the caller's, it was just being done by borrowed workers.

onnx-runtime-ep-cpu lib suite: 1453 passed, 0 failed on merged latest main. fmt and clippy clean for this crate. (cargo fmt --check currently reports five sites repo-wide, all inherited from main and none in the one file this PR touches; #1393 repairs them.)

justinchuby and others added 2 commits August 19, 2026 07:47
The bounded, pinned decode Rayon pool is built eagerly on the first
`MatMulNBits` projection and then, for a pure-decode workload, never runs
any model arithmetic: `ONNX_GENAI_PROFILE_OPS` attributes 99.95% of the
forward pass to a single `MatMulNBits` per iteration, and the task runtime
takes 1.01 dispatches per iteration for it. The pool is a pass-through --
16 resident workers whose only cost is the per-dispatch install/wake.

That cost is not idle spin. Holding the iteration count fixed and
stretching the inter-token gap 100us -> 4000us grows wall time 2.9x while
the pool's CPU stays flat at ~350ms, so it is ~1.8ms of CPU per dispatch
paid to enter and leave a pool that computes nothing.

Defer building it. `with_decode_pool_lazy` runs the kernel inline while
publishing the width the pool *would* have had, and
`parallel_output_rows_repeated` installs the pool on demand if a fan-out
actually routes `Wide`. Routing is therefore unchanged: `wide` reads the
same value it read inside the installation, so no shape moves between the
task runtime and Rayon.

Scope is what makes this provable rather than hopeful. The deferral is
gated on `m == 1`, and at `m == 1` every kernel reachable from the
borrowed int4 closure parallelises solely through
`parallel_output_rows_repeated`. The one branch that drives
`par_chunks_mut` directly, `borrowed_affine_int4_matmul_prefill`, is gated
on `m >= 2` and so is unreachable; prefill keeps installing eagerly, as do
the seven other call sites whose kernels reach Rayon by other routes
(`packed_nbits_gemm`, `int8_matmul`, `parallel_n16_output_rows`,
`parallel_kai_output_rows`). A one-shot probe at every `with_decode_pool`
call site confirms this one is the only builder for a decode workload.

Measured at a 16-core budget, interleaved against the unmodified binary:
the process drops from 48 threads to 32, the 16 `onnx-genai-decode-*`
workers and their 340-490ms disappear entirely, and voluntary context
switches per iteration fall from ~18.7 to ~14.8. Latency is not claimed
here: the host was at load average 26 for these runs, well outside the
noise guard, so p50/p90 are deferred to a quiet window.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ust the route

Opus review caught that `output_chunk_len` reads
`rayon::current_num_threads()` directly, so it was the one input to the
fan-out that the deferral hint did not cover. The executor choice was
preserved but the *grain* was not: with no explicit budget the global pool
is `available_parallelism()` wide while the decode pool is narrower, so a
deferred fan-out would partition for the wrong pool and, at the boundary,
flip serial and parallel. Results stay correct either way -- row sharding
is associative -- but "routing is unchanged" was not literally true.

Route both the executor choice and the grain through one
`effective_fan_out_width()`, so a deferred fan-out reproduces the
installed case exactly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Opus review: no bugs, one nuance closed

Independent Opus review traced every kernel reachable from the three converted sites at m == 1, checked the guard/error propagation, and confirmed the aarch64 paths. Verdict: no correctness bugs, race conditions or reachability holes; routing class genuinely preserved; the m == 1 gate provably safe; the guard panic-safe.

It did find one real imprecision, now fixed in 635d93771:

output_chunk_len reads rayon::current_num_threads() directly — it was not switched to the deferred hint the way wide was. So while deferred, chunk reflects the ambient pool's width, not the decode pool's.

It rated this non-blocking (unreachable in the production scoped path, and never changes results because row sharding is associative), but it made the PR's own "routing is unchanged" claim not literally true, so I closed it rather than caveat it. Both the executor choice and the grain now go through one effective_fan_out_width(). New test a_deferred_fan_out_partitions_for_the_pool_it_would_have_installed; verified as a falsifier (reverting output_chunk_len makes it fail with chunk 64 instead of 4096).

Latency: measured on a quiet window, and it is neutral

The Draft deferred p50/p90 because the host was at load average 26. It quieted (cpu/wall back to 9.5–10.0, in band), so here is the interleaved A/B with alternating arm order, 400 iters, 16-core budget, plus an A/A null control run in the same window:

A/A null control (identical binaries, separate processes): p50 ratios 1.000, 0.828, 0.997 → null band 0.83–1.21.

A/B (gated / baseline), alternating order:

rep order baseline p50 gated p50 ratio
1 rt first 1.7315 1.6250 0.938
2 gated first 1.6028 1.8909 1.180
3 rt first 1.6029 1.8113 1.130
4 gated first 1.6477 1.8375 1.115
5 rt first 1.8067 1.8632 1.031
6 gated first 1.5936 1.6353 1.026

Median 1.073, every rep inside the null band. So: no detectable latency change in either direction. I am not claiming a speedup, and there is no regression.

The win is structural, and those metrics are load-insensitive. From a clean census pair in the same window (cpu/wall 9.56 vs 9.53):

metric baseline this PR
threads 48 32
onnx-genai-decode-* 16 workers, 350 ms none
voluntary ctxsw/iter 24.61 16.12
total CPU 3.660 cpu-s 3.580 cpu-s
RSS 213.7 MB 213.5 MB
dispatches/iter 1.30 1.08

Why the 350 ms does not reappear elsewhere

Worth stating explicitly, because the thread census looks like it moved: the caller group's CPU rises by roughly what the decode pool loses. That is attribution, not new work. pool.install(closure) used to run the kernel body on a decode worker; now it runs on the calling thread. Total process CPU goes down by ~2%, which is the install/wake crossing we stopped paying. The global Rayon pool is not picking this up — confirmed by converting all ten call sites in a scratch build, which produced the same 32 threads and the same total CPU.

Taking out of Draft: Opus green, full onnx-runtime-ep-cpu lib suite green on merged latest main (1453 passed), fmt and clippy clean. Auto-merge will be enabled and will wait for required CI — no bypass.

justinchuby added a commit that referenced this pull request Aug 19, 2026
## What

Unbreak the two CI lanes that are **red on `main` right now**. Three
independent breakages, all pre-existing and all reproduced on an
unmodified
`dbade34c1` checkout with the same stable 1.97.1 toolchain CI installs:

| # | gate | breakage | fix |
|---|------|----------|-----|
| 1 | `cargo fmt --all -- --check` | 5 sites / 4 files | rustfmt |
| 2 | `Rust quality` clippy | `clippy::unnecessary_map_or` —
`dispatch.rs:26` | `map_or(true, f)` → `is_none_or(f)` |
| 3 | `Fast (Linux x86_64)` clippy `--all-targets` |
`clippy::inconsistent_digit_grouping` — `cost-model/model.rs:314` |
`2_000_000_000_000_0` → `20_000_000_000_000` |

Both lanes build with `RUSTFLAGS: -D warnings`, so #2 and #3 are hard
errors,
not warnings. **Every PR that merges `main` inherits all three** —
verified on
#1434 and #1420. Nothing in the queue can go green until this lands.

#3 is worth calling out: it is invisible to a plain `cargo clippy`
because the
literal lives in a `#[cfg(test)]` module. Only the `--all-targets`
invocation
in the Fast lane sees it.

## Semantics

Both non-fmt changes are provably value-preserving:

- `is_none_or(f)` is the rewrite the lint itself suggests, and is
definitionally
  `map_or(true, f)`: `None` → `true`, `Some(v)` → `f(v)`.
`gqa_shape_capacity_bound_enabled()` is unchanged — unset stays enabled,
the
  falsey spellings stay disabled.
- `20_000_000_000_000 == 2_000_000_000_000_0` (both 2e13), which is what
the
  test's own comment already claims — *"2e13 FLOP / 2e13 = 1 s"*.
  `op_cost_takes_roofline_max` still asserts the compute term dominates.

## Validation

Ran locally per the delayed-Actions directive, on this head merged with
`origin/main` @ `dbade34c1`:

| gate | result |
|------|--------|
| `cargo fmt --all -- --check` | **0 diffs** |
| `cargo clippy --locked --all-targets $(workspace_test_packages.py
cargo-args offline-linux) -- -D warnings` | **exit 0** |
| `cargo test --locked $(… offline-linux)` | **3943 passed, 0 failed,
exit 0** |
| `scripts/check_cross_compile.sh` | **PASS** — x86_64 + aarch64 full
offline set |
| `benchmark_muse_native_local.py --self-test --require-numpy` | 43
cases passed |
| `check_publish_order.py` / `check_profile_table.py` /
`check_platform_naming.py` | PASS |
| `check_dispatch_reachability.py` / `check_feature_gate_coverage.py` |
PASS |
| `check_dispatch_manifest.py` (`--self-test` and plain) | PASS |
| `workspace_test_packages.py verify` | PASS |
| `verify_documented_env_vars.py` | PASS — 113 documented, 13
known-unimplemented |
| MLAS cfg: `-p onnx-runtime-ep-cpu --no-default-features --features
mlas` | `moe::` 19 passed · `qlinear_matmul::` 30 passed ·
`optimization_registry_excludes_nchwc_without_cnn_ops` 1 passed |
| `cargo clippy -p onnx-genai-engine --features native-backend` | clean
|
| `cargo build -p onnx-runtime-ep-cpu-plugin --features mlas` | clean |

### Windows ARM64: not validated locally — stated as a blocker, then
bounded

I could not run `Rust (Windows ARM64)` here and I am **not** claiming it
as a
pass. Two routes were attempted, both fail *identically on unmodified
`main`*,
so neither can discriminate this PR from baseline:

- **`cargo-xwin` / clang-cl** — installed, MSVC CRT + SDK downloaded,
correctly
targeting `aarch64-pc-windows-msvc`. Fails in vendored `mlasi.h` on NEON
intrinsics (`veorq_s32`, `vdupq_n_f32`, …) that MSVC supplies but
clang-cl in
  MSVC mode does not.
- **`aarch64-unknown-linux-gnu` + GNU cross toolchain** as an ARM64-NEON
proxy —
gets much further, compiles most of the ARM64 MLAS source set, then
fails on
  `activate_fp16.cpp`.

What makes this safe to merge anyway is **dependency-graph
disjointness**, not a
judgement call. That lane builds only `mlas-sys` and
`onnx-runtime-ep-cpu-plugin --features mlas`. This PR touches
`onnx-genai-engine` (tests), `onnx-runtime-ep-cuda` and
`onnx-runtime-session`:

```
$ cargo tree -p onnx-runtime-ep-cpu-plugin --features mlas -e normal --prefix none \
    | sort -u | grep -cE "onnx-runtime-ep-cuda|onnx-runtime-session|onnx-genai-engine"
0
```

Zero of the crates this PR modifies are in that lane's graph, so it
cannot
observe this change.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Local validation (delayed-Actions directive)

Head merged with origin/main @ 1557a355d.

The 48 -> 32 reduction, re-measured from scratch on latest main

Same worktree, same release binary recipe, same model, same host. The only
variable is this PR's single file (matmul_nbits.rs), swapped via
git checkout origin/main -- <file> and rebuilt, so the two arms differ by
nothing else. Peak /proc/<pid>/status:Threads sampled at 10 ms, with a
comm snapshot taken at the peak.

Model gemm_nbits_llama3_8b_mlp_t1.onnx (m = 1, the decode shape),
bench_generic --native-only --runs 400, pure native, no ORT, no MLAS:

ONNX_GENAI_CPU_DECODE_THREADS main this PR delta
4 9 9 —
8 17 17 —
16 48 32 −16

The comm attribution shows exactly what left. At 16 on main the peak
carries a onnx-genai-deco… cohort; at 16 on this PR it does not:

main   @16: 48 threads
PR     @16: 32 threads = 17 process/runtime + 15 nxrt-task-0 … nxrt-task-14
                         and zero onnx-genai-deco

So the decode pool is removed by construction on this path, not merely
idled — and the 16 saved threads are precisely the decode pool's width.

That the 4 and 8 rows are unchanged is the expected shape, not a miss: below
the routing width the fan-out is supposed to stay where it was, so there is no
pool to elide.

Scope honesty

The saving is specific to the covered path. On the same model with the decode
width left at its default the two arms are identical (39 threads, 6
onnx-genai-deco both sides) — the deferral only pays when the kernel does not
end up needing the fan-out it would have built the pool for. This PR claims no
more than that.

Behavioural falsifiers

cargo test -p onnx-runtime-ep-cpu --lib defer — 8 passed, 0 failed:

  • only_the_decode_shaped_borrowed_int4_path_defers_the_pool — the gate is narrow
  • a_deferred_fan_out_partitions_for_the_pool_it_would_have_installed — grain preserved
  • a_deferred_wide_fan_out_runs_on_the_decode_pool — wide work still gets the pool
  • a_deferred_decode_kernel_still_dispatches_to_the_task_runtime — routing intact
  • deferring_the_decode_pool_does_not_change_results — numerics unchanged
  • deferred_decode_width_nests_and_restores_on_panic — nesting + unwind safe
  • affinity_defer_routing_child, auto_default_with_explicit_affinity_defers_to_flat

Repository gates

gate result
cargo test --locked $(… offline-linux) 3996 passed, 0 failed, exit 0
cargo clippy --locked --all-targets $(… offline-linux) -- -D warnings exit 0
cargo fmt --all -- --check 0 diffs
scripts/check_cross_compile.sh PASS (x86_64 + aarch64 full offline set)
9 guard scripts PASS
benchmark_muse_native_local.py --self-test --require-numpy 43 cases
MLAS cfgs (moe:: / qlinear_matmul:: / registry) 19 / 30 / 1 passed
clippy -p onnx-genai-engine --features native-backend clean
build -p onnx-runtime-ep-cpu-plugin --features mlas clean

@justinchuby
justinchuby merged commit 6501bc6 into main Aug 19, 2026
6 checks passed
@justinchuby
justinchuby deleted the squad/sebastian-lazy-decode-pool branch August 19, 2026 17:39
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/medium_generic_f16_threads=8/32x512x512 30.17 µs 123.41 µs +309.0%
🔴 matmul/small_generic_f32_threads=8/1x256x256 33.10 µs 104.04 µs +214.3%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 40.41 µs 94.81 µs +134.6%
🔴 matmul/small_generic_f16_threads=8/1x256x256 31.17 µs 71.98 µs +130.9%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 31.87 µs 65.24 µs +104.7%
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 61.65 µs 105.93 µs +71.8%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 380.20 µs 638.39 µs +67.9%
🔴 gather/large_f32_threads=1-internal/131072 34.10 µs 50.47 µs +48.0%
⚠️ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 508.59 µs 638.70 µs +25.6%
⚠️ matmul/large_generic_bf16_threads=8/32x1024x1024 1.30 ms 1.61 ms +24.1%
⚠️ tokenization/encode_tokens_per_second 372.41 µs 458.46 µs +23.1%
⚠️ tokenization/decode_tokens_per_second 6.17 ms 7.50 ms +21.6%
⚠️ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 42.11 µs 51.19 µs +21.6%
⚠️ matmul/medium_generic_f32_threads=1/32x512x512 2.36 ms 2.77 ms +17.7%
⚠️ gather/large_bf16_threads=1-internal/131072 13.22 µs 15.42 µs +16.7%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.55 ms 4.09 ms +15.0%
✅ qwen3_sampling_processors/top_k_partial_selection 149.82 µs 166.96 µs +11.4%
✅ sampling_latency/min_p_per_token 217.14 µs 240.46 µs +10.7%
✅ add/small_f16_threads=1-internal/1024 495.7 ns 548.3 ns +10.6%
✅ sampling_latency/top_p_per_token 407.66 µs 447.84 µs +9.9%
✅ add/large_f32_threads=1-internal/4194304 633.44 µs 694.92 µs +9.7%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 541.95 µs 592.21 µs +9.3%
✅ matmul/small_generic_f32_threads=1/1x256x256 37.49 µs 40.72 µs +8.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.26 ms 2.41 ms +6.7%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.34 ms 9.87 ms +5.8%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 375.76 µs 396.07 µs +5.4%
✅ logit_processing/seven_processor_chain_per_step 354.18 µs 368.49 µs +4.0%
✅ add/large_bf16_threads=1-internal/4194304 1.61 ms 1.67 ms +3.7%
✅ matmul/medium_generic_f32_threads=8/32x512x512 924.97 µs 951.24 µs +2.8%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.99 ms 2.04 ms +2.4%
✅ gather/small_f32_threads=1-internal/4096 670.2 ns 684.8 ns +2.2%
✅ gather/large_f16_threads=1-internal/131072 13.42 µs 13.60 µs +1.3%
✅ reduce_mean/large_f32_threads=1-internal/262144 979.41 µs 991.88 µs +1.3%
✅ gather/medium_bf16_threads=1-internal/32768 2.39 µs 2.40 µs +0.5%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 576.99 µs 578.70 µs +0.3%
✅ matmul/small_generic_bf16_threads=1/1x256x256 31.74 µs 31.75 µs +0.0%
✅ gather/small_f16_threads=1-internal/4096 483.4 ns 482.4 ns -0.2%
✅ add/medium_bf16_threads=1-internal/262144 103.54 µs 103.26 µs -0.3%
✅ sampling_latency/top_k_per_token 58.25 µs 58.01 µs -0.4%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 89.09 µs 88.25 µs -0.9%
✅ add/small_bf16_threads=1-internal/1024 446.3 ns 441.2 ns -1.2%
✅ gather/medium_f16_threads=1-internal/32768 2.44 µs 2.41 µs -1.2%
✅ add/medium_f16_threads=1-internal/262144 105.61 µs 104.21 µs -1.3%
✅ add/medium_f32_threads=1-internal/262144 26.50 µs 26.14 µs -1.4%
✅ gather/medium_f32_threads=1-internal/32768 4.11 µs 4.05 µs -1.5%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 81.67 µs 79.99 µs -2.1%
✅ matmul/small_generic_f16_threads=1/1x256x256 30.45 µs 29.82 µs -2.1%
✅ add/large_f16_threads=1-internal/4194304 1.73 ms 1.70 ms -2.1%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.36 µs 14.96 µs -2.6%
✅ reduce_mean/medium_f32_threads=1-internal/65536 252.89 µs 242.30 µs -4.2%
✅ matmul/medium_generic_f16_threads=1/32x512x512 31.37 µs 29.79 µs -5.0%
✅ add/small_f32_threads=1-internal/1024 230.0 ns 216.3 ns -6.0%
✅ grammar_masking/llguidance_compute_mask/32 90.46 µs 84.86 µs -6.2%
✅ kv_cache/alloc_dealloc_pages 48.40 µs 44.62 µs -7.8%
✅ sampling_latency/greedy_per_token 4.02 µs 3.64 µs -9.5%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 7.39 ms 6.54 ms -11.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 839.21 µs 739.29 µs -11.9%
✅ gather/small_bf16_threads=1-internal/4096 547.1 ns 481.3 ns -12.0%
🟢 matmul/large_generic_f32_threads=8/32x1024x1024 5.03 ms 3.90 ms -22.6%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.76 3.62 5.06 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@codecov

codecov Bot commented Aug 20, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.64%. Comparing base (4a9f4ec) to head (3babf71).
⚠️ Report is 67 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1434      +/-   ##
==========================================
+ Coverage   82.10%   82.64%   +0.54%     
==========================================
  Files          12       12              
  Lines        5471     5475       +4     
  Branches     5471     5475       +4     
==========================================
+ Hits         4492     4525      +33     
+ Misses        780      757      -23     
+ Partials      199      193       -6     
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (?)
cli-ort-windows 82.10% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant