Skip to content

perf(plugin): stop routing host intermediates through ORT scratch, and recycle them - #1073

Merged
justinchuby merged 11 commits into
mainfrom
deckard/f16-contrib-assignment
Aug 16, 2026
Merged

justinchuby merged 11 commits into
mainfrom
deckard/f16-contrib-assignment

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 16, 2026 •

Copy link
Copy Markdown
Owner

Problem

Every multi-node subgraph this EP claims was paying roughly 10x for each
intermediate tensor it produced.

Routed subgraphs allocate their intermediates with
KernelContext_GetScratchBuffer, using the memory info returned by
device_mem_info. On a device EP that is correct and necessary: a device kernel
handed a host pointer dereferences it as device memory. On a host EP,
device_mem_info finds no device-resident input and falls through to the memory
info of kernel-context input 0 — CPU memory — so the host path was going through
ORT's scratch allocator too.

That allocator services the request through an aligned allocation, and glibc
maps a fresh region for an aligned request of this size regardless of the
M_MMAP_THRESHOLD tuning that makes ordinary malloc reuse the heap. So every
intermediate was a brand-new mapping: first-touch page faults charged to the
producing kernel's write, and an unmap when the Run ended. Nothing was ever
reused, and an N-node chain touched N cold megabytes per Run.

The cost is invisible in a single-node subgraph — which is why it survived this
long — and it is exactly the shape ORT hands us for the float16 contrib
activations, because ORT has no float16 CPU kernel for them and inlines their
function bodies into 15-node Cast/Mul/Add/Tanh primitive graphs.

How it was found

An 8-node Relu chain, f32, one thread. Same kernel, same size, but position in
the subgraph changed the time by more than 10x:

position in subgraph time per node, 262144 elements
first (ORT input → buffer output) 37 us
interior (buffer → buffer) 463 us
last (buffer → ORT output) 460 us
ORT CPU EP, same graph 27 us

Phase timing inside compute_execute put 98% of it inside
execute_with_workspace, not in the allocation call — consistent with page
faults being charged to the kernel's first write rather than to the allocator.
Instrumenting try_dense_elementwise confirmed the SIMD fast path was firing
for every node and the overlap guard never tripped, so it was not a dispatch
regression. Forcing the host-buffer arm dropped the same graph from 3246 us to
383 us.

Change

Both parts are confined to host-resident partitions. Device EPs keep the exact
behaviour they have today.

  1. intermediate_scratch — intermediates go through ORT scratch only when
    the resolved memory info is a device (mem_info_is_device). Otherwise they
    are plain Vec<u8>. scratch_mem_info itself is unchanged and still feeds
    prepare_workspace's placement fallback.

  2. Liveness-based recycling — last_reader_per_buffer computes, from the
    routing table, the highest node index that reads each intermediate. A buffer
    is retired the moment its last reader has run, so the next allocation gets
    storage that is still in cache, and a bounded thread-local pool carries that
    storage across Runs. An unread buffer or an out-of-range index maps to
    None, which means "nothing keeps this alive" — the conservative answer,
    since a buffer is only ever retired after its recorded last reader.

Reused storage is not re-zeroed in release builds. That is not a weakening of what kernels may
assume: an output routed to an ORT sink already arrives as whatever
KernelContext_GetOutput returned, which ORT does not zero, and the
single-kernel path has always worked that way. Making buffer sinks behave the
same way means a kernel that fails to write its whole output now fails
identically wherever it sits in a partition, instead of only when it happens to
be last. Re-zeroing costs a full memset per intermediate — measured at 0.586x
of ORT against 0.801x without it on the 1 MiB chain.

Debug builds do write reused storage, with 0xFF poison rather than zeros —
NaN in every float width, -1 in every signed integer width, so it cannot hide
inside a tolerance comparison. A kernel that leaves part of its output unwritten
therefore fails loudly in every test run instead of inheriting something
plausible.

Benchmarks

AMD EPYC 9V74, AVX2 (no AVX-512, masked by the hypervisor), 1 intra-op thread,
taskset -c 0-15, 15 interleaved A/B rounds per point, ORT 1.28.0. Ratio is
ORT CPU EP session latency / plugin session latency, so >1.0 means the plugin
is faster
.

"Before" is this commit's parent, b0fd8a040 — not git merge-base HEAD origin/main. The merge base is the first parent of the f16 elementwise merge
this branch carries, so measuring float16 there reports ~0.10x and attributes
another PR's win to this one.

Numbers are steady state. The first large tensor a process touches pays a
one-off penalty the per-session warmup does not fully absorb, so a single
low-repetition run can read ~2x low on whichever size happens to go first.

8-node f32 Relu chain (a routed subgraph with 7 intermediates)

elements before after change
1024 0.183 0.623 3.4x
16384 0.081 0.782 9.7x
262144 0.066 0.801 12.1x

float16 activations that ORT inlines into multi-node bodies

21 interleaved rounds. p90 within 0.005 of p50 on every cell.

op elements before after
FastGelu 3072 0.452 0.715
FastGelu 262144 0.433 0.783
FastGelu 1048576 0.515 0.778
QuickGelu 3072 0.549 0.660
QuickGelu 262144 0.735 0.721
QuickGelu 1048576 0.733 0.735
Gelu(tanh) 3072 1.470 1.443
Gelu(tanh) 262144 1.531 1.534
Gelu(tanh) 1048576 1.451 1.460

QuickGelu's inlined body is shorter and produces fewer intermediates, so it
moves less. Gelu(tanh) is unchanged, as expected — it was already a single
claimed subgraph whose ratio is set by the kernels, not the plumbing.

Single-node subgraphs are untouched by construction (no intermediates): f32
Relu at 262144 is 0.944 before and after.

Limitations

  • This does not make everything win. The 8-node chain is at 0.80x, not 1.0x.
    ORT's allocation planner reuses two arena slots for the whole chain and its
    elementwise kernels thread across the intra-op pool while ours are
    single-threaded. Those are separate problems and are not claimed to be fixed
    here.
  • Measured on one machine and one ISA (AVX2, no AVX-512). The mechanism —
    aligned large allocations bypassing allocator reuse — is glibc-specific in its
    details, so the size of the win will differ elsewhere. The direction should
    not: reusing warm storage cannot be worse than mapping cold storage.
  • The pool is bounded at 8 buffers per thread. A partition with a wider live set
    than that falls back to fresh allocation for the excess rather than growing
    without limit.

Tests

Eight new unit tests in compute.rs:

  • last_reader_marks_the_final_consumer_of_each_buffer
  • last_reader_takes_the_highest_index_when_a_buffer_is_read_twice — the
    falsifier for the dangerous failure mode, a buffer freed before its second
    reader
  • last_reader_is_none_for_unread_and_out_of_range_buffers
  • recycled_intermediate_storage_is_reused_without_reallocating — asserts
    address reuse, which is the entire point
  • a_recycled_buffer_serves_a_smaller_request_at_the_requested_length — the
    length must be the requested one, since byte_len bounds every
    from_raw_parts built from the buffer
  • a_request_larger_than_every_pooled_buffer_allocates_fresh_zeroed_storage
  • scratch_backed_buffers_are_not_pooled and the_pool_is_bounded

Each pool test drains the thread-local pool first, so they are deterministic
under both the parallel harness and --test-threads=1.

One new end-to-end test against real ORT, in plugin_ort_e2e.rs:

  • conformance_chain_add_mul_repeated_runs_do_not_leak_stale_intermediates —
    the three-node Add/Mul/Add fixture, six Runs with changing inputs. The
    first Run gets freshly zeroed storage; every later one is served recycled,
    dirty buffers, so an element a kernel failed to write would surface as the
    previous iteration's answer. In debug builds it is served 0xFF poison
    instead, i.e. NaN, which no tolerance admits.

  • cargo test -p onnx-runtime-ep-plugin --lib → 224 passed, parallel and
    --test-threads=1.

  • NXRT_REQUIRE_ORT_TESTS=1 cargo test -p onnx-runtime-ep-cpu-plugin → 32
    passed against real ORT (31 before, plus the new one), with the debug poison
    active.

  • cargo test -p onnx-runtime-session --lib → 161 passed.

  • cargo clippy -p onnx-runtime-ep-plugin --all-targets → clean.

  • cargo fmt --all --check → clean.

Independent review

Reviewed independently: GO WITH FINDINGS, no blockers. The reviewer built
both sides in a separate worktree and reproduced every cell:

case claimed before measured before claimed after measured after
Relu chain=8, 1024 0.183 0.197 0.623 0.644
Relu chain=8, 16384 0.081 0.081 0.782 0.830
Relu chain=8, 262144 0.066 0.072 0.801 0.796
FastGelu f16, 3072 0.452 0.464 0.715 0.728
FastGelu f16, 262144 0.433 0.432 0.783 0.788
FastGelu f16, 1048576 0.515 0.519 0.778 0.772
QuickGelu f16, 1048576 0.733 0.742 0.735 0.736
single-node Relu f32, 262144 — 0.915 ~0.944 0.934

They also:

  • Walked the lifetime transmute at compute.rs:2065 against every new
    .take() and confirmed the TensorView<'static> aliases are dead before any
    retirement runs — moving a Vec into the pool preserves its heap address, and
    a drop can only happen after the last use.
  • Audited the kernels for anything relying on zeroed output and found none
    reachable as an interior node; confirmed absent_scratch still allocates
    zeroed (compute.rs:2117).
  • Ran their own numerical validation across Relu chain=8 and chain=15,
    Sigmoidx12, Tanhx10, and f16 FastGelu/QuickGelu, 50 iterations each so
    the pool serves dirty storage throughout — all matched ORT (f32 max diff 0,
    f16 within 2e-3).
  • Re-ran the plugin unit tests three times parallel and three times with
    --test-threads=1 looking for pool-related flakiness; none.

Findings applied: the debug poison fill and the repeated-run e2e test are theirs
(they asked for a guardrail on the non-zeroing relaxation); the benchmark
section now names the correct before-commit and warns about first-touch
dispersion; the test count is corrected.

justinchuby and others added 9 commits August 16, 2026 09:30
`x * 0.5` and `x + b` ran at ~7.7 ns/element, about 60x slower than
ORT's CPU kernels. That is not a corner case: when ORT has no kernel for
a contrib op at a given dtype it inlines the ONNX function, and the
inlined body is almost entirely scalar-broadcast Mul and Add.

Two separate causes, both structural.

First, `binary_contiguous` only fires when all three shapes are
identical, so any broadcast -- including a scalar operand -- fell to
`broadcast_apply`, which per element does a dot product over the tensor
rank, a `next_index` carry chain and a closure call, and additionally
allocates a whole-tensor accumulator that it then walks three times.
`Add` had no non-mlas fast path at all, so even a same-shape f32 Add
took that route.

Add `binary_broadcast_contiguous`, which recognises an operand whose
shape is a right-aligned suffix of the output shape -- a scalar, or the
row-bias case `[B, S, C] op [C]` -- and walks it as a repeated
contiguous block. An interior unit axis such as `[B, 1, C]` is not a
suffix and is still declined to the general path. `Add` now shares this
primitive rather than growing a second copy of the walk.

Second, `BinOp::apply` matches on a runtime value, and left inside the
element loop that match is re-evaluated per element and blocks
auto-vectorisation outright. With only the first fix an f32 scalar
multiply was still 1.1 ns/element. `dispatch_binop!` resolves the
combiner once, outside the loop, for both the new broadcast walk and the
existing same-shape one.

Measured, 1 thread, interleaved, ratio = ORT ns / plugin ns:

  Mul by scalar n=1048576   0.016 -> 0.961   (8577 us -> 137 us)
  Add by scalar n=1048576   0.017 -> 0.967   (8058 us -> 137 us)

Equivalence with the general walk is pinned bitwise: every case is run
twice over the same logical values, once with a contiguous broadcast
operand and once with a strided view of padded storage that
`dense_operand` rejects, and the raw output bytes must match.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The two new fast-path tests asserted that the process-global
`ADD_SCALAR_TEST_HITS` counter did *not* advance across an
`AddKernel::execute` call. Cargo runs tests in parallel and several other
tests in the same binary legitimately increment that counter, so the
equality assertion raced and failed intermittently (observed ~1 run in 3
in debug).

Assert the dispatch predicate directly instead: call
`add_dense_fast_path` and require it to accept (and produce the right
values), which is deterministic. `AddKernel::execute` only reaches the
counter after that predicate declines, and the two arms ahead of it
(`mlas`, vDSP) both require identical operand shapes, so a broadcasting
input the predicate accepts provably never reaches the fallback.

Adds the decline half as its own test: an interior unit axis must be
refused outright and must leave the output untouched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A 16-bit binary elementwise op computes
`from_acc(fold(to_acc(a), to_acc(b)))` with `Acc = f32`, so every element
paid two `half` software conversions in and one out. That was ~3.6
ns/element -- roughly 12x slower than ORT -- and it dominated every f16
graph, including the bodies ORT inlines for the f16 contrib activations
(FastGelu, QuickGelu, Gelu), which are mostly Cast/Mul/Add.

Widen both operands in `HALF_STAGE_CHUNK`-element passes with the
existing F16C/AVX2 bulk converters, fold in f32, and narrow back. The
scalar-operand shape -- what an inlined activation emits -- skips the
second staging buffer entirely and folds against a register constant.

f16 @1m elements, vs ORT, interleaved A/B, 1 thread:
  Mul dense   0.082 -> 1.278
  Add dense   0.100 -> 1.308
  Mul scalar  0.086 -> 1.170
  Add scalar  0.083 -> 1.167
f16 FastGelu 0.030 -> 0.524, QuickGelu 0.030 -> 0.730.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ementwise

# Conflicts:
#	crates/onnx-runtime-ep-cpu/src/kernels/elementwise.rs
The comment claimed three f32 buffers totalling 12 KiB, but the staged
loop allocates two (one per operand); the narrow step writes back through
the left buffer. Corrected to two buffers / 8 KiB and cross-referenced
F16_STAGE_CHUNK in dense_elementwise, which stages the unary paths with
the same chunk size.

Found by independent review.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…d recycle them

A multi-node partition was paying ~10x for every intermediate it produced.
Intermediates were allocated with KernelContext_GetScratchBuffer using the
memory info of kernel-context input 0, which on a host EP is CPU memory.
ORT services that through an aligned allocation, and glibc always maps a
fresh region for an aligned request of this size, so each intermediate was
a new mapping: first-touch page faults while the producing kernel wrote it,
and an unmap at the end of the Run. Nothing was ever reused, and a chain of
N nodes touched N cold megabytes.

Two changes, both confined to host-resident partitions:

* Intermediates go through ORT scratch only when the resolved memory info
  is a *device*. That is the case scratch exists for - a device kernel
  handed a host pointer dereferences it as device memory - and it is
  untouched. Host partitions take a plain Vec instead.
* Host intermediates are recycled by liveness. A buffer is retired as soon
  as the last node that reads it has run, so the next allocation reuses
  storage that is still in cache, and a thread-local pool carries the
  storage across Runs.

Measured on an 8-node f32 Relu chain, one thread, 15 interleaved rounds,
as a fraction of ORT CPU EP session latency:

  elements   before   after
     1024    0.183    0.623
    16384    0.081    0.782
   262144    0.066    0.801

f16 FastGelu, which ORT inlines into a 15-node primitive body, goes from
0.43-0.52x to 0.72-0.78x.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 16, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.96403% with 7 lines in your changes missing coverage. Please review.
✅ Project coverage is 79.73%. Comparing base (bf0ae66) to head (7ae60da).
⚠️ Report is 3 commits behind head on main.

Files with missing lines Patch % Lines
crates/onnx-runtime-ep-plugin/src/compute.rs 94.96% 2 Missing and 5 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1073      +/-   ##
==========================================
+ Coverage   79.71%   79.73%   +0.01%     
==========================================
  Files         369      369              
  Lines      160706   160844     +138     
  Branches   160706   160844     +138     
==========================================
+ Hits       128113   128242     +129     
- Misses      27860    27865       +5     
- Partials     4733     4737       +4     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.40% <ø> (+0.09%) ⬆️
mlas 83.39% <ø> (+0.11%) ⬆️
offline 79.52% <94.96%> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-ep-plugin/src/compute.rs 77.41% <94.96%> (+0.65%) ⬆️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Aug 16, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 40.36 µs 92.38 µs +128.9%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 378.56 µs 695.20 µs +83.6%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 888.73 µs 1.44 ms +62.0%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 43.37 µs 69.99 µs +61.4%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 28.90 µs 44.50 µs +54.0%
🔴 gather/large_f16_threads=1-internal/131072 11.78 µs 17.39 µs +47.6%
🔴 matmul/small_generic_f32_threads=8/1x256x256 31.90 µs 46.60 µs +46.1%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.24 ms 1.76 ms +41.7%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 498.08 µs 693.20 µs +39.2%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 433.17 µs 587.93 µs +35.7%
🔴 matmul/medium_generic_bf16_threads=1/32x512x512 489.44 µs 659.44 µs +34.7%
🔴 matmul/small_generic_f16_threads=1/1x256x256 28.16 µs 36.67 µs +30.2%
🔴 matmul/small_generic_f32_threads=1/1x256x256 34.11 µs 44.35 µs +30.0%
⚠️ gather/large_bf16_threads=1-internal/131072 12.01 µs 15.43 µs +28.5%
⚠️ matmul/large_generic_f32_threads=8/32x1024x1024 3.69 ms 4.74 ms +28.5%
⚠️ matmul/small_generic_bf16_threads=1/1x256x256 29.76 µs 35.79 µs +20.3%
⚠️ matmul/large_generic_f32_threads=1/32x1024x1024 8.63 ms 10.34 ms +19.9%
⚠️ matmul/small_generic_f16_threads=8/1x256x256 28.97 µs 34.56 µs +19.3%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 73.02 µs 86.90 µs +19.0%
⚠️ grammar_masking/llguidance_compute_mask/32 69.74 µs 82.44 µs +18.2%
⚠️ matmul/medium_generic_f32_threads=1/32x512x512 2.13 ms 2.48 ms +16.4%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 86.30 µs 100.07 µs +16.0%
⚠️ matmul/small_generic_bf16_threads=8/1x256x256 29.39 µs 34.05 µs +15.9%
✅ matmul/medium_generic_f16_threads=1/32x512x512 27.71 µs 30.83 µs +11.2%
✅ logit_processing/seven_processor_chain_per_step 297.37 µs 325.50 µs +9.5%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 155.87 µs 170.00 µs +9.1%
✅ add/large_f32_threads=1-internal/4194304 553.55 µs 574.42 µs +3.8%
✅ add/medium_bf16_threads=1-internal/262144 95.60 µs 98.92 µs +3.5%
✅ gather/medium_bf16_threads=1-internal/32768 2.34 µs 2.39 µs +2.3%
✅ gather/small_f16_threads=1-internal/4096 437.7 ns 447.6 ns +2.2%
✅ add/large_bf16_threads=1-internal/4194304 1.52 ms 1.53 ms +0.9%
✅ gather/medium_f32_threads=1-internal/32768 3.44 µs 3.47 µs +0.9%
✅ reduce_mean/medium_f32_threads=1-internal/65536 225.74 µs 226.83 µs +0.5%
✅ tokenization/decode_tokens_per_second 5.72 ms 5.71 ms -0.0%
✅ gather/medium_f16_threads=1-internal/32768 2.23 µs 2.23 µs -0.1%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.93 ms 1.92 ms -0.4%
✅ kv_cache/alloc_dealloc_pages 36.70 µs 36.54 µs -0.4%
✅ add/medium_f32_threads=1-internal/262144 22.62 µs 22.52 µs -0.5%
✅ gather/small_f32_threads=1-internal/4096 625.6 ns 621.7 ns -0.6%
✅ reduce_mean/large_f32_threads=1-internal/262144 927.30 µs 920.52 µs -0.7%
✅ gather/large_f32_threads=1-internal/131072 34.40 µs 34.08 µs -0.9%
✅ add/large_f16_threads=1-internal/4194304 1.54 ms 1.52 ms -1.1%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.33 µs 14.04 µs -2.0%
✅ add/small_f16_threads=1-internal/1024 431.1 ns 422.1 ns -2.1%
✅ tokenization/encode_tokens_per_second 361.94 µs 353.87 µs -2.2%
✅ add/medium_f16_threads=1-internal/262144 99.27 µs 96.72 µs -2.6%
✅ add/small_bf16_threads=1-internal/1024 428.0 ns 410.3 ns -4.1%
✅ sampling_latency/top_k_per_token 51.32 µs 48.55 µs -5.4%
✅ gather/small_bf16_threads=1-internal/4096 466.8 ns 440.5 ns -5.6%
✅ qwen3_sampling_processors/top_k_partial_selection 146.43 µs 133.67 µs -8.7%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.92 ms 5.37 ms -9.4%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 535.24 µs 482.08 µs -9.9%
✅ sampling_latency/greedy_per_token 3.35 µs 3.01 µs -10.2%
✅ qwen3_sampling_processors/top_k_top_p_fast 683.57 µs 608.18 µs -11.0%
✅ sampling_latency/top_p_per_token 403.43 µs 358.49 µs -11.1%
✅ sampling_latency/min_p_per_token 220.75 µs 193.23 µs -12.5%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.27 ms 1.95 ms -13.8%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.78 ms 3.25 ms -14.1%
🟢 add/small_f32_threads=1-internal/1024 254.4 ns 179.5 ns -29.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.62 3.50 4.45 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby and others added 2 commits August 16, 2026 13:13
… end

Two guardrails for the non-zeroing relaxation, both from review.

Debug builds fill a reused buffer with 0xFF before handing it back: NaN in
every float width, -1 in every signed integer width, so a kernel that
leaves part of its output unwritten cannot have the gap absorbed by a
tolerance comparison. Release builds still skip the write, which is the
cost the change exists to avoid.

The new end-to-end test runs the three-node Add/Mul/Add fixture six times
against real ORT with changing inputs. Only the first Run gets zeroed
storage; every later one is served recycled buffers, so a partial write
would surface as the previous iteration's answer. A single Run cannot
catch that.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ssignment

# Conflicts:
#	crates/onnx-runtime-ep-cpu/src/kernels/elementwise.rs
@justinchuby
justinchuby marked this pull request as ready for review August 16, 2026 19:56
@justinchuby
justinchuby merged commit f8f6f50 into main Aug 16, 2026
12 of 16 checks passed
@justinchuby
justinchuby deleted the deckard/f16-contrib-assignment branch August 16, 2026 19:56
justinchuby added a commit that referenced this pull request Aug 16, 2026
## What

`cargo fmt --all -- --check` currently **fails on `origin/main`**
(`bce03cabb`) in five
places across `crates/onnx-runtime-ep-cpu/src/kernels/gemm.rs` and
`crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs`. This is the `cargo
fmt --all` output
and nothing else.

## Why it happened

Nobody wrote badly-formatted code. #1073, #1079 and #1080 each touched
these two files and
each was fmt-clean against its own base. Squash-merging them produced a
combined text that
rustfmt formats differently:

- `TRANSPOSE_TEST_LOCK.lock().unwrap_or_else(|e| e.into_inner())` is 61
characters, which
exceeds rustfmt's default `chain_width` of 60 once it sits at test-body
indentation (two
  sites).
- The `onnx_runtime_ir` import list grew past `max_width` and now wants
braces on their own
  lines.
- One `Gemm`/`transB` attribute chain became short enough to fit on a
single line after a
  neighbouring edit.
- One stray double blank line at end of a `mod tests`.

`main` is unprotected, so no required check re-ran fmt on the merge
result and the breakage
landed silently.

## How it was found

It blocked #1086: that PR's `Fast (Linux x86_64)` and `Rust quality`
jobs failed on the
PR **merge ref** with diffs in files #1086 does not touch. Reproduced
independently by
checking out `origin/main` into a clean worktree and running `cargo fmt
--all -- --check`.

## Verification

- `cargo fmt --all -- --check` -> clean (was: 5 diffs).
- Diff is whitespace/line-breaking only; `git diff -w` on the two files
is empty apart from
  the import-brace move. No logic, no behaviour, no test changes.

## Risk

None. Mechanical formatter output.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant