Repository navigation
perf(cpu-ep): give QLinearMatMul a native integer GEMM - #1194
Merged
Merged
Conversation
The default (non-MLAS) build had no integer GEMM. `execute` widened both
operands to `Vec<i32>` on every call -- 16 MiB for a 2048x2048 weight --
and then ran a scalar rank-1 update that re-streamed all of `B` once per
row of `A`. Section 3 of the matmul doc calls per-call packing "fixed",
but that only ever applied to `--features mlas`; the build we ship was
paying 12x ORT at both decode and prefill, the largest single loss in the
matmul family.
Add `kernels::qgemm_native`: the same arithmetic on the operand bytes,
with an AVX2 kernel behind a portable reference.
`vpmaddwd` is the instruction that fits: it takes eight `i16` pairs and
sums each pair into an `i32`, sixteen multiply-accumulates per
instruction. It is exact here rather than merely close -- a centred `a`
is in [-255, 255] and a raw `b` in [-128, 255], so a pair sum cannot
exceed 130050 -- and unlike `vpmaddubsw` it does not saturate, so no
operand has to be translated into another sign domain first.
Two kernels, chosen by `m`:
- `m > MR`: `B` is packed into 16-column, k-pair-interleaved tiles in a
256 KiB L2-resident panel that every row block re-reads. The pack is
SIMD and costs about 1% of the GEMM it feeds.
- `m <= MR` (decode): no pack at all. At one row block a panel would be
written once and read once, and writing `2 * k * n` bytes to serve a
GEMV that reads `k * n` is most of the call. The interleave happens
in registers instead, accumulators stay in registers across a `k`
block, and the column permutation `vpunpcklwd` implies is undone once
per block rather than once per iteration.
Every accumulation is wrapping `i32`, which is exactly arithmetic mod
2^32 and therefore associative and commutative, so neither the blocking
nor the thread count can change an output bit -- on overflow included.
The tests assert that directly against an integer oracle rather than
assuming it.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A column block owns a packed panel, so splitting the columns further to reach every worker shrinks the panel and re-walks `B`. Splitting the rows instead duplicates only the pack, which is around a percent of the GEMM it feeds. Left at one row block, an `n` of 2048 offers eight tasks, so a sixteen-worker pool leaves half of itself idle and spinning: 128x2048x2048 measured 2.69 ms at sixteen threads against 1.62 ms at eight. With the row split it is 1.99 ms, and the session A/B at four threads goes from 1.95x to 1.47x. Also document the native path in the matmul performance file, including the fact that the `QLinearMatMul` rows in the ranges table are `--features mlas` measurements that never described the build we ship. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…D grid The note still said a task owns whole rows of its column block, which was true before the row split. Describe the rectangle it actually owns, and derive the tile width from the block extent so a `block_width` that is not a multiple of `NR` cannot make the pack and the store disagree. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
marked this pull request as ready for review
August 18, 2026 04:17
justinchuby
enabled auto-merge (squash)
August 18, 2026 04:17
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Union-resolved docs/performance/CPU_MATMUL_ASSIGNMENT.md: kept both the branch's 3b (native QGEMM) and main's section 4 (f32 M=1 GEMV default).
`degenerate_extents_do_nothing` called `qgemm` with an empty `b_zero_points` and `n == 4`. `qgemm` asserts one zero point per column on entry, so that call is not one the function accepts. It passed only because `debug_assert` compiles out under `--release`, which is how this branch was being validated locally; a debug test profile fails it with `left: 0, right: 4`. Fixed the call rather than the assertion. The assertion states the contract the kernel's indexing depends on, and a caller whose `m` is zero still has `n` columns and still knows their zero points. `-p onnx-runtime-ep-cpu --lib` in a **debug** profile: 1440 passed, 0 failed (was 1439 passed, 1 failed). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1194 +/- ##
===========================================
- Coverage 82.10% 80.87% -1.23%
===========================================
Files 12 378 +366
Lines 5471 167778 +162307
Branches 5471 167778 +162307
===========================================
+ Hits 4492 135696 +131204
- Misses 780 27234 +26454
- Partials 199 4848 +4649
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
… to end The AVX2 tile constants (NR/MR/NC/KC/FUSED_KC) are only read by the x86 kernels, so an aarch64 build warned on all five and would have failed CI's -D warnings. Gate them to x86 alongside the code that uses them. Also extend qlinear_matmul_reordered_accumulation_is_bit_identical with a (4, 1029, 1100) shape. Every existing m <= 4 case sat below PARALLEL_MIN_WORK, so the pack-free kernel's column split had never been exercised end to end through requantize_rows -- only at the kernel level. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This was referenced Aug 20, 2026
justinchuby
added a commit
that referenced
this pull request
Aug 20, 2026
…es mlas` The guard asserted a *total* borrow count of exactly `before + 1`. Once QLinearMatMul got its native integer GEMM (#1194), `B` is taken through `dense_bytes` as well, so the count is 2 and the test fails on `main` whenever the `mlas` feature is on -- while the property it guards (the sign-flip route never writes through to the caller's `A`) still holds. Assert non-vacuity on `A` itself instead: `A` is in the borrow domain, and the call borrowed something. How many *other* operands a route borrows is a routing detail this test does not fix. The failure escapes CI because the mlas lanes are name-filtered. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 20, 2026
…es mlas` (#1523) ## What `kernels::qlinear_matmul::tests::the_sign_flip_route_never_writes_through_to_the_callers_input` fails on plain `origin/main` whenever the (non-default) `mlas` feature is on. ``` $ cargo test --release -p onnx-runtime-ep-cpu --features mlas --lib the_sign_flip_route_never_writes_through assertion `left == right` failed: the flip route no longer borrows A, so this test proves nothing left: 2 right: 1 ``` ## Why The test's non-vacuity guard counted `BORROWED_INPUT_CALLS` across *all* operands and asserted exactly `before + 1`. When QLinearMatMul got its native integer GEMM (#1194), `B` started going through `dense_bytes` too, so the count is now 2. The property the test actually guards — the MLAS sign-flip route goes through `Cow::to_mut` and so never writes through to the caller's `A` — **still holds**. The second assertion (`a.bytes == untouched`) passes. Only the guard is wrong. ## Fix Assert non-vacuity on `A` itself: - `dense_bytes(&a.view())` returns `Cow::Borrowed` — `A` is in the borrow domain, which is the precondition that makes the property non-trivial; - the call borrowed *something* (`> borrows_before`). How many *other* operands a route borrows is a routing detail this test does not fix, so it no longer asserts on it. ## Why CI missed it The `mlas` lanes are name-filtered, so this test is not selected there. Found while validating #1365, where an unrelated `--features mlas` route assertion had to be cfg-gated; I said in that PR I would file this separately. ## Validation ``` cargo test --release -p onnx-runtime-ep-cpu --features mlas --lib qlinear_matmul # 32 passed, 0 failed cargo test --release -p onnx-runtime-ep-cpu --lib qlinear_matmul # 17 passed, 0 failed cargo clippy --release -p onnx-runtime-ep-cpu --features mlas --lib --tests -- -D warnings # clean cargo fmt --all --check ``` Test-only change: no production code touched, no `mlas` symbol reaches the default artifact. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 20, 2026
`widen16(signed: bool, ..)` -- the u8 -> i16 widen every byte of `B` passes through in `qgemm_native` -- took signedness as a runtime argument and branched on it in the innermost loop, once per 16 bytes. `Operand::signed` is fixed for the whole call, but the `#[target_feature]` boundary stops the compiler hoisting the test out. It is now a `const SIGNED: bool` on `widen16`, `fused_strip`, `accumulate_fused` and `pack_panel`, with the single runtime `match` moved out to the block dispatcher that already matched on `m`. No arithmetic changed: accumulation is still wrapping `i32`, so the output is bit-identical, which the existing both-sign-domain oracle asserts. Kernel A/B, two prebuilt test binaries alternated over three repetitions with the harness's `portable` drift control steady to 1.6%: 1.13x at 1x3584x3584, 1.14x at 1x1024x3072 and 1.03x on the packed 128x3584x3584 path, with the two m=1 ranges non-overlapping between arms. End-to-end the effect is below what this host can resolve. A first attempt across two separately built bench binaries read 1.43x and is withdrawn -- a 1.13x kernel cannot yield a 1.43x call. A null control (one binary against itself) puts the paired end-to-end floor at +/-10%, and re-running each binary against its own ORT reference gives medians of 2.43x (base) and 2.58x (new). The change is kept for the mechanism, not for an end-to-end number: the branch is provably loop-invariant and the output is bit-identical. This was found by the first `QLinearMatMul` A/B this repository could run. `scripts/ort_ab/` had no generator for the op, so the ledger's `QLinearMatMul` rows -- `--features mlas` numbers under a caveat claiming 11.8x-12.0x on the default build -- had never been re-measured after #1194 landed the native integer GEMM. `gen_qlinear.py` closes that hole, with cells straddling both of the kernel's own dispatch gates. The corrected picture: 1.13x at u8 M=128, 0.11x at i8 M=1 (a 9.4x win, because ORT's own i8 path is 35x slower than its u8 path), and a real loss of ~2.4x-3.2x confined to u8 M=1, which this change does not close. Also corrects two stale claims the measurements invalidated: the ledger's 11.8x/12.0x `QLinearMatMul` caveat, and `scripts/ort_ab/README.md`'s instruction to set `ONNX_GENAI_CPU_MM_HALF_GEBP=0` to reach the half decode GEMV, which #1613 made unnecessary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 20, 2026
…#1616) ## What `widen16(signed: bool, ..)` — the `u8 -> i16` widen every byte of `B` passes through in `qgemm_native` (the native integer GEMM behind `QLinearMatMul`) — took signedness as a **runtime argument** and branched on it in the innermost loop, once per 16 bytes of `B`. `Operand::signed` comes from the input dtype and is fixed for the whole call, but the `#[target_feature]` boundary stops the compiler hoisting the test out. It is now a `const SIGNED: bool` on `widen16`, `fused_strip`, `accumulate_fused` and `pack_panel`, with the single runtime `match` moved out to the block dispatcher that already matched on `m`. **No arithmetic changed.** Accumulation is still wrapping `i32`, so the output is bit-identical — asserted by the existing `check(m, k, n, a_signed, b_signed)` oracle across both sign domains. ## Measurement Kernel A/B, two prebuilt test binaries alternated over three repetitions, `bench_qgemm_ab`'s `portable` arm as a drift control (steady to 1.6% across all six runs): | shape | base | new | ratio | | --- | ---: | ---: | ---: | | 1x3584x3584 | 19.97 GMACS | 22.52 | **1.13x** | | 1x1024x3072 | 17.52 | 19.89 | **1.14x** | | 128x3584x3584 | 61.92 | 64.00 | 1.03x | | *`portable` control* | 3.82 | 3.82 | *1.00x* | The two `m = 1` ranges do not overlap between arms. **End to end the effect is below what this host can resolve, and an earlier claim is retracted.** A first A/B across two separately built `bench_generic` binaries read 1.43x; that is withdrawn, because a 1.13x kernel cannot produce a 1.43x call (`1 / (f/R_k + 1 - f) <= R_k`). A null control (one binary against itself) puts the paired end-to-end floor at ±10%, and re-running each binary against its own ORT reference gives medians of 2.43x (base) and 2.58x (new). The change is kept for the mechanism — provably loop-invariant branch, bit-identical output, clean separation at the level where it acts — not for an end-to-end number. ## Why this was found now `scripts/ort_ab/` had **no `QLinearMatMul` generator**, so the ledger's `QLinearMatMul` rows (`--features mlas`, under a caveat claiming 11.8x–12.0x on the default build) had never been re-measured after #1194 landed the native integer GEMM. `gen_qlinear.py` closes that hole, with cells straddling both of the kernel's own dispatch gates (`PARALLEL_MIN_WORK` and the `m <= MR` fused/packed split). The corrected picture on the shipped build, one thread, parity `PASS` everywhere: **1.13x at u8 M=128**, **0.11x at i8 M=1** (a 9.4x win — ORT's own i8 path is 35x slower than its u8 path), and a real loss of **2.4x–3.2x confined to u8 M=1**, which this change does not close. Also localised, for the follow-ups: at `m = 1` a 1 MB L2-resident weight runs at the same 19.5 GB/s as a 12.85 MB one and every aspect ratio at equal footprint lands within 17–20 GB/s, so the kernel is **instruction-bound, not memory-bound**, up to L3. ## Also in here Two stale claims the measurements invalidated: the ledger's 11.8x/12.0x `QLinearMatMul` caveat, and `scripts/ort_ab/README.md`'s instruction to set `ONNX_GENAI_CPU_MM_HALF_GEBP=0` to reach the half decode GEMV, which #1613 made unnecessary. ## Still open (recorded, not fixed here) - The **~2.4x u8 M=1 gap** itself. - **Parallel scaling at m=1**: 8 threads buy 1.3x where ORT gets 5.8x. The fused path splits columns, handing every worker a page-crossing strided walk of `B`; at m=1 there is no `B` reuse to protect, so a `k` split with private accumulators is the shape that streams. - The single-thread residual is an instruction budget: exact full-range 8-bit needs `vpmaddwd`; `vpmaddubsw` would halve the uops but saturates unless one operand stays within ±64, which is why ORT's quantizer ships `reduce_range` for non-VNNI AVX2. Full record: `docs/benchmarks/2026-08-21-qlinearmatmul-m1-signedness.md`, ledger section 16. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The build we ship had no integer GEMM at all
QLinearMatMulin the default build did this per call:read_quantizedwidened operandAto aVec<i32>, then did it again for operandB. For a2048x2048
Bthat is a 16 MiB allocation and fill on every single call, thrown away at the endof it.
Arow by row, and each row re-streamed the whole ofB. Atm = 128that is 512 MiB of traffic for 1 GFLOP of work.The result was 11.8x ORT at
m = 1and 12.1x atm = 128— the largest single loss on thex86-64 CPU EP. The performance doc's
QLinearMatMulrows never described this build: they weretaken with
--features mlas, which is a research build we do not ship. That is now called out inthe doc.
This adds
kernels/qgemm_native.rs, a native byte-operand integer GEMM, and pointsqlinear_matmul.rsat it. Nothing defers, and nothing falls back.Two kernels, chosen by
mm <= 4(decode)Ameans a packed panel ofBis never reused, so packing is pure cost. Accumulators stay in registers across a 256-rowkblock.m > 4(prefill)KC/2pairs ofNCcolumns (KC = 512,NC = 256) is 256 KiB ofB, which stays in L2 while every row ofAsweeps it.The inner tile is
vpmaddwdoverNR = 16columns andMR = 4rows,kconsumed two rows at atime.
Why
vpmaddwdand notvpmaddubswMLAS gets 32 MACs from two instructions using
vpmaddubsw, which saturates: it needs asign-domain translation of
Band its intermediate is only nominally exact.vpmaddwdneeds fourinstructions for the same 32 MACs, but with centred
ain[-255, 255]and rawbin[-128, 255]a product is at most 65025 and a pair sum at most 130050, so it cannot saturate andcannot overflow. No sign-domain flip, no reasoning about clamped intermediates.
That instruction-count difference is the whole of the residual gap at
m = 1. Closing it meansgiving up exact integer arithmetic, which is not a trade I am willing to make for a quantized
kernel whose entire value is that it is exact.
Determinism is structural, not tested-in
The kernel computes
sum_k (a - za)(b - zb)assum_k (a - za) * b - zb * sum_k (a - za), withevery accumulation a wrapping
i32add. Wrapping addition is arithmetic mod 2^32, which isassociative and commutative, so any blocking, tiling, column split, row split or thread count
gives bit-identical output — including on overflow, where the wrap itself is reproducible.
wrapping_overflow_is_reordering_invariantandthe_thread_count_cannot_change_the_resultassertexactly that, and the SIMD path is checked bit-for-bit against a portable scalar oracle
(
the_simd_kernel_is_bit_identical_to_the_portable_loop, and separately for the fused path).Numbers
Session A/B against plain ORT,
K = N = 2048, u8 x u8, ratio isours / ORT, lower is better,p50 of 61 iterations. ORT's own timings moved under 1.5% between the two arms at 1 and 4 threads,
which is the control that makes the comparison mean anything.
i8_m1goes 0.206 ms to 0.049 ms at one thread.Kernel-level scaling (
bench_qgemm_ab,taskset -c 0-15), with the portable scalar arm as thecontrol:
The task grid splits rows as well as columns. Columns alone gave only
n / NCtasks — eight forn = 2048— so a sixteen-worker pool left half of itself spinning;128x2048x2048was 2.69 ms atsixteen threads against 1.62 ms at eight. Splitting columns further would shrink the panel and
re-walk
B; splitting rows duplicates only the pack, about a percent of the GEMM it feeds.Things I measured and rejected
Brows (PREFETCH_ROWS = 8): a consistent 8% regressionwith a stable
m = 128control. The hardware prefetcher already has the sequential stream.with a single
vperm2i128fixup perk-block flush. Saves eight instructions per 32 MACs.Left open, deliberately
Bpacked cache. The pack is repeated per call. Caching it would remove it fromprefill entirely, but any new weight-derived cache has to go through
kernels/governed_weight_cache.rsto satisfy the "New weight-derived caches must be governed"gate. That is a separate PR with its own eviction story, not a rider on this one.
alone does 0.090 ms, and past eight threads both arms get worse. That is the pre-existing
oversubscription item — it is present before and after this change, so it is not a regression
here, and it is the next thing I am working on.
Validation
cargo test --release -p onnx-runtime-ep-cpu --lib— 1340 passed, 0 failed.onnx-runtime-ep-cpu-pluginsuite withNXRT_REQUIRE_ORT_TESTS=1, including the 53-testplugin_ort_e2eORT conformance suite with CPU fallback disabled.cargo clippy -p onnx-runtime-ep-cpu --all-targetsclean,cargo fmt --all --checkclean.cargo check -p onnx-runtime-ep-cpu --lib --features mlas— the research build still compiles.edge-extent claims above; no blockers, two documentation fixes applied.
Refreshed against
main(2026-08-18)The branch was behind
mainand its red CI wall came from that, not from thischange:
crates/onnx-runtime-session/src/executor/mod.rs:175failed-D dead-codeon current stable, fixed on
mainbyca32b3adf(#1239) after this branch forked.origin/main(c55a3fab3) is merged in — no rebase, no force-push.One conflict, in
docs/performance/CPU_MATMUL_ASSIGNMENT.md, resolved as aunion: this branch's
#### 3b(the native integer GEMM) andmain's### 4(the f32M = 1GEMV becoming the default, #1091) were both new sectionsappended after 3a. Both are kept, in that order. Taking either side would have
silently deleted the other's record.
Revalidated on the merge commit, AVX2/FMA host, no AVX-512:
cargo test --release -p onnx-runtime-ep-cpu --lib— 1424 passed, 0 failed,18 ignored, including
qgemm_i32_matches_the_integer_oracle_for_every_signedness,the_simd_kernel_is_bit_identical_to_the_portable_loop,wrapping_overflow_is_reordering_invariantandthe_thread_count_cannot_change_the_result.The measurements in this PR were taken before the merge; nothing in the merged
range touches
qgemm_native.rs,qlinear_matmul.rs, or the CPU threadpool, sothey stand as recorded. The
mainchange that did land in this range (#1091'sf32
M = 1GEMV default) is on a different kernel family and is documented inthe section-4 text kept above.
Refreshed again against
main@6a855d5e0, and a real branch bug foundmainmoved again while this was queued (#1346/#1352/#1361 on the quality lane,#1154/#1232/#1238 on the CPU side). Merged in normally — no rebase — and
revalidated.
The revalidation caught something the earlier ones had not. Running
-p onnx-runtime-ep-cpu --libin a debug profile rather than--releasefails:degenerate_extents_do_nothingcalledqgemmwith an emptyb_zero_pointsand
n == 4.qgemmopens withdebug_assert_eq!(b_zero_points.len(), n),so that call is not one the function accepts — the test was exercising the
m == 0early return through an argument list the contract forbids. It passedevery previous run here only because
debug_assertcompiles out under--release, which is how I had been validating this branch locally. A debugtest profile fails it, and this is branch-caused:
qgemm_native.rsis new inthis PR.
Fixed in
9ca99e538by sizing the test's zero points tomandn, not byweakening the assertion — the assertion states the contract the kernel's
indexing depends on, and a caller whose
mis zero still hasncolumns andstill knows their zero points.
-p onnx-runtime-ep-cpu --lib, debug profile: 1440 passed, 0 failed (was1439 passed, 1 failed).
This is the second time on this stack that the profile a test runs under
decided whether it caught anything. Worth remembering:
--releasesilentlydisables every
debug_assertin the crate under test, so a localcargo test --releaseis not a substitute for what CI runs.Re-validated on latest
main(e0aedd0fa), 2026-08-19Latest
mainmerged in normally (no rebase). Full re-measurement, 1 threadpinned,
K = N = 2048, 61 iters / 10 warmup, 2 reps,ours_p50 / ort_p50:mainoursmainratiobench_qlinear_u8_m1bench_qlinear_u8_m128bench_qlinear_i8_m1ORT-side drift between the two arms was 0.7% at
m = 128and 0.0% oni8,which is the control that makes the comparison mean anything.
At
m = 128we are now 2.7x faster than ORT outright, andm = 1closesfrom 11.8x to 1.16x. These are better than the numbers originally posted above
because the dispatch work in #1077 landed in between.
Review fixes (
987aa0c5c)An independent review found no blockers but two things worth fixing:
NR/MR/NC/KC/FUSED_KCare readonly by the x86 kernels, so every non-x86 target warned on all five. CI
builds with
-D warnings, so this was a branch-caused CI failure waiting tohappen; the local x86 clippy run could never have caught it. Now
#[cfg]-gatedalongside the code that uses them: 0 warnings on both x86-64 and aarch64.
m <= 4shapein
qlinear_matmul_reordered_accumulation_is_bit_identicalsat belowPARALLEL_MIN_WORK, so the pack-free kernel's column split was only everchecked at the kernel level, never through
requantize_rows. Added(4, 1029, 1100), which forks both.The review independently re-derived the register-shuffle math in numpy
(
cvtep*_epi16,permute4x64_epi64(0xD8),unpacklo/hi_epi16,madd_epi16,permute2x128) against a plain per-column dot product over 2000 tiles withextreme values — 0 mismatches — and confirmed the
vpmaddwdnon-saturationbound for all four operand combos, the wrapping-add determinism claim, and the
absence of out-of-bounds access in every tail path.
Validation on the merged base
cargo fmtclean;cargo clippy --all-targets -D warningscleanonnx-runtime-ep-cputests, debug profile (sodebug_asserts are live)NXRT_REQUIRE_ORT_TESTS=1, release)every_assigned_node_is_also_executed_by_this_epandevery_fixture_loads_with_cpu_fallback_disabledgreen — nothing defers,nothing falls back to the ORT CPU EP