Skip to content

ci: stop an unrelated failure from skipping the EP-assignment falsifiers - #1096

Merged
justinchuby merged 2 commits into
mainfrom
deckard/ci-plugin-gate
Aug 16, 2026
Merged

justinchuby merged 2 commits into
mainfrom
deckard/ci-plugin-gate

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 16, 2026 •

Copy link
Copy Markdown
Owner

The EP-assignment falsifiers have not been running in CI

While verifying a review finding on #1093 — "could these new plugin E2E falsifiers pass vacuously?" — I checked whether the lane that runs them is actually green. It is not, and the answer turned out to be worse than vacuous passing: they are not executing at all.

The Test onnx-runtime-ep-cpu-plugin (ORT gate) step carries:

if: runner.os == 'Linux'

GitHub Actions inserts an implicit success() into every if that does not already contain a status function, so this evaluates as success() && runner.os == 'Linux'. Any earlier failing step in the CLI ORT job skips it.

That is the state of main today. Step 12, Test ORT-backed workspace crates, fails on a pre-existing, unrelated onnx-genai-engine test, so the plugin step is reported skipped:

step name conclusion
11 Test onnx-genai-cli with coverage success
12 Test ORT-backed workspace crates failure
13 Test onnx-genai-engine native backend skipped
14 Test onnx-runtime-ep-cpu-plugin (ORT gate) skipped

(CLI ORT (Linux x86_64), run 31975532936, job 95234327113, on 21cd05b3d.)

Why this particular step matters

CLI ORT is the only job that sets NXRT_REQUIRE_ORT_TESTS: "1", and the only one that runs the plugin suite with ORT actually present. That combination is what makes the EP-assignment falsifiers fail-closed instead of silently skipping when ORT or the cdylib is missing.

To be precise about scope: onnx-runtime-ep-cpu-plugin is compiled and tested in other jobs too (Fast, Rust coverage, Rust (Windows ARM64), via workspace_test_packages.py). But in those jobs ORT is absent and the gate is unset, so the ORT-dependent falsifiers skip themselves internally. cli-ort is therefore the only place they genuinely execute and assert.

Those falsifiers are load-bearing for the CPU plugin's assignment honesty: they assert that size/dtype ranges measured slower than ORT are not advertised as EP assignments, that declined ops actually land on CPUExecutionProvider, and that performance deferrals still claim when CPU fallback is disabled. Silently not running is precisely the failure mode they were written to prevent.

The change

if: ${{ !cancelled() && runner.os == 'Linux' }}

This does not weaken any gate:

  • the job still fails if either step fails;
  • the pre-existing engine failure is untouched and stays red;
  • !cancelled() still short-circuits on a cancelled workflow, so it does not burn runner time on aborted runs.

For the test steps, this simply stops one failure from hiding another, so both are reported independently.

Tradeoff, stated explicitly: !cancelled() tolerates failure of any earlier step, including setup. If the ORT download in Build onnx-genai-cli fails, this step will now run and go red for that reason instead of being skipped. It cannot produce a false green — the falsifiers fail closed under the gate — so the cost is one extra red line on a job that is already red, in exchange for never silently losing the assertions. Scoping the condition to a specific step's conclusion would avoid the noise but couples the step to a step ID for little benefit.

The unrelated failure, for the record

Not fixed here — it is outside this change's scope and belongs to the engine's KV-budget logic, not the CPU EP:

thread 'pipeline_load_rejects_kv_page_pool_before_fixed_state_when_host_budget_cannot_fit_either'
panicked at crates/onnx-genai-engine/tests/decode_position_and_state_e2e.rs:132:5:
failed to resolve the shared pipeline KV memory budget

It reproduces locally on a clean checkout of main (2 passed; 1 failed), so it is environment-independent and genuinely pre-existing rather than a runner artefact.

Verification

  • .github/workflows/ci.yml parses, and the step's if/env resolve as intended (NXRT_REQUIRE_ORT_TESTS=1 on the job).
  • Plugin suite passes locally under the gate: the plugin_ort_e2e binary reports 47 passed; 0 failed with NXRT_REQUIRE_ORT_TESTS=1 on this machine . This branch's own CI run reports 40 passed for the same binary; the difference is exactly accounted for: this branch is based on main at 21cd05b3d, which predates perf(ep-cpu): vectorise Exp, and stop claiming the unary ops ORT still wins #1093 (merged as af04a613c) and its 7 added tests — 6 assignment falsifiers plus one no-fallback override test. 40 + 7 = 47.
  • Proof it works, from this PR's own CI (run 31977040589, job 95237979877): step 12 failure, step 13 skipped, step 14 success — the plugin step now executes, with real assertions rather than *** SKIPPED *** (assignment_policy_defers_float32_activations_to_ort ... ok, assignment_policy_claims_float16_gelu_without_inlining_it ... ok, assignment_policy_always_claims_bfloat16_activations ... ok, assignment_policy_yields_when_cpu_fallback_is_disabled ... ok), zero *** SKIPPED *** lines in the step's log, and the job's overall conclusion is still failure — confirming the gate is intact and nothing was masked.

Limitations

  • Does not fix the engine test, so CLI ORT stays red overall until that is addressed by its owner.
  • Only this step is changed. Test onnx-genai-engine native backend keeps the default success() semantics; widening that is a broader policy call and not needed to restore the falsifiers.
  • Fast (Linux x86_64) is the only required check, so this lane's colour does not gate merges either way — the value here is signal, not enforcement.

The `Test onnx-runtime-ep-cpu-plugin (ORT gate)` step carried only
`if: runner.os == 'Linux'`, which GitHub Actions evaluates as
`success() && runner.os == 'Linux'`. Any earlier failing step in the
`CLI ORT` job therefore skips it.

That is currently happening on `main`. `Test ORT-backed workspace crates`
fails in `onnx-genai-engine`'s
`pipeline_load_rejects_kv_page_pool_before_fixed_state_when_host_budget_cannot_fit_either`,
so step 14 has been reported as `skipped`, and the plugin suite — the
only place the EP-assignment falsifiers execute with
`NXRT_REQUIRE_ORT_TESTS=1` — has not run in that lane.

Those falsifiers exist specifically to prove that ranges measured slower
than ORT are not advertised as EP assignments and that declined ops
really do land on `CPUExecutionProvider`. Skipping them silently is the
one failure mode they were written to prevent, and it removes the
fail-closed guarantee the `NXRT_REQUIRE_ORT_TESTS` gate is meant to give.

Using `!cancelled()` keeps both steps reported independently. It does not
weaken the gate: the job still fails if either step fails, and the
pre-existing engine failure is untouched and still red. It only stops one
failure from hiding another.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 16, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 79.90%. Comparing base (21cd05b) to head (6e39e08).
⚠️ Report is 3 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1096      +/-   ##
==========================================
+ Coverage   79.87%   79.90%   +0.03%     
==========================================
  Files         370      370              
  Lines      162510   162862     +352     
  Branches   162510   162862     +352     
==========================================
+ Hits       129800   130133     +333     
- Misses      27957    27971      +14     
- Partials     4753     4758       +5     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.31% <ø> (ø)
mlas 84.17% <ø> (+0.11%) ⬆️
offline 79.69% <ø> (+0.03%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Review pointed out two imprecisions in the explanation, both fixed here;
the condition itself is unchanged.

The comment claimed this was the only job running the plugin crate. It is
not: `Fast`, `Rust coverage` and `Rust (Windows ARM64)` also compile and
run it via `workspace_test_packages.py`. What is unique to `cli-ort` is
that ORT is present *and* `NXRT_REQUIRE_ORT_TESTS=1`, which is what makes
the falsifiers fail closed instead of skipping themselves.

Also record the tradeoff: `!cancelled()` tolerates failure of any earlier
step, setup included, so an ORT download failure will now surface here as
a red step rather than a skip. It cannot false-green, so this costs one
extra red on an already-red job and buys never silently dropping the
assertions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 16, 2026 23:18
@justinchuby
justinchuby merged commit 744d867 into main Aug 16, 2026
12 of 16 checks passed
@justinchuby
justinchuby deleted the deckard/ci-plugin-gate branch August 16, 2026 23:18
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/large_generic_f16_threads=8/32x1024x1024 78.05 µs 207.19 µs +165.5%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 45.12 µs 102.08 µs +126.2%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 386.76 µs 871.27 µs +125.3%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 27.83 µs 50.10 µs +80.0%
🔴 matmul/medium_generic_bf16_threads=1/32x512x512 487.05 µs 641.37 µs +31.7%
🔴 grammar_masking/llguidance_compute_mask/32 70.00 µs 91.56 µs +30.8%
⚠️ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 43.49 µs 54.26 µs +24.8%
⚠️ gather/large_bf16_threads=1-internal/131072 9.27 µs 11.53 µs +24.4%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.82 ms 2.06 ms +13.1%
✅ matmul/medium_generic_f32_threads=8/32x512x512 914.05 µs 1.00 ms +9.7%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.72 ms 4.00 ms +7.7%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 540.90 µs 580.00 µs +7.2%
✅ logit_processing/seven_processor_chain_per_step 299.46 µs 317.89 µs +6.2%
✅ tokenization/encode_tokens_per_second 354.36 µs 375.44 µs +5.9%
✅ sampling_latency/top_p_per_token 358.94 µs 371.58 µs +3.5%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.31 ms 1.35 ms +3.3%
✅ matmul/small_generic_f32_threads=8/1x256x256 32.34 µs 33.07 µs +2.3%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 74.30 µs 75.30 µs +1.4%
✅ matmul/small_generic_f16_threads=1/1x256x256 27.65 µs 28.00 µs +1.3%
✅ matmul/small_generic_f16_threads=8/1x256x256 28.25 µs 28.59 µs +1.2%
✅ matmul/small_generic_bf16_threads=1/1x256x256 28.63 µs 28.90 µs +0.9%
✅ sampling_latency/min_p_per_token 192.89 µs 193.91 µs +0.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 612.19 µs 615.41 µs +0.5%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 1.97 ms 1.96 ms -0.3%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 482.34 µs 479.99 µs -0.5%
✅ sampling_latency/greedy_per_token 3.01 µs 2.98 µs -1.1%
✅ matmul/medium_generic_f16_threads=1/32x512x512 28.20 µs 27.87 µs -1.1%
✅ qwen3_sampling_processors/top_k_partial_selection 138.00 µs 136.14 µs -1.3%
✅ sampling_latency/top_k_per_token 49.25 µs 48.58 µs -1.3%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.17 ms 2.14 ms -1.4%
✅ kv_cache/alloc_dealloc_pages 36.67 µs 36.16 µs -1.4%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 159.83 µs 157.02 µs -1.8%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.32 ms 3.26 ms -1.8%
✅ matmul/small_generic_f32_threads=1/1x256x256 34.81 µs 33.96 µs -2.5%
✅ gather/small_f16_threads=1-internal/4096 478.2 ns 465.4 ns -2.7%
✅ matmul/small_generic_bf16_threads=8/1x256x256 30.77 µs 29.36 µs -4.6%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.54 ms 5.25 ms -5.3%
✅ add/small_f16_threads=1-internal/1024 441.6 ns 416.7 ns -5.6%
✅ tokenization/decode_tokens_per_second 5.88 ms 5.52 ms -6.1%
✅ gather/medium_f16_threads=1-internal/32768 2.38 µs 2.21 µs -7.1%
✅ gather/small_bf16_threads=1-internal/4096 480.3 ns 444.4 ns -7.5%
✅ reduce_mean/medium_f32_threads=1-internal/65536 247.68 µs 228.78 µs -7.6%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 484.23 µs 437.98 µs -9.6%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.97 ms 9.81 ms -10.6%
✅ gather/medium_bf16_threads=1-internal/32768 2.50 µs 2.23 µs -11.0%
✅ gather/small_f32_threads=1-internal/4096 728.6 ns 631.9 ns -13.3%
✅ gather/large_f16_threads=1-internal/131072 13.31 µs 11.48 µs -13.7%
🟢 add/small_bf16_threads=1-internal/1024 484.2 ns 406.7 ns -16.0%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.09 ms 910.85 µs -16.3%
🟢 reduce_mean/small_f32_threads=1-internal/4096 17.08 µs 14.05 µs -17.7%
🟢 add/medium_f32_threads=1-internal/262144 28.58 µs 22.82 µs -20.2%
🟢 gather/medium_f32_threads=1-internal/32768 4.37 µs 3.47 µs -20.7%
🟢 gather/large_f32_threads=1-internal/131072 29.08 µs 22.59 µs -22.3%
🟢 add/large_bf16_threads=1-internal/4194304 1.98 ms 1.52 ms -23.3%
🟢 add/large_f32_threads=1-internal/4194304 819.96 µs 585.89 µs -28.5%
🟢 add/small_f32_threads=1-internal/1024 258.4 ns 182.7 ns -29.3%
🟢 add/large_f16_threads=1-internal/4194304 2.21 ms 1.53 ms -30.7%
🟢 add/medium_bf16_threads=1-internal/262144 146.76 µs 96.75 µs -34.1%
🟢 add/medium_f16_threads=1-internal/262144 186.09 µs 96.57 µs -48.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.32 4.22 5.25 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 17, 2026
## What this changes

When our CPU EP is selected, it must not hand work to ORT's CPU EP. This
PR
removes the **three** mechanisms by which it was doing so for the
activation
and normalization families. `GetCapability` runs three independent
fail-closed
filters and a claim must clear all of them; each of the last two was
found only
after this PR had already claimed the job was done.

### 1. The performance-based decline policy is deleted

`assignment_policy.rs` (2118 lines) measured whether we beat ORT on a
given
shape/dtype and, where we lost, returned `ClaimPreference::defer` so
ORT's CPU
EP would take the node. That whole file is gone, along with the
`claim_preference` override in `provider.rs`. `claim_preference_node`
now
returns `Claim` immediately.

Beyond the architectural rule, the policy could not have worked as
intended:

- **A deferral splits the graph.** Every declined node is a partition
boundary,
which costs fusion, prepacking and buffer reuse across it — none of
which the
  per-node threshold accounted for.
- **It cannot see the thread count.** Capability runs before the
session's
intra-op pool is known. `Sqrt` at 64 Ki wins 1.9x against a
single-threaded
ORT and loses at 0.30x against 16 threads. One number cannot be right
for
  both.
- **Sometimes there is nothing to defer to.** ORT has no bf16 kernel for
these
ops and no f16 kernel for most. Declining a bf16 `Gelu` does not get a
faster
  kernel, it gets a load failure.
- **The thresholds were tuned on one host.** Every number was measured
on a
single EPYC 9V74. Shipping it made every other machine's latency a
guess.

Removing the override is also a small capability-time win: the default
adapter
deep-clones every input `Shape` and collects dtypes for every node in
order to
build the metadata the policy consumed.

### 2. The shape-inference filter was declining the same ops anyway

Deleting the policy did not, by itself, achieve the goal.
`GetCapability` runs
a second, independent fail-closed filter
(`onnx-runtime-ep-plugin/src/ep.rs`):
it drops any claim containing a node whose `ShapeInference::for_node`
returns
`Declined`, and that match ends in `_ => Declined`. **An op we register
a
kernel for, but which is absent from that table, is silently handed to
ORT no
matter what `supports_op` answers.** This is the same mechanism that
made the
`com.microsoft` activations unreachable until #1082.

The trigonometric, hyperbolic and remaining activation ops were all in
that
gap. Now listed: `Sin`, `Cos`, `Tan`, `Asin`, `Acos`, `Atan`, `Sinh`,
`Cosh`,
`Asinh`, `Acosh`, `Atanh`, `ThresholdedRelu`, `Swish`,
`com.microsoft::Silu`
and `PRelu`.

`Silu` had been deliberately excluded with the comment that ORT has no
kernel
for it. That reasoning was backwards: an op ORT cannot run is precisely
the one
we must never hand over.

`GroupNormalization` was in the same gap and is now covered too.

### 3. The dtype filter was declining a different set of ops

Review found a third filter, and it was still handing over one of the
very ops
section 2 had just fixed.

`node_passes_dtype_filter` looks the node's op up in the plugin's
`KernelRegistryEntry` list and returns `false` when there is no entry.
That
list is built from `build_cpu_registry_with_descriptors`, which recorded
keys
as they were registered — but `register_cnn_ops` takes `&mut OpRegistry`
and
writes *past* the recording wrapper. Eighteen ops were in the registry
and
absent from the descriptors, so `supports_op` claimed each one and
capability
then dropped it:

> `PRelu`, `BatchNormalization`, `InstanceNormalization`,
`GroupNormalization`,
> `Conv`, `ConvTranspose`, `MaxPool`, `AveragePool`, `GlobalMaxPool`,
> `GlobalAveragePool`, `GlobalLpPool`, `LpPool`, `Resize`, `GridSample`,
> `AffineGrid`, `Col2Im`, `CenterCropPad`, `SpaceToDepth`

Four are activations or normalizations this EP owns. **`PRelu` is the
sharpest
case: section 2 gave it a shape rule, so it cleared filter two, and the
dtype
filter declined it anyway.** Both pure-Rust inventory tests passed while
real
ORT ran the node.

Descriptors are now derived from `OpRegistry::keys()` instead of a
parallel
recorded list, making the two sets identical by construction rather than
by
convention. They are also sorted: they get leaked into a `'static` slice
ORT
reads, and hash-map iteration order would make any snapshot diff flap.

`descriptors_derived_from_real_registry_not_hand_maintained` had been
asserting
this bug as *correct behaviour* — it allowed a delta of up to 50 entries
and
named CNN ops as the expected difference. It now asserts set equality.

The lesson, having now been caught twice: an inventory test is only as
good as
its source of truth, and two review rounds passed on tests that
enumerated the
wrong set. The only check that cannot be fooled this way is the
end-to-end one
that asks real ORT which EP got the node.

## Scope — what this does *not* fix

**64 registered ops remain in the shape-inference gap.** This PR closes
the
activation/elementwise families; it does not close the gap universally.
The
full list is asserted exactly by the new inventory test and summarised
in
`docs/performance/CPU_ACTIVATION_GAPS.md`:

- **20 data-dependent** — output shape is a function of an input's
*values*
(`NonZero`, `Unique`, `Compress`, `Expand`, `Tile`, `Pad`, `TopK`,
`Split`,
  `Unsqueeze`, `Resize`, `AffineGrid`, `Col2Im`, `CenterCropPad`, ...).
  Correctly declined today, though most carry a constant initializer in
practice, so a pass that resolves initializer values at capability time
could
  claim them.
- **10 internal fusion ops** created after capability, never candidates
(`FusedGemm`, `FusedAttention`, `FusedMatMulBias`, the `pkg.nxrt` ops).
- **34 inferrable but unwritten — this is the work.** Ten are one-line
  shape-preserving rules (`QuantizeLinear`, `DequantizeLinear`,
  `CastLike`, `ScatterND`, `Trilu`, `CumSum`, ...); nine
  are pooling/CNN geometry (`MaxPool`, `AveragePool`, `Global*Pool`,
  `ConvTranspose`, `GridSample`, `SpaceToDepth`) inferrable exactly as
`build_conv` already does for `Conv`; eight more follow from attributes
(`ArgMax`, `Flatten`, `GatherElements`, `Size`, ...); the rest are
contrib
  and model ops.

Two entries deserve singling out. `com.microsoft::Attention` is *the*
attention
op in exported GenAI models, and the existing opset-23 arm is guarded to
the
default domain, so we hand it over. `LinearAttention` (both domains) and
`com.microsoft::CausalConvWithState` are the Qwen3.5 / Qwen3-Next hybrid
linear-attention primitives — **ORT has no kernel for them at all**, so
declining them does not get a faster implementation, it gets a load
failure.

Closing that remainder is the next PR. It is a different domain from the
activation kernels and each entry needs its own numeric test.

### Corrections to my earlier figures

I published two wrong counts before this test existed, in opposite
directions,
and the test found both.

**66, then 52, now 66 again — for different reasons each time.** The
first
figure came from a scratch script that text-matched the table and missed
its
guard arms, so it reported the whole `Reduce*` family (handled by
`op_name if
is_reduction(op_name)`) as a gap. Correcting that, I over-corrected and
claimed
the pooling family "has rules already" — it does not; `compute.rs` has
no
pooling arm at all. The 52 figure was also built on
`build_cpu_registry_with_descriptors`, which is **not** the registry:
`register_cnn_ops` writes straight to the inner `OpRegistry`, so 14 CNN
ops and
`PRelu` never appear in the descriptors at all. The test now enumerates
`OpRegistry::keys()` — the same set `supports_op` consults.

**`MoE`/`QMoE`/`LinearAttention`/`CausalConvWithState` were
misclassified as
internal fusion ops.** They are read from exported models —
`deepseek_v2_tiny_qmoe_native_e2e.rs` asserts a loaded graph contains a
`com.microsoft::QMoE` node, and the linear-attention pair are Qwen3.5
primitives. All are real gaps.

**Three gaps I had missed entirely:** `Unsqueeze` (declines whenever
`axes` is
input[1], i.e. every opset-13+ graph), `com.microsoft::Attention`, and
`EyeLike`.

This is the argument for the inventory test: hand-maintained prose about
which
ops reach ORT was wrong three times in a row, in both directions, and
each time
it read as confident.

## Tests

### Inventory - the test that would have caught this class of bug

`every_registered_op_has_a_shape_rule_or_is_a_known_gap` enumerates
`OpRegistry::keys()` and asserts the set of ops that decline shape
inference
*exactly*. Registering an op without a shape rule fails; adding a shape
rule
without removing the op from the list also fails. Neither direction can
pass
silently, and the allowlist doubles as the gap inventory. It needs no
ORT, so
it runs in every job rather than only the ORT-gated one.

The probe is a sweep, not a point: opsets {1, 13, 18, 22, 23} x arities
1..4 x
ranks 1..4, counting an op as declining only when nothing in the matrix
produces a rule. A single one-input rank-2 probe reported `Conv` as a
gap,
because `build_conv` reads `input_shapes[1][0]` for its output channel
count
and needs rank >= 3. Sweeping removes that artifact and keeps the
allowlist from
encoding one arbitrary opset.

Verified both directions falsify: deleting `| "Cos"` fails with `[("",
"Cos")]`;
adding `| "Trilu"` fails with `these ops now have a shape rule but are
still
listed as declined: [("", "Trilu")]`.

`no_activation_or_norm_op_is_left_to_ort` is a standing guard on the 39
activation and norm ops this EP owns - the families #1082, #1093 and
#1097 made
reachable. It asserts in both directions: a filter of the form
`registered.contains(op) && declines(op)` would let a *deleted
registration*
pass silently, which is the same hand-off by a different route. Verified
by
renaming the `Silu` kernel key, which now fails with `these ops are no
longer
registered by the CPU EP, so ORT will execute them: [("com.microsoft",
"Silu")]`. That direction immediately caught two entries I had wrong:
`Swish`
is registered in the default domain, not `com.microsoft`, and
`HardSwish` has
no kernel at all.

### Assignment sweeps

Two sweeps replace the 13 deleted deferral tests. Both iterate 21
fixtures:

- `no_supported_node_is_ever_left_to_the_ort_cpu_ep` - for every
fixture, the
  op under test appears in *our* EP's node list and in no other EP's.
- `every_fixture_loads_with_cpu_fallback_disabled` - loads each fixture
with
  `session.disable_cpu_ep_fallback`, so a silent hand-off becomes a load
  failure rather than a slow success.

Verified the sweep falsifies too: with `| "Sin"` removed from the table
it
fails with `[sin_assignment_f32] ours=[], others=["Sin"]`.

Both run fail-closed in CI: `conformance_setup` panics rather than
skipping when
`NXRT_REQUIRE_ORT_TESTS=1`, which the `CLI ORT` job sets. (That job's
plugin
step was itself being skipped whenever an earlier step failed - fixed
separately in #1096.)

`erf_reference` would have become dead code when the deferral tests were
deleted. Rather than remove it, it is now used by
`float16_biasgelu_runs_on_our_ep_with_correct_numerics`, which checks
*our*
kernel's numerics - more important now that we always execute `BiasGelu`
instead of sometimes declining it.

## Docs

`CPU_MATMUL_ASSIGNMENT.md` is reframed from claim/defer to win/gap.
Every
measurement is unchanged — the numbers still say exactly where we are
slower
than ORT; they are now a work list rather than a decline table.

`CPU_ACTIVATION_GAPS.md` is new: every range where we still lose, and
the two
root causes. Neither is polynomial accuracy:

1. **Our elementwise kernels are single-threaded** while ORT splits
across its
intra-op pool. This is the flat ~0.7-0.8x plateau across the f32
activation
   family, and it is worth more than any further approximation work.
   `KernelContext_ParallelFor` is the untried lever.
2. **~1.2 us of fixed per-node plugin dispatch overhead**, which is what
the
   ~0.75-0.8x at n=1 measures.

## Validation

`cargo fmt`, clippy clean on the three affected crates, 1266 `ep-cpu`
lib tests, 224 `ep-plugin` unit tests, 2 inventory tests, and 37 plugin
E2E tests under `NXRT_REQUIRE_ORT_TESTS=1` with real ORT.



## Update — third decline path (commit `67a611353`)

Opus review round 4 returned **NO-GO** on the grounds that a fourth
decline
path existed and the PR's central claim was therefore still false. It
was
right. See section 3 above.

Independently verified before fixing: 18 of the 177 unique `(domain,
op_type)`
pairs in the registry had no descriptor, including `PRelu`,
`BatchNormalization`, `InstanceNormalization` and `GroupNormalization`.
Counting individual registry entries (`op_type` + `domain` +
`since_version`,
the unit `OpRegistry::len()` reports) the registry holds 208, and
descriptors
now match it exactly —
`descriptors_derived_from_real_registry_not_hand_maintained`
asserts that equality.

New guards, each verified to fail when its invariant is broken:

| guard | falsified by | observed failure |
| --- | --- | --- |
| `every_registered_op_has_a_kernel_registry_entry` | filtering `PRelu`
out of descriptors | `1 registered ops have no kernel-registry entry ...
["::PRelu"]` |
| `activation_and_norm_ops_clear_every_capability_filter` | same |
`::PRelu: no kernel-registry entry (dtype filter declines it)` |
| `prelu_assignment_f32` (real ORT) | same | `ours=[], others=["PRelu"]`
→ `'PRelu' must run on this EP` |
| `descriptors_derived_from_real_registry_not_hand_maintained` | same |
names the missing ops instead of tolerating a delta of 50 |

After the fix, real ORT reports `ours=["PRelu"], others=[]` and
`ours=["GroupNormalization"], others=[]`.

Writing `activation_and_norm_ops_clear_every_capability_filter`
immediately
found two more real gaps: **`Celu` and `Mish` have no kernel at all**.
That is a
missing feature rather than a decline, so they are excluded from that
test with
a comment naming them, and recorded in `CPU_ACTIVATION_GAPS.md` under a
new
"Activations with no kernel at all" section rather than quietly dropped.

Re-validated: `cargo fmt --all --check`, scoped clippy clean, 1267
`ep-cpu` lib
tests, 224 `ep-plugin` unit tests, and 57 plugin tests under
`NXRT_REQUIRE_ORT_TESTS=1` (37 E2E + 4 coverage + 9 + 6 + 1).

## Update — Opus round 5: `GO WITH FINDINGS`

Round 5 confirmed the central claim now holds, verified against real ORT
(`ours=["PRelu"], others=[]` and `ours=["GroupNormalization"],
others=[]`), and
traced every node-removing gate in `ep_get_capability_inner` to confirm
no
fifth decline path exists for the activation/norm families. Two
findings, both
addressed:

**Finding 1 (minor, real) — `Conv` advertised a dtype its kernel
rejects.**
Giving `Conv` a `KernelRegistryEntry` was itself an over-claim:
`supported_dtypes_for_op("Conv", "")` returned `FLOAT_DTYPES`, which
includes
f64, but `ConvKernel::execute` rejects anything outside f32/f16/bf16.
Before
this PR `Conv` had no descriptor so an f64 `Conv` was declined; after
it, the
node would clear both filters, compile, and then fail at `Run`. Fixed by
adding
`MLAS_FLOAT_DTYPES` and giving `Conv` its own arm.

This is not a re-introduced fallback. Advertising a dtype we cannot
execute is
a different thing from declining one we can: f64 `Conv` is genuinely
unsupported, and reporting that honestly at capability time is correct.
The
rest of the CNN family really does dispatch f64 through
`dispatch_float!` and
keeps `FLOAT_DTYPES`. Pinned by
`conv_does_not_advertise_a_dtype_its_kernel_rejects`.

**Finding 2 (minor, docs) — stale counts in this description.** The
breakdown
said 20 + 10 + 36 = 66 while the total said 65, and still listed
`GroupNormalization` among the unwritten shape-preserving rules even
though
this PR wrote one. Corrected to 35 above. The "177" figure was unique
`(domain, op_type)` pairs, not `OpRegistry::len()` (208) — both are now
stated
explicitly. The committed `CPU_ACTIVATION_GAPS.md` was already correct.

---

## Merge with `main`

`main` moved under this branch and #1101 independently hit the *same*
dtype-filter
bug class, for `MatMulNBits` and `QLinearMatMul`, adding
`FLOAT_COMPUTE_DTYPES` —
byte-for-byte the same f32/f16/bf16 set as this branch's
`MLAS_FLOAT_DTYPES`.
Two people finding the same trap independently is the strongest evidence
that
the fail-closed dtype filter needed the systematic fix in this PR rather
than
another per-op patch. `kernels/mod.rs` was resolved by taking main's
block
verbatim, deleting the duplicate constant, pointing `Conv` at
`FLOAT_COMPUTE_DTYPES` and folding this branch's rationale into main's
doc
comment.

The merge then made the inventory test earn its keep twice:

**The probe was lying about `com.microsoft::MatMulNBits`.** `declines()`
supplied
no attributes, deliberately: an attribute-dependent rule should fall
back to the
ONNX default when the attribute is absent, and the empty bundle keeps
that
honest. But `MatMulNBits` derives its output width from `N`, and a node
without
`N` is *malformed*, not defaulted — `for_node` declines it, correctly.
No real
graph reaches that path, so the test was reporting a gap that does not
exist.
The probe now sweeps a plausible attribute bundle **alongside** the
empty one,
so default-fallback rules are still checked against absence while rules
that
cannot default are modelled the way production sees them.

**With that fixed, the first assert stopped masking the second.**
`QLinearMatMul` gained a real rule in main and was stale in this
branch's
`DECLINED` list. Removed. That is the drift check doing exactly the job
it was
written for — the list cannot quietly stop describing reality in either
direction.

Gap count 65 → 64; group 3, 35 → 34. `CPU_ACTIVATION_GAPS.md` updated to
match.

Post-merge validation, all green: `cargo fmt --all -- --check`; scoped
clippy
(`onnx-runtime-ep-cpu`, `-ep-plugin`, `-ep-cpu-plugin`, `--all-targets`)
clean;
1280 `ep-cpu` lib tests; and the full plugin suite under
`NXRT_REQUIRE_ORT_TESTS=1` — 40 real-ORT E2E assignment tests plus the 4
inventory tests.

> The `Fast` lane on this branch will stay red on `cargo fmt --all --
--check`
> until #1107 lands. That failure is `main`'s, from #1101, in
> `matmul_nbits.rs` — a file this PR does not touch. The fix is
deliberately
> **not** duplicated here; this branch will pick it up by merging `main`
once
> #1107 is in.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant