Skip to content

fix(cpu-ep): stop handing attention and MoE nodes to ORT's CPU EP - #1129

Merged
justinchuby merged 7 commits into
mainfrom
squad/leon-p5-no-decline
Aug 17, 2026
Merged

justinchuby merged 7 commits into
mainfrom
squad/leon-p5-no-decline

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 17, 2026 •

Copy link
Copy Markdown
Owner

What

Selecting this CPU EP is supposed to keep work off ORT's CPU EP. It did not, for
exactly the operators this EP exists for. This PR makes attention, MoE and
KV-cache nodes actually reach our kernels, proves it against a real ORT 1.27
session with fallback disabled, and checks the two recovered ops produce ORT's
numbers.

The project rule this implements:

我们的 cpu ep 不要分配到 ort cpu ep 上。用了我们的 cpu ep 就不要用 ort 的 cpu ep。但凡比 ort 慢的都要想方设法比 ort 快。

A losing range is a kernel to fix, not a node to give away.

Methodology

The perf-based decline (assignment_policy.rs) was already deleted and
claim_preference_node already returns Claim unconditionally, so the question
was what that actually guaranteed. Answer: nothing for attention.

plugin_ort_e2e's ASSIGNMENT_FIXTURES covered 23 activation and normalisation
graphs and zero attention, MoE, KV-cache, Softmax, Transpose or RoPE graphs.
The rule was asserted where it was easy. Fourteen fixtures close that gap
(generate_attention_fixtures.py), and
no_supported_node_is_ever_left_to_the_ort_cpu_ep now collects every failure and
asserts once, so a run reports the whole matrix instead of stopping at the first
decline.

Every measurement below is a real ORT 1.27 CPU session with
session.record_ep_graph_assignment_info=1, read back through
Session_GetEpGraphAssignmentInfo.

Evidence

Three ways a node was still reaching ORT

1. GetCapability's fail-closed shape filter — the silent layer. Three
independent gates must all pass. Deleting the policy addressed one; this third
one answers to neither of the others and was giving away
com.microsoft::Attention, MoE, QMoE, PackedMultiHeadAttention,
ScatterND, ScatterElements and Trilu. ORT has no CPU kernel for MoE,
QMoE or PackedMHA
, so the "fallback" bought a load failure, not a slower run.
All seven now have rules; shape_inference_coverage.rs asserts the allowlist
exactly in both directions, so this cannot silently regrow.

2. The per-op dtype union declined the two most important decode ops. With
the shape rules in, the first real run said:

2 of 32 fixtures were handed to ORT's CPU EP.
  rotary_assignment_f32: 'RotaryEmbedding'     — ours=[], ORT was given ["RotaryEmbedding"]
  gqa_assignment_f32:    'GroupQueryAttention' — ours=[], ORT was given ["GroupQueryAttention"]

The plugin advertises one dtype set per op and tests every input slot against
it. Attention ops map to FLOAT_DTYPES, but RotaryEmbedding's position_ids
is int64 and GroupQueryAttention's seqlens_k / total_sequence_length are
int32, so the integer slots failed the float test and the claim was dropped.
input_dtype_constraints_for_op already existed for the opposite problem (a
union too wide for MatMulNBits) and simply had no attention entries. The
com.microsoft and ai.onnx RoPE slot orders differ, so they need separate
tables.

Consequence worth flagging: any plugin-path RoPE or GQA measurement taken
before this fix — including #1078's — was measuring ORT, not us.

3. ORT stamps schema defaults on the node, and one of them is not 0. With
the dtypes fixed, GQA still failed, at kernel construction:

STAGE [CreateSession] FAILED: get_kernel failed for node '' (GroupQueryAttention):
  GroupQueryAttention: smooth_softmax is not yet supported (got -1)

The fixture never sets it. The contrib schema default is -1; ORT's own kernel
enables the feature only for the exact value 1, so -1 means off. Our != 0
test refused every GQA node ORT ever resolved. scale is the same hazard
inverted — ORT's kernels and ours both read 0 as "use 1/sqrt(head_size)", so
a stamped zero taken literally would multiply every score by zero and return a
silently wrong answer instead of an error. Fixed in attention.rs,
msft_attention.rs, multi_head_attention.rs and group_query_attention.rs.

Assignment matrix — all 37 fixtures, never partial

fixture op before after
softmax_assignment_f32 Softmax no fixture ours
transpose_assignment_f32 Transpose no fixture ours
kv_concat_assignment_f32 Concat no fixture ours
kv_scatternd_assignment_f32 ScatterND silently declined ours
scatter_elements_assignment_f32 ScatterElements silently declined ours
trilu_assignment_f32 Trilu silently declined ours
rotary_assignment_f32 RotaryEmbedding ORT (dtype union) ours
mha_assignment_f32 MultiHeadAttention no fixture ours
gqa_assignment_f32 GroupQueryAttention ORT (dtype union) ours
gqa_rotary_pos_assignment_f32 GroupQueryAttention + int64 position_ids ORT (dtype union, slot 9) ours
msft_attention_assignment_f32 com.microsoft::Attention silently declined ours
packed_mha_assignment_f32 PackedMultiHeadAttention declined, then ORT (dtype union) ours
moe_assignment_f32 MoE silently declined ours
moe_assignment_f16 MoE float16 declined, then ORT (f32-only union) ours
qmoe_assignment_f32 QMoE declined, then ORT (dtype union) ours
23 pre-existing activation/norm fixtures — ours ours

every_fixture_loads_with_cpu_fallback_disabled runs the same 38 with
session.disable_cpu_ep_fallback=1, so ORT is forbidden from placing a
supported node on its own CPU EP: 38/38 load, 38/38 assigned to us. For
PackedMultiHeadAttention that is also the only way the graph loads at all —
ORT has no CPU kernel for it, so declining it produced Could not find an implementation for PackedMultiHeadAttention(1) rather than a slower run.

Three fixtures are deliberately outside that set — qmoe_columnwise_f32,
moe_sparse_mixer_f32 and gqa_smooth_softmax_f32 — because we have no
kernel for those configurations. See Round 4.

Round 2: what independent review changed

The first round shipped nine fixtures — one per op I thought was affected.
Review asked for one per rescued op instead, and three of the four extra
fixtures immediately found a decline that was still live:

  • GroupQueryAttention + position_ids. That is optional input 9. I
    listed slots 5 and 6 because those were the slots my fixtures exercised, so a
    do_rotary node with explicit int64 positions still went to ORT. Our kernel
    runs that config and rotary_explicit_position_ids_apply_to_query_and_key has
    covered it all along — the claim path just never reached it.
  • QMoE (uint8-packed experts and zero points) and
    PackedMultiHeadAttention (int32 token_offset /
    cumulative_sequence_length) failed the same union.

The lesson is the specific one: shape_inference_coverage.rs builds synthetic
nodes and never opens a session, so it is blind to both the dtype filter and
the kernel factory. Only a real session per op finds these. That is the third
time that has cost something in this file's history.

Two corrections review forced, which I would rather state than bury:

  • The scale rationale was wrong. I claimed ORT stamps scale = 0. It does
    not — scale is OPTIONAL_VALUE with no schema default. The reviewer
    instrumented a real CreateSession over these fixtures and got
    scale_attr=None for GQA, MHA and com.microsoft::Attention (while
    smooth_softmax_attr=Some(-1) on a node that never set it, so that half
    holds), and removing the guard left the numerics unchanged at 1.006e-7. The
    guard stays as defence, not a fix, is now described that way, and now also
    covers packed_multi_head_attention.rs — the one attention kernel this change
    made reachable without it.
  • §23.4 said ORT has no CPU kernel for com.microsoft::Attention. It does
    (contrib_ops/cpu/bert/attention.cc). Round 3 showed the remainder of that
    statement was wrong too.

Round 3: what the second independent review changed

Round 2's findings are all resolved. The second review found three more, one of
which was a regression this PR itself introduced.

  • MoE was still declining float16 and bfloat16. Its kernel widens both to
    f32 and narrows on the way out, but supported_dtypes_for_op advertised
    F32_ONLY and MoE had no per-slot table, so the union decided it. The f32
    fixture passed while every production mixture — which is half precision —
    went to ORT. Union is now FLOAT_COMPUTE_DTYPES; moe_assignment_f16 proves
    it. One dtype's worth of coverage is not coverage.

  • Column-wise QMoE was a regression, now fixed. ORT leaves block_size
    without a schema default, so an absent attribute means the column-wise form —
    and ORT runs that on CPU; the reviewer loaded and ran it under 1.27 and
    1.28. Before this PR we declined QMoE and ORT ran such models. After Round
    2 we claimed it, and the kernel factory then rejected block_size = 0.
    A factory rejection arrives after ORT has compiled the node onto this EP,
    and no fallback recovers from it, so a previously-working model died at
    CreateSession.

    qmoe::unsupported_reason now mirrors the factory's limit in supports_op,
    where a decline is still recoverable, and
    column_wise_qmoe_is_declined_at_claim_time_not_failed_late asserts the node
    lands on ORT and the session loads. The general lesson is written into
    CPU_ACTIVATION_GAPS.md: any capability limit a factory enforces must be
    mirrored at claim time, because claiming an op is not free.

    To be explicit about the policy this PR exists to enforce: this is the one
    deliberate decline in the suite, and it is a capability answer, not a
    performance one. We have no column-wise implementation; the choice is between
    ORT running the model and nobody running it. It should stop being an exception
    as soon as that kernel exists, and it is not a precedent for declining ranges
    where we are merely slower.

  • "ORT has no CPU kernel for MoE or QMoE" was false. The reviewer loaded
    and ran these very fixtures on ORT's CPUExecutionProvider. Only
    PackedMultiHeadAttention genuinely lacks one, which its falsifier shows
    directly: reverting PACKED_MHA_SLOTS does not hand the node to ORT, it fails
    session creation. Corrected in four places — compute.rs, plugin_ort_e2e.rs,
    the fixture generator and §23.1/§23.4/§23.6 — because "ORT can't run it
    either" was being used to make a decline sound harmless when it was simply us
    not running an operator we implement.

Round 4: what the third independent review changed

The reviewer reproduced every round-3 number and ran both falsifiers, then
found that the column-wise QMoE fix was a point patch of a class. Two more
instances were live, on attributes production models set, both of which ORT's
CPU EP runs today — proven end-to-end through the real-ORT harness, not
inferred:

op attribute who sets it what happened before this commit
MoE / QMoE use_sparse_mixer=1 Phi-3.5-MoE, GRIN-MoE STAGE [CreateSession] FAILED: ... MoE: use_sparse_mixer=1 is unsupported
GroupQueryAttention smooth_softmax=1 Gemma-style attention sink STAGE [CreateSession] FAILED: ... GroupQueryAttention: smooth_softmax is not yet supported

Round 1 had narrowed the smooth_softmax guard from != 0 to == 1 to stop
rejecting ORT's -1 schema default — which fixed the default case and left the
genuinely-set case as an unrecoverable claim.

The fix is structural, not three more conditions. Each kernel's attribute
validation now lives in a single function, and the claim-time guard is that
function's error:

pub(crate) fn unsupported_reason(node: &Node) -> Option<String> {
    attributes_from_node(node).err().map(|e| e.to_string())
}

so a limit cannot be added to a factory without appearing at claim time —
drift is impossible by construction rather than by discipline. supports_op
consults it for MoE, QMoE, GroupQueryAttention, MultiHeadAttention,
com.microsoft::Attention, ai.onnx::Attention and
PackedMultiHeadAttention, which also sweeps up the reviewer's MAJOR (GQA's
quantized-KV and qk_output rejections) plus msft Attention's do_rotary /
past_present_share_buffer, ai.onnx::Attention's qk_matmul_output_mode and
QMoE's expert_weight_bits / quant_type.

Two tests pin it:

  • provider::tests::every_factory_attribute_rejection_is_mirrored_at_claim_time
    — pure Rust, so it runs everywhere. For eleven hostile nodes it asserts
    both that the factory rejects and that supports_op declines, failing if
    the two ever diverge.
  • plugin_ort_e2e::factory_only_capability_limits_are_declined_at_claim_time
    — real ORT, three fixtures: node lands on ORT, session loads. Disabling the
    guard produces a hard CreateSession failure for each of the three; I ran
    that falsifier arm-by-arm and all three fire.

To restate the policy boundary, since this PR exists to enforce it: these are
the only deliberate declines in the suite and every one is a capability
answer, not a performance one
. We have no column-wise, sparse-mixer or
smooth-softmax implementation; for each, the choice is between ORT running the
model and nobody running it. They are not exceptions to the "never hand a slow
range to ORT" rule — they are an admission that we owe three kernels, and each
stops being an exception the moment its kernel lands.

Round 4's reviewer found no blocking issue and reproduced every number:
1300 / 1318 / 231 / 45 tests, all four falsifiers firing, and — checking the
docs rather than the code — loading and running both new fixtures on ORT 1.28's
CPUExecutionProvider to confirm it really does run use_sparse_mixer=1 and
smooth_softmax=1. Two MINORs, both fixed in 34584d70e:

  • The PackedMultiHeadAttention row of the pure-Rust test passed for the
    wrong reason
    — its claim-time check inspects inputs, shapes and dtypes
    before reaching the attribute mirror, so an attribute-only node declined with
    "Q, K, V, token_offset, and cumulative_sequence_length are required" and the
    row still passed with the mirror deleted. It now builds a fully-formed node
    and asserts the decline reason; deleting the mirror fails it.
  • "Drift is impossible by construction" was too broad. It holds within a
    wired op; wiring an op is still discipline. The reviewer audited the rest of
    the EP and found ten other factories rejecting things supports_op does not
    pre-check (MatMulNBits bits/block_size, Resize, Pad, GridSample,
    LpNormalization, Unique, BitShift, ConstantOfShape, two pkg.nxrt
    internals). None is a live regression — each refuses only schema-invalid
    values or configurations ORT's own CPU kernel also refuses. Both the claim and
    that audit are now in the docs so it does not have to be redone.

The reviewer also independently reproduced every §23 number and ran each
falsifier — reverting the dtype rows, restoring smooth_softmax != 0, forcing
Attention back to Declined, and perturbing one baseline input to confirm the
numerics test is not vacuous (it failed at rel err 0.458). All fired as claimed.

Numerics — assignment is not the guarantee

A claimed node that computes the wrong answer is worse than the deferral it
replaced. rope_and_gqa_execute_on_our_ep_and_match_ort_numerics runs each
fixture twice over the same model file with the same input bytes: once with our
EP appended and fallback disabled, once with our EP simply not appended so ORT
resolves the node to its own contrib CPU kernel.

fixture output values max rel. error vs ORT
rotary_assignment_f32 output 4 096 1.182e-7
gqa_assignment_f32 output 4 096 1.006e-7
gqa_assignment_f32 present_key 1 048 576 0
gqa_assignment_f32 present_value 1 048 576 0

Both attention outputs are at float32 rounding; the KV cache is bit-identical.

Limitations

This makes the losing ranges reachable. It does not make them fast. The
fused attention region is still 3–15x short of ORT at 8–16 threads (§15), the
Transpose and MoE gaps in §18–20 are unchanged, and #1078's own evidence has
float32 RoPE losing 12/12 cells at 1.53–17.21x. Under the rule above every one
of those is a kernel to fix with no exit; RoPE is first because it is the op that
was being given away. Tracked in §23.6 of the benchmark doc.

Two further caveats:

  • The GQA fixture's declared present_* shape is the non-shared-buffer form,
    so ORT logs a benign MergeShapeInfo ... lenient merge warning (it infers the
    shared-buffer form). Both sides still produce identical output; our static
    rule is an upper bound on what the kernel writes, so it over-allocates by one
    row rather than under-allocating.
  • A capability gap this exposed: QMoE's fixture needs block_size = 32,
    because our kernel implements only the blocked form while an absent
    block_size selects the column-wise one. ORT can run that form, so the
    right follow-up is to implement it here rather than to leave it declined.
    Recorded in §23.6.

Docs

§10's assignment matrix and §15's "decline the whole region to ORT" conclusion
presented fallback as a solution. Both are marked withdrawn in place — the
ratios are kept because they are real evidence and they are the work queue, but
the conclusions drawn from them are void. New §23 records the rule, the audit,
both matrices and what is left. CPU_ACTIVATION_GAPS.md gains the two failure
modes this found (a dtype union that is too narrow, and the post-assignment
kernel factory).

Validation

  • cargo test -p onnx-runtime-ep-cpu --lib — 1300 passed;
    --features mlas — 1318 passed
  • cargo test -p onnx-runtime-ep-plugin — 231 + 2 + 9 passed (one stale test
    asserting com.microsoft::Attention is Declined updated to assert its new
    rule, and renamed to say what it checks)
  • cargo test -p onnx-runtime-ep-cpu-plugin with NXRT_REQUIRE_ORT_TESTS=1
    against real ORT 1.27 — all suites green, plugin_ort_e2e 45 passed / 1
    ignored, with all 38 fixtures
  • cargo fmt --all -- --check clean; cargo clippy --all-targets clean on the
    three touched crates

The CLI ORT (Linux/Windows) CI jobs — the ones that set
NXRT_REQUIRE_ORT_TESTS=1 — fail on main already, so the ORT suites above were
run locally against libonnxruntime.so 1.27 rather than relied on from CI.

The rule is that selecting this EP keeps work off ORT's CPU EP. The
perf-based decline was already withdrawn, but auditing what that
guaranteed turned up three ways attention and MoE nodes were still
reaching ORT, none of which any existing test could see.

`ASSIGNMENT_FIXTURES` covered 23 activation and normalisation graphs and
zero attention, MoE, KV-cache, Softmax, Transpose or RoPE graphs: the
guarantee was asserted where it was easy, not where it mattered. Nine
fixtures close that, and the assignment test now collects every failure
and asserts once so a run reports the whole matrix.

With them wired in, a real ORT 1.27 session found:

  * `GetCapability`'s fail-closed shape filter silently declined
    `com.microsoft::Attention`, `MoE`, `QMoE`, `PackedMultiHeadAttention`,
    `ScatterND`, `ScatterElements` and `Trilu`. ORT has no CPU kernel for
    three of those, so the "fallback" bought a load failure, not a
    slower run. All seven now have shape rules.

  * The per-op dtype *union* declined `RotaryEmbedding` and
    `GroupQueryAttention` outright, because their integer slots
    (`position_ids` int64, `seqlens_k` / `total_sequence_length` int32)
    were tested against `FLOAT_DTYPES`. `input_dtype_constraints_for_op`
    existed for the opposite problem and simply had no attention
    entries; the two RoPE domains need separate tables because their
    slot orders differ.

  * ORT stamps schema defaults onto a node before an EP sees it, and the
    contrib default for `smooth_softmax` is -1, not 0. Testing it for
    `!= 0` refused every GQA node ORT ever resolved. `scale` is the same
    hazard in the other direction: both ORT's kernels and ours read 0 as
    "use 1/sqrt(head_size)", so a stamped zero taken literally would
    zero every score silently.

All 32 fixtures now load and are assigned to this EP with
`session.disable_cpu_ep_fallback=1`. Assignment is not the guarantee,
so the two recovered ops are also checked against the kernels they
displaced: same model, same input bytes, our EP versus ORT's own contrib
kernels. RoPE 1.18e-7 and GQA 1.01e-7 max relative error, with
`present_key` and `present_value` bit-identical.

Docs: §10's and §15's "defer to ORT" conclusions are marked withdrawn
and §23 records the rule, the audit and the matrices. The measurements
stand; they are now a work queue rather than a justification.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 78.07971% with 121 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.80%. Comparing base (857dc8b) to head (2a704fc).

Files with missing lines Patch % Lines
crates/onnx-runtime-ep-plugin/src/compute.rs 21.11% 69 Missing and 2 partials ⚠️
...untime-ep-cpu/src/kernels/group_query_attention.rs 77.33% 13 Missing and 4 partials ⚠️
.../onnx-runtime-ep-cpu/src/kernels/msft_attention.rs 84.72% 10 Missing and 1 partial ⚠️
crates/onnx-runtime-ep-cpu/src/kernels/qmoe.rs 78.12% 3 Missing and 4 partials ⚠️
...rates/onnx-runtime-ep-cpu/src/kernels/attention.rs 88.09% 0 Missing and 5 partials ⚠️
...-ep-cpu/src/kernels/packed_multi_head_attention.rs 82.75% 0 Missing and 5 partials ⚠️
crates/onnx-runtime-ep-cpu/src/kernels/moe.rs 71.42% 3 Missing and 1 partial ⚠️
...runtime-ep-cpu/src/kernels/multi_head_attention.rs 97.29% 0 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1129      +/-   ##
==========================================
+ Coverage   79.95%   80.80%   +0.85%     
==========================================
  Files         368      368              
  Lines      160934   161233     +299     
  Branches   160934   161233     +299     
==========================================
+ Hits       128670   130288    +1618     
+ Misses      27542    26214    -1328     
- Partials     4722     4731       +9     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.40% <ø> (+0.09%) ⬆️
mlas 85.80% <ø> (+0.16%) ⬆️
offline 80.60% <78.07%> (+0.89%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-ep-cpu/src/kernels/mod.rs 95.01% <100.00%> (+0.02%) ⬆️
crates/onnx-runtime-ep-cpu/src/provider.rs 92.62% <100.00%> (+2.49%) ⬆️
...runtime-ep-cpu/src/kernels/multi_head_attention.rs 79.80% <97.29%> (+0.30%) ⬆️
crates/onnx-runtime-ep-cpu/src/kernels/moe.rs 86.27% <71.42%> (+0.43%) ⬆️
...rates/onnx-runtime-ep-cpu/src/kernels/attention.rs 87.93% <88.09%> (+0.05%) ⬆️
...-ep-cpu/src/kernels/packed_multi_head_attention.rs 75.26% <82.75%> (+6.91%) ⬆️
crates/onnx-runtime-ep-cpu/src/kernels/qmoe.rs 84.48% <78.12%> (+0.15%) ⬆️
.../onnx-runtime-ep-cpu/src/kernels/msft_attention.rs 80.31% <84.72%> (+0.31%) ⬆️
...untime-ep-cpu/src/kernels/group_query_attention.rs 92.17% <77.33%> (+0.33%) ⬆️
crates/onnx-runtime-ep-plugin/src/compute.rs 78.92% <21.11%> (-1.02%) ⬇️

... and 13 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 reduce_mean/large_f32_threads=1-internal/262144 994.29 µs 1.50 ms +50.4%
🔴 gather/medium_bf16_threads=1-internal/32768 2.46 µs 3.68 µs +49.8%
🔴 gather/medium_f32_threads=1-internal/32768 4.54 µs 6.62 µs +45.8%
⚠️ add/medium_bf16_threads=1-internal/262144 102.96 µs 133.24 µs +29.4%
⚠️ gather/medium_f16_threads=1-internal/32768 2.95 µs 3.79 µs +28.5%
⚠️ add/medium_f16_threads=1-internal/262144 110.90 µs 140.33 µs +26.5%
⚠️ gather/small_bf16_threads=1-internal/4096 473.7 ns 599.3 ns +26.5%
⚠️ matmul/small_generic_f16_threads=8/1x256x256 35.30 µs 43.24 µs +22.5%
⚠️ reduce_mean/medium_f32_threads=1-internal/65536 253.17 µs 308.61 µs +21.9%
⚠️ gather/large_bf16_threads=1-internal/131072 11.88 µs 14.19 µs +19.5%
⚠️ matmul/small_generic_f32_threads=1/1x256x256 37.92 µs 44.82 µs +18.2%
⚠️ gather/large_f16_threads=1-internal/131072 12.45 µs 14.39 µs +15.5%
✅ gather/large_f32_threads=1-internal/131072 27.88 µs 31.97 µs +14.7%
✅ add/large_f32_threads=1-internal/4194304 721.11 µs 815.72 µs +13.1%
✅ add/large_f16_threads=1-internal/4194304 1.73 ms 1.92 ms +10.9%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.98 µs 16.27 µs +8.6%
✅ sampling_latency/min_p_per_token 242.25 µs 260.09 µs +7.4%
✅ gather/small_f16_threads=1-internal/4096 480.1 ns 512.4 ns +6.7%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.12 ms 4.32 ms +4.8%
✅ gather/small_f32_threads=1-internal/4096 680.1 ns 710.8 ns +4.5%
✅ tokenization/decode_tokens_per_second 7.73 ms 8.06 ms +4.2%
✅ add/small_bf16_threads=1-internal/1024 500.3 ns 511.2 ns +2.2%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 587.93 µs 589.84 µs +0.3%
✅ qwen3_sampling_processors/top_k_partial_selection 163.61 µs 161.54 µs -1.3%
✅ logit_processing/seven_processor_chain_per_step 375.39 µs 370.19 µs -1.4%
✅ matmul/small_generic_f16_threads=1/1x256x256 32.36 µs 31.90 µs -1.4%
✅ qwen3_sampling_processors/top_k_top_p_fast 765.80 µs 754.82 µs -1.4%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 487.48 µs 475.39 µs -2.5%
✅ sampling_latency/greedy_per_token 3.70 µs 3.58 µs -3.3%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.63 ms 6.40 ms -3.3%
✅ sampling_latency/top_p_per_token 451.13 µs 434.00 µs -3.8%
✅ add/medium_f32_threads=1-internal/262144 30.23 µs 28.79 µs -4.8%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.97 ms 9.38 ms -5.9%
✅ kv_cache/alloc_dealloc_pages 46.66 µs 43.77 µs -6.2%
✅ add/large_bf16_threads=1-internal/4194304 1.93 ms 1.81 ms -6.5%
✅ matmul/small_generic_bf16_threads=1/1x256x256 41.87 µs 39.12 µs -6.6%
✅ add/small_f16_threads=1-internal/1024 565.8 ns 527.0 ns -6.9%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.54 ms 2.35 ms -7.4%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 182.34 µs 168.44 µs -7.6%
✅ sampling_latency/top_k_per_token 64.38 µs 59.41 µs -7.7%
✅ tokenization/encode_tokens_per_second 449.83 µs 405.70 µs -9.8%
✅ matmul/small_generic_bf16_threads=8/1x256x256 46.19 µs 41.18 µs -10.8%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.25 ms 2.00 ms -11.2%
✅ matmul/medium_generic_f16_threads=1/32x512x512 36.62 µs 32.36 µs -11.6%
✅ grammar_masking/llguidance_compute_mask/32 90.97 µs 79.14 µs -13.0%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 632.76 µs 547.65 µs -13.5%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 56.14 µs 47.81 µs -14.8%
🟢 matmul/large_generic_f16_threads=1/32x1024x1024 95.60 µs 80.71 µs -15.6%
🟢 add/small_f32_threads=1-internal/1024 278.6 ns 231.8 ns -16.8%
🟢 matmul/medium_generic_f32_threads=1/32x512x512 3.27 ms 2.69 ms -17.8%
🟢 matmul/large_generic_f32_threads=8/32x1024x1024 4.72 ms 3.83 ms -18.9%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 61.38 µs 49.13 µs -20.0%
🟢 matmul/small_generic_f32_threads=8/1x256x256 53.78 µs 41.91 µs -22.1%
🟢 matmul/large_generic_f16_threads=8/32x1024x1024 118.00 µs 87.34 µs -26.0%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 1.77 ms 1.30 ms -26.5%
🟢 matmul/medium_generic_bf16_threads=8/32x512x512 526.08 µs 374.79 µs -28.8%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 45.77 µs 31.24 µs -31.7%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.94 ms 1.28 ms -34.1%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.00 ms 570.21 µs -43.0%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 6.32 4.75 6.08 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby and others added 2 commits August 17, 2026 16:50
Review asked for one real-ORT fixture per rescued op rather than a
representative sample, and three of the four extra fixtures immediately
found a decline that was still live:

  * `GroupQueryAttention`'s `position_ids` is optional input *9*. The
    first pass listed only slots 5 and 6 -- the ones the fixtures
    happened to exercise -- so a `do_rotary` node with explicit int64
    positions still failed the float union and went to ORT. Our kernel
    runs that config; `rotary_explicit_position_ids_apply_to_query_and_key`
    has covered it all along.

  * `QMoE`'s uint8-packed expert weights and zero points, and
    `PackedMultiHeadAttention`'s int32 `token_offset` /
    `cumulative_sequence_length`, failed the same union. ORT has no CPU
    kernel for either, so declining bought `Could not find an
    implementation for PackedMultiHeadAttention(1)`.

The inventory test cannot see any of this: it builds synthetic nodes and
never opens a session, so it is blind to both the dtype filter and the
kernel factory. That is the third time that has cost something.

Also from review:

  * The `scale` rationale was wrong and is corrected in place. ORT does
    *not* stamp `scale` -- it is OPTIONAL_VALUE with no schema default,
    and instrumenting a real CreateSession prints `scale_attr=None` for
    GQA, MHA and `com.microsoft::Attention`, while `smooth_softmax`
    prints `Some(-1)` on a node that never set it. The guard stays as
    defence, not as a fix, and now also covers
    `packed_multi_head_attention.rs`, which this change made reachable
    and which was the one attention kernel left without it.

  * §23.4 claimed ORT has no CPU kernel for `com.microsoft::Attention`.
    It does. The claim is true only of MoE, QMoE and PackedMHA.

`QMoE`'s fixture needs `block_size = 32`: our kernel implements only the
blocked form, and ORT's schema default of 0 selects the column-wise one.
That is a real capability gap, recorded in §23.6 rather than papered
over -- and not one ORT can cover either.

37/37 fixtures assigned to this EP with `session.disable_cpu_ep_fallback=1`.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ot run

Round 2 of independent review found three problems with the previous commit.

`MoE` advertised float32 only while its kernel widens float16 and bfloat16 to
f32 and narrows on the way out, so the f32 fixture passed while every realistic
half-precision mixture was still handed to ORT. Its union is now
`FLOAT_COMPUTE_DTYPES`, with `moe_assignment_f16` to prove it.

Making `QMoE` reachable was a regression. Its factory rejects `block_size = 0`
-- the column-wise form, which is what an absent attribute means, and which ORT
does run on CPU -- but a factory rejection lands *after* ORT has compiled the
node onto this EP, so a model that used to work died at `CreateSession` with no
fallback. `qmoe::unsupported_reason` now mirrors that limit in `supports_op`,
where a decline is still recoverable, and
`column_wise_qmoe_is_declined_at_claim_time_not_failed_late` pins it. This is a
capability gap we owe a kernel for, not a performance decline.

Finally, "ORT has no CPU kernel for MoE or QMoE" was false; review loaded and
ran these fixtures on ORT's CPUExecutionProvider under 1.27 and 1.28. Only
`PackedMultiHeadAttention` genuinely has none. Corrected in both docs, the
shape table and the fixture generator.

38/38 fixtures load and are assigned to us with ORT CPU fallback disabled.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Verified your central claim independently on main before reading the fix, and it is worse than the sentence conveys.

crates/onnx-runtime-ep-cpu-plugin/tests/plugin_ort_e2e.rs has 32 assignment fixtures. Counting operator mentions in that file:

operator mentions
MatMulNBits 23
Attention 0
GroupQueryAttention 0
MoE 0
Softmax 0
Transpose 0
RotaryEmbedding 0

So the suite that exists to prove node ownership covers none of the operators this EP exists for. It is the same defect shape as several others found this week: the check was present, was green, and could not fail on the case that mattered -- like a decline test built around a single instantiation (#1100), or a memory invariant asserted through a proxy field (#1027), or my own CI guard whose first regex matched neither of the two commits that motivated it.

That is why "23 activation and normalisation graphs and zero attention graphs" is the most valuable line in this PR. The kernels are the fix; the fixture gap is the reason nobody knew there was one.

The rule you are implementing is the owner's, verbatim: 用了我们的 cpu ep 就不要用 ort 的 cpu ep。但凡比 ort 慢的都要想方设法比 ort 快 -- a losing range is a kernel to fix, not a node to give away. Worth stating explicitly that this cuts both ways: once we claim these nodes we own their performance, so each newly-claimed op needs a measured comparison against ORT's kernel on the same graph, and a plan where we lose.

I have not run the validation yet -- another agent is building on this box for #1138 and a plugin/ORT run needs the machine, so any timing I took now would be unusable (I discarded eight consecutive runs on #1116 for exactly this reason). What I would want before merge:

  1. The parity evidence you mention, per newly-claimed op, against a real ORT session with CPU fallback disabled -- that disabled-fallback detail is what makes the claim falsifiable rather than a green test that would pass either way.
  2. A per-op ours/ORT ratio for the ops this newly claims, so we know where we are now losing on purpose. Claiming a node we run slower than ORT is still an improvement in ownership terms, but it should be a recorded debt, not a surprise.
  3. A guard that the fixture list keeps covering the EP's own operator set -- a test asserting every op in the kernel registry's transformer group appears in ASSIGNMENT_FIXTURES would have prevented this gap and will prevent the next one. Coverage of a hand-maintained list decays silently; a list derived from the registry cannot.

Ping me when the box is free and I will validate.

justinchuby and others added 2 commits August 17, 2026 20:18
The column-wise QMoE fix was a point patch of a class. A third review round
found two more live instances, both on attributes production models set and
both of which ORT's CPU EP runs today: `use_sparse_mixer=1` (Phi-3.5-MoE,
GRIN-MoE) and `smooth_softmax=1` (Gemma-style attention sink). In each case we
claimed the node and the kernel factory then rejected it -- which lands after
ORT has compiled the node onto this EP, where no fallback recovers, so a model
that ran fine died at `CreateSession`.

The fix is structural rather than three more conditions. Each kernel's attribute
validation now lives in one function and the claim-time guard *is* that
function's error, so a limit cannot be added to a factory without appearing at
claim time too:

    pub(crate) fn unsupported_reason(node: &Node) -> Option<String> {
        attributes_from_node(node).err().map(|e| e.to_string())
    }

`supports_op` consults it for MoE, QMoE, GroupQueryAttention,
MultiHeadAttention, com.microsoft::Attention, ai.onnx::Attention and
PackedMultiHeadAttention, which also sweeps up GQA's quantized-KV and qk_output
rejections, msft Attention's do_rotary and past_present_share_buffer,
ai.onnx::Attention's qk_matmul_output_mode and QMoE's expert_weight_bits /
quant_type.

Two tests pin it. `every_factory_attribute_rejection_is_mirrored_at_claim_time`
is pure Rust and asserts, for eleven hostile nodes, both that the factory
rejects and that `supports_op` declines, so it fails if they ever diverge.
`factory_only_capability_limits_are_declined_at_claim_time` proves it against
real ORT for three fixtures; disabling the guard turns each into a hard
`CreateSession` failure, which is what the falsifier run showed.

These remain capability declines, never performance ones: each is a kernel we
owe, and each should stop being an exception as soon as it exists.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Round-4 review found the `PackedMultiHeadAttention` row of
`every_factory_attribute_rejection_is_mirrored_at_claim_time` passing for the
wrong reason: its claim-time check inspects inputs, shapes and dtypes before it
reaches the attribute mirror, so an attribute-only node declined with "Q, K, V,
token_offset, and cumulative_sequence_length are required" and the row still
passed with the mirror deleted. It now builds a fully-formed node and asserts
the decline reason is the scale one; removing the mirror fails it.

Also narrows a doc claim. "Drift is impossible by construction" is true within
a wired op, but wiring an op is still discipline, and review's EP-wide audit
found ten other factories rejecting things `supports_op` does not pre-check.
None is a live regression -- each refuses only schema-invalid values or configs
ORT's own CPU kernel also refuses -- and that audit is now recorded so it does
not have to be redone.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 17, 2026 21:42
@justinchuby
justinchuby merged commit 8d4401d into main Aug 17, 2026
13 of 18 checks passed
@justinchuby
justinchuby deleted the squad/leon-p5-no-decline branch August 17, 2026 23:00
justinchuby added a commit that referenced this pull request Aug 18, 2026
…den (#1150)

Lands DeepSeek-V2-Lite as **supported + numerically correct** on the
native CUDA EP. Combines two reviewer-gated artifacts (planner fix
authored by Leon, oracle correction + tests authored by Luv; both gated
🟢 by Rachael).

## What
1. **MoE workspace planner shape-resolution** (`executor/bindings.rs`):
the real V2-Lite blocker was a MoE-gate
`Reshape([-1,hidden])→Cast→MatMul` whose `batch*sequence` was not
symbol-bound at workspace-reservation time. Fix walks deterministic
`Cast/CastLike/Identity/Reshape` producer chains to resolve runtime
shapes before reservation. Proven dense-safe (Qwen3 reservation
byte-identical `persistent=983040 step=0`; recovery only fires for V2
MoE gates).
2. **Golden lock rebased to the native-CUDA stream**
(`deepseek_v2_lite_decode_lock.rs`): CPU `…207,17,15,1012` → native-CUDA
`…207,16,24,1012`. This is a **correctness fix, not a regression**: vs
an f64 reference, native CUDA is *closer to truth* than the CPU fold on
every measured QMoE/dense case (e.g. dense int8 block32: CUDA/f64
1.26e-6 vs CPU/f64 3.60e-4 = 285×). The prior CPU lock pinned to the
*less* accurate stream. Golden verified identical on graphs-OFF (GPU0)
and graphs-ON (GPU1).
3. **f64-bounded QMoE test**
(`qmoe_gpu.rs::qmoe_int4_identity_expert_gemv_within_f64_roundoff`):
drives the real QMoE kernel and asserts CUDA within the f64
tree-reduction roundoff bound.
4. **Env-gated router probe** `ONNX_GENAI_QMOE_ROUTE_DUMP` (default-OFF)
for future divergence forensics.

## Why the CPU≠CUDA divergence is benign
Token-5 divergence is f32 accumulation-ORDER drift (CPU sequential fold
vs CUDA 256-lane striped + 8-step tree reduction), not a bug. The
token-5 expert-set swap (CPU expert 61 / CUDA expert 1 in the 6th top-k
slot) is a below-fp32-resolution near-tie: router-logit delta ~5e-5 <
reassociation drift ~4.7–6.4e-5. Rachael independently judged an epsilon
tie-break arbitrary (would risk re-pinning to the less-accurate CPU
stream) and accepted the boundary.

## Verification (rebased onto current main 11043a0)
- `deepseek_v2_lite_decode_lock`: PASS graphs OFF (GPU0) + ON (GPU1);
golden unchanged through #1129's MoE-claiming change
- `qmoe_gpu`: 30/30 PASS · new f64 test PASS
- `matmul_nbits_gpu` f64 dense: 5/5 PASS
- `cargo fmt --check` clean · `cargo build --release --features cuda`
clean

Reviewer gate: Rachael 🟢 (planner regression + oracle re-gate). Author
lockout honored throughout (planner author distinct from oracle author).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant