Repository navigation
fix(cpu-ep): stop handing attention and MoE nodes to ORT's CPU EP - #1129
Conversation
The rule is that selecting this EP keeps work off ORT's CPU EP. The
perf-based decline was already withdrawn, but auditing what that
guaranteed turned up three ways attention and MoE nodes were still
reaching ORT, none of which any existing test could see.
`ASSIGNMENT_FIXTURES` covered 23 activation and normalisation graphs and
zero attention, MoE, KV-cache, Softmax, Transpose or RoPE graphs: the
guarantee was asserted where it was easy, not where it mattered. Nine
fixtures close that, and the assignment test now collects every failure
and asserts once so a run reports the whole matrix.
With them wired in, a real ORT 1.27 session found:
* `GetCapability`'s fail-closed shape filter silently declined
`com.microsoft::Attention`, `MoE`, `QMoE`, `PackedMultiHeadAttention`,
`ScatterND`, `ScatterElements` and `Trilu`. ORT has no CPU kernel for
three of those, so the "fallback" bought a load failure, not a
slower run. All seven now have shape rules.
* The per-op dtype *union* declined `RotaryEmbedding` and
`GroupQueryAttention` outright, because their integer slots
(`position_ids` int64, `seqlens_k` / `total_sequence_length` int32)
were tested against `FLOAT_DTYPES`. `input_dtype_constraints_for_op`
existed for the opposite problem and simply had no attention
entries; the two RoPE domains need separate tables because their
slot orders differ.
* ORT stamps schema defaults onto a node before an EP sees it, and the
contrib default for `smooth_softmax` is -1, not 0. Testing it for
`!= 0` refused every GQA node ORT ever resolved. `scale` is the same
hazard in the other direction: both ORT's kernels and ours read 0 as
"use 1/sqrt(head_size)", so a stamped zero taken literally would
zero every score silently.
All 32 fixtures now load and are assigned to this EP with
`session.disable_cpu_ep_fallback=1`. Assignment is not the guarantee,
so the two recovered ops are also checked against the kernels they
displaced: same model, same input bytes, our EP versus ORT's own contrib
kernels. RoPE 1.18e-7 and GQA 1.01e-7 max relative error, with
`present_key` and `present_value` bit-identical.
Docs: §10's and §15's "defer to ORT" conclusions are marked withdrawn
and §23 records the rule, the audit and the matrices. The measurements
stand; they are now a work queue rather than a justification.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Review asked for one real-ORT fixture per rescued op rather than a
representative sample, and three of the four extra fixtures immediately
found a decline that was still live:
* `GroupQueryAttention`'s `position_ids` is optional input *9*. The
first pass listed only slots 5 and 6 -- the ones the fixtures
happened to exercise -- so a `do_rotary` node with explicit int64
positions still failed the float union and went to ORT. Our kernel
runs that config; `rotary_explicit_position_ids_apply_to_query_and_key`
has covered it all along.
* `QMoE`'s uint8-packed expert weights and zero points, and
`PackedMultiHeadAttention`'s int32 `token_offset` /
`cumulative_sequence_length`, failed the same union. ORT has no CPU
kernel for either, so declining bought `Could not find an
implementation for PackedMultiHeadAttention(1)`.
The inventory test cannot see any of this: it builds synthetic nodes and
never opens a session, so it is blind to both the dtype filter and the
kernel factory. That is the third time that has cost something.
Also from review:
* The `scale` rationale was wrong and is corrected in place. ORT does
*not* stamp `scale` -- it is OPTIONAL_VALUE with no schema default,
and instrumenting a real CreateSession prints `scale_attr=None` for
GQA, MHA and `com.microsoft::Attention`, while `smooth_softmax`
prints `Some(-1)` on a node that never set it. The guard stays as
defence, not as a fix, and now also covers
`packed_multi_head_attention.rs`, which this change made reachable
and which was the one attention kernel left without it.
* §23.4 claimed ORT has no CPU kernel for `com.microsoft::Attention`.
It does. The claim is true only of MoE, QMoE and PackedMHA.
`QMoE`'s fixture needs `block_size = 32`: our kernel implements only the
blocked form, and ORT's schema default of 0 selects the column-wise one.
That is a real capability gap, recorded in §23.6 rather than papered
over -- and not one ORT can cover either.
37/37 fixtures assigned to this EP with `session.disable_cpu_ep_fallback=1`.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ot run Round 2 of independent review found three problems with the previous commit. `MoE` advertised float32 only while its kernel widens float16 and bfloat16 to f32 and narrows on the way out, so the f32 fixture passed while every realistic half-precision mixture was still handed to ORT. Its union is now `FLOAT_COMPUTE_DTYPES`, with `moe_assignment_f16` to prove it. Making `QMoE` reachable was a regression. Its factory rejects `block_size = 0` -- the column-wise form, which is what an absent attribute means, and which ORT does run on CPU -- but a factory rejection lands *after* ORT has compiled the node onto this EP, so a model that used to work died at `CreateSession` with no fallback. `qmoe::unsupported_reason` now mirrors that limit in `supports_op`, where a decline is still recoverable, and `column_wise_qmoe_is_declined_at_claim_time_not_failed_late` pins it. This is a capability gap we owe a kernel for, not a performance decline. Finally, "ORT has no CPU kernel for MoE or QMoE" was false; review loaded and ran these fixtures on ORT's CPUExecutionProvider under 1.27 and 1.28. Only `PackedMultiHeadAttention` genuinely has none. Corrected in both docs, the shape table and the fixture generator. 38/38 fixtures load and are assigned to us with ORT CPU fallback disabled. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Verified your central claim independently on
So the suite that exists to prove node ownership covers none of the operators this EP exists for. It is the same defect shape as several others found this week: the check was present, was green, and could not fail on the case that mattered -- like a decline test built around a single instantiation (#1100), or a memory invariant asserted through a proxy field (#1027), or my own CI guard whose first regex matched neither of the two commits that motivated it. That is why "23 activation and normalisation graphs and zero attention graphs" is the most valuable line in this PR. The kernels are the fix; the fixture gap is the reason nobody knew there was one. The rule you are implementing is the owner's, verbatim: 用了我们的 cpu ep 就不要用 ort 的 cpu ep。但凡比 ort 慢的都要想方设法比 ort 快 -- a losing range is a kernel to fix, not a node to give away. Worth stating explicitly that this cuts both ways: once we claim these nodes we own their performance, so each newly-claimed op needs a measured comparison against ORT's kernel on the same graph, and a plan where we lose. I have not run the validation yet -- another agent is building on this box for #1138 and a plugin/ORT run needs the machine, so any timing I took now would be unusable (I discarded eight consecutive runs on #1116 for exactly this reason). What I would want before merge:
Ping me when the box is free and I will validate. |
The column-wise QMoE fix was a point patch of a class. A third review round
found two more live instances, both on attributes production models set and
both of which ORT's CPU EP runs today: `use_sparse_mixer=1` (Phi-3.5-MoE,
GRIN-MoE) and `smooth_softmax=1` (Gemma-style attention sink). In each case we
claimed the node and the kernel factory then rejected it -- which lands after
ORT has compiled the node onto this EP, where no fallback recovers, so a model
that ran fine died at `CreateSession`.
The fix is structural rather than three more conditions. Each kernel's attribute
validation now lives in one function and the claim-time guard *is* that
function's error, so a limit cannot be added to a factory without appearing at
claim time too:
pub(crate) fn unsupported_reason(node: &Node) -> Option<String> {
attributes_from_node(node).err().map(|e| e.to_string())
}
`supports_op` consults it for MoE, QMoE, GroupQueryAttention,
MultiHeadAttention, com.microsoft::Attention, ai.onnx::Attention and
PackedMultiHeadAttention, which also sweeps up GQA's quantized-KV and qk_output
rejections, msft Attention's do_rotary and past_present_share_buffer,
ai.onnx::Attention's qk_matmul_output_mode and QMoE's expert_weight_bits /
quant_type.
Two tests pin it. `every_factory_attribute_rejection_is_mirrored_at_claim_time`
is pure Rust and asserts, for eleven hostile nodes, both that the factory
rejects and that `supports_op` declines, so it fails if they ever diverge.
`factory_only_capability_limits_are_declined_at_claim_time` proves it against
real ORT for three fixtures; disabling the guard turns each into a hard
`CreateSession` failure, which is what the falsifier run showed.
These remain capability declines, never performance ones: each is a kernel we
owe, and each should stop being an exception as soon as it exists.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Round-4 review found the `PackedMultiHeadAttention` row of `every_factory_attribute_rejection_is_mirrored_at_claim_time` passing for the wrong reason: its claim-time check inspects inputs, shapes and dtypes before it reaches the attribute mirror, so an attribute-only node declined with "Q, K, V, token_offset, and cumulative_sequence_length are required" and the row still passed with the mirror deleted. It now builds a fully-formed node and asserts the decline reason is the scale one; removing the mirror fails it. Also narrows a doc claim. "Drift is impossible by construction" is true within a wired op, but wiring an op is still discipline, and review's EP-wide audit found ten other factories rejecting things `supports_op` does not pre-check. None is a live regression -- each refuses only schema-invalid values or configs ORT's own CPU kernel also refuses -- and that audit is now recorded so it does not have to be redone. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…den (#1150) Lands DeepSeek-V2-Lite as **supported + numerically correct** on the native CUDA EP. Combines two reviewer-gated artifacts (planner fix authored by Leon, oracle correction + tests authored by Luv; both gated 🟢 by Rachael). ## What 1. **MoE workspace planner shape-resolution** (`executor/bindings.rs`): the real V2-Lite blocker was a MoE-gate `Reshape([-1,hidden])→Cast→MatMul` whose `batch*sequence` was not symbol-bound at workspace-reservation time. Fix walks deterministic `Cast/CastLike/Identity/Reshape` producer chains to resolve runtime shapes before reservation. Proven dense-safe (Qwen3 reservation byte-identical `persistent=983040 step=0`; recovery only fires for V2 MoE gates). 2. **Golden lock rebased to the native-CUDA stream** (`deepseek_v2_lite_decode_lock.rs`): CPU `…207,17,15,1012` → native-CUDA `…207,16,24,1012`. This is a **correctness fix, not a regression**: vs an f64 reference, native CUDA is *closer to truth* than the CPU fold on every measured QMoE/dense case (e.g. dense int8 block32: CUDA/f64 1.26e-6 vs CPU/f64 3.60e-4 = 285×). The prior CPU lock pinned to the *less* accurate stream. Golden verified identical on graphs-OFF (GPU0) and graphs-ON (GPU1). 3. **f64-bounded QMoE test** (`qmoe_gpu.rs::qmoe_int4_identity_expert_gemv_within_f64_roundoff`): drives the real QMoE kernel and asserts CUDA within the f64 tree-reduction roundoff bound. 4. **Env-gated router probe** `ONNX_GENAI_QMOE_ROUTE_DUMP` (default-OFF) for future divergence forensics. ## Why the CPU≠CUDA divergence is benign Token-5 divergence is f32 accumulation-ORDER drift (CPU sequential fold vs CUDA 256-lane striped + 8-step tree reduction), not a bug. The token-5 expert-set swap (CPU expert 61 / CUDA expert 1 in the 6th top-k slot) is a below-fp32-resolution near-tie: router-logit delta ~5e-5 < reassociation drift ~4.7–6.4e-5. Rachael independently judged an epsilon tie-break arbitrary (would risk re-pinning to the less-accurate CPU stream) and accepted the boundary. ## Verification (rebased onto current main 11043a0) - `deepseek_v2_lite_decode_lock`: PASS graphs OFF (GPU0) + ON (GPU1); golden unchanged through #1129's MoE-claiming change - `qmoe_gpu`: 30/30 PASS · new f64 test PASS - `matmul_nbits_gpu` f64 dense: 5/5 PASS - `cargo fmt --check` clean · `cargo build --release --features cuda` clean Reviewer gate: Rachael 🟢 (planner regression + oracle re-gate). Author lockout honored throughout (planner author distinct from oracle author). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
What
Selecting this CPU EP is supposed to keep work off ORT's CPU EP. It did not, for
exactly the operators this EP exists for. This PR makes attention, MoE and
KV-cache nodes actually reach our kernels, proves it against a real ORT 1.27
session with fallback disabled, and checks the two recovered ops produce ORT's
numbers.
The project rule this implements:
A losing range is a kernel to fix, not a node to give away.
Methodology
The perf-based decline (
assignment_policy.rs) was already deleted andclaim_preference_nodealready returnsClaimunconditionally, so the questionwas what that actually guaranteed. Answer: nothing for attention.
plugin_ort_e2e'sASSIGNMENT_FIXTUREScovered 23 activation and normalisationgraphs and zero attention, MoE, KV-cache, Softmax, Transpose or RoPE graphs.
The rule was asserted where it was easy. Fourteen fixtures close that gap
(
generate_attention_fixtures.py), andno_supported_node_is_ever_left_to_the_ort_cpu_epnow collects every failure andasserts once, so a run reports the whole matrix instead of stopping at the first
decline.
Every measurement below is a real ORT 1.27 CPU session with
session.record_ep_graph_assignment_info=1, read back throughSession_GetEpGraphAssignmentInfo.Evidence
Three ways a node was still reaching ORT
1.
GetCapability's fail-closed shape filter — the silent layer. Threeindependent gates must all pass. Deleting the policy addressed one; this third
one answers to neither of the others and was giving away
com.microsoft::Attention,MoE,QMoE,PackedMultiHeadAttention,ScatterND,ScatterElementsandTrilu. ORT has no CPU kernel for MoE,QMoE or PackedMHA, so the "fallback" bought a load failure, not a slower run.
All seven now have rules;
shape_inference_coverage.rsasserts the allowlistexactly in both directions, so this cannot silently regrow.
2. The per-op dtype union declined the two most important decode ops. With
the shape rules in, the first real run said:
The plugin advertises one dtype set per op and tests every input slot against
it. Attention ops map to
FLOAT_DTYPES, butRotaryEmbedding'sposition_idsis int64 and
GroupQueryAttention'sseqlens_k/total_sequence_lengthareint32, so the integer slots failed the float test and the claim was dropped.
input_dtype_constraints_for_opalready existed for the opposite problem (aunion too wide for
MatMulNBits) and simply had no attention entries. Thecom.microsoftandai.onnxRoPE slot orders differ, so they need separatetables.
Consequence worth flagging: any plugin-path RoPE or GQA measurement taken
before this fix — including #1078's — was measuring ORT, not us.
3. ORT stamps schema defaults on the node, and one of them is not 0. With
the dtypes fixed, GQA still failed, at kernel construction:
The fixture never sets it. The contrib schema default is -1; ORT's own kernel
enables the feature only for the exact value
1, so-1means off. Our!= 0test refused every GQA node ORT ever resolved.
scaleis the same hazardinverted — ORT's kernels and ours both read
0as "use1/sqrt(head_size)", soa stamped zero taken literally would multiply every score by zero and return a
silently wrong answer instead of an error. Fixed in
attention.rs,msft_attention.rs,multi_head_attention.rsandgroup_query_attention.rs.Assignment matrix — all 37 fixtures, never partial
softmax_assignment_f32Softmaxtranspose_assignment_f32Transposekv_concat_assignment_f32Concatkv_scatternd_assignment_f32ScatterNDscatter_elements_assignment_f32ScatterElementstrilu_assignment_f32Trilurotary_assignment_f32RotaryEmbeddingmha_assignment_f32MultiHeadAttentiongqa_assignment_f32GroupQueryAttentiongqa_rotary_pos_assignment_f32GroupQueryAttention+ int64position_idsmsft_attention_assignment_f32com.microsoft::Attentionpacked_mha_assignment_f32PackedMultiHeadAttentionmoe_assignment_f32MoEmoe_assignment_f16MoEfloat16qmoe_assignment_f32QMoEevery_fixture_loads_with_cpu_fallback_disabledruns the same 38 withsession.disable_cpu_ep_fallback=1, so ORT is forbidden from placing asupported node on its own CPU EP: 38/38 load, 38/38 assigned to us. For
PackedMultiHeadAttentionthat is also the only way the graph loads at all —ORT has no CPU kernel for it, so declining it produced
Could not find an implementation for PackedMultiHeadAttention(1)rather than a slower run.Three fixtures are deliberately outside that set —
qmoe_columnwise_f32,moe_sparse_mixer_f32andgqa_smooth_softmax_f32— because we have nokernel for those configurations. See Round 4.
Round 2: what independent review changed
The first round shipped nine fixtures — one per op I thought was affected.
Review asked for one per rescued op instead, and three of the four extra
fixtures immediately found a decline that was still live:
GroupQueryAttention+position_ids. That is optional input 9. Ilisted slots 5 and 6 because those were the slots my fixtures exercised, so a
do_rotarynode with explicit int64 positions still went to ORT. Our kernelruns that config and
rotary_explicit_position_ids_apply_to_query_and_keyhascovered it all along — the claim path just never reached it.
QMoE(uint8-packed experts and zero points) andPackedMultiHeadAttention(int32token_offset/cumulative_sequence_length) failed the same union.The lesson is the specific one:
shape_inference_coverage.rsbuilds syntheticnodes and never opens a session, so it is blind to both the dtype filter and
the kernel factory. Only a real session per op finds these. That is the third
time that has cost something in this file's history.
Two corrections review forced, which I would rather state than bury:
scalerationale was wrong. I claimed ORT stampsscale = 0. It doesnot —
scaleisOPTIONAL_VALUEwith no schema default. The reviewerinstrumented a real
CreateSessionover these fixtures and gotscale_attr=Nonefor GQA, MHA andcom.microsoft::Attention(whilesmooth_softmax_attr=Some(-1)on a node that never set it, so that halfholds), and removing the guard left the numerics unchanged at 1.006e-7. The
guard stays as defence, not a fix, is now described that way, and now also
covers
packed_multi_head_attention.rs— the one attention kernel this changemade reachable without it.
com.microsoft::Attention. It does(
contrib_ops/cpu/bert/attention.cc). Round 3 showed the remainder of thatstatement was wrong too.
Round 3: what the second independent review changed
Round 2's findings are all resolved. The second review found three more, one of
which was a regression this PR itself introduced.
MoEwas still declining float16 and bfloat16. Its kernel widens both tof32 and narrows on the way out, but
supported_dtypes_for_opadvertisedF32_ONLYandMoEhad no per-slot table, so the union decided it. The f32fixture passed while every production mixture — which is half precision —
went to ORT. Union is now
FLOAT_COMPUTE_DTYPES;moe_assignment_f16provesit. One dtype's worth of coverage is not coverage.
Column-wise
QMoEwas a regression, now fixed. ORT leavesblock_sizewithout a schema default, so an absent attribute means the column-wise form —
and ORT runs that on CPU; the reviewer loaded and ran it under 1.27 and
1.28. Before this PR we declined
QMoEand ORT ran such models. After Round2 we claimed it, and the kernel factory then rejected
block_size = 0.A factory rejection arrives after ORT has compiled the node onto this EP,
and no fallback recovers from it, so a previously-working model died at
CreateSession.qmoe::unsupported_reasonnow mirrors the factory's limit insupports_op,where a decline is still recoverable, and
column_wise_qmoe_is_declined_at_claim_time_not_failed_lateasserts the nodelands on ORT and the session loads. The general lesson is written into
CPU_ACTIVATION_GAPS.md: any capability limit a factory enforces must bemirrored at claim time, because claiming an op is not free.
To be explicit about the policy this PR exists to enforce: this is the one
deliberate decline in the suite, and it is a capability answer, not a
performance one. We have no column-wise implementation; the choice is between
ORT running the model and nobody running it. It should stop being an exception
as soon as that kernel exists, and it is not a precedent for declining ranges
where we are merely slower.
"ORT has no CPU kernel for MoE or QMoE" was false. The reviewer loaded
and ran these very fixtures on ORT's
CPUExecutionProvider. OnlyPackedMultiHeadAttentiongenuinely lacks one, which its falsifier showsdirectly: reverting
PACKED_MHA_SLOTSdoes not hand the node to ORT, it failssession creation. Corrected in four places —
compute.rs,plugin_ort_e2e.rs,the fixture generator and §23.1/§23.4/§23.6 — because "ORT can't run it
either" was being used to make a decline sound harmless when it was simply us
not running an operator we implement.
Round 4: what the third independent review changed
The reviewer reproduced every round-3 number and ran both falsifiers, then
found that the column-wise
QMoEfix was a point patch of a class. Two moreinstances were live, on attributes production models set, both of which ORT's
CPU EP runs today — proven end-to-end through the real-ORT harness, not
inferred:
MoE/QMoEuse_sparse_mixer=1STAGE [CreateSession] FAILED: ... MoE: use_sparse_mixer=1 is unsupportedGroupQueryAttentionsmooth_softmax=1STAGE [CreateSession] FAILED: ... GroupQueryAttention: smooth_softmax is not yet supportedRound 1 had narrowed the
smooth_softmaxguard from!= 0to== 1to stoprejecting ORT's
-1schema default — which fixed the default case and left thegenuinely-set case as an unrecoverable claim.
The fix is structural, not three more conditions. Each kernel's attribute
validation now lives in a single function, and the claim-time guard is that
function's error:
so a limit cannot be added to a factory without appearing at claim time —
drift is impossible by construction rather than by discipline.
supports_opconsults it for
MoE,QMoE,GroupQueryAttention,MultiHeadAttention,com.microsoft::Attention,ai.onnx::AttentionandPackedMultiHeadAttention, which also sweeps up the reviewer's MAJOR (GQA'squantized-KV and
qk_outputrejections) plus msftAttention'sdo_rotary/past_present_share_buffer,ai.onnx::Attention'sqk_matmul_output_modeandQMoE'sexpert_weight_bits/quant_type.Two tests pin it:
provider::tests::every_factory_attribute_rejection_is_mirrored_at_claim_time— pure Rust, so it runs everywhere. For eleven hostile nodes it asserts
both that the factory rejects and that
supports_opdeclines, failing ifthe two ever diverge.
plugin_ort_e2e::factory_only_capability_limits_are_declined_at_claim_time— real ORT, three fixtures: node lands on ORT, session loads. Disabling the
guard produces a hard
CreateSessionfailure for each of the three; I ranthat falsifier arm-by-arm and all three fire.
To restate the policy boundary, since this PR exists to enforce it: these are
the only deliberate declines in the suite and every one is a capability
answer, not a performance one. We have no column-wise, sparse-mixer or
smooth-softmax implementation; for each, the choice is between ORT running the
model and nobody running it. They are not exceptions to the "never hand a slow
range to ORT" rule — they are an admission that we owe three kernels, and each
stops being an exception the moment its kernel lands.
Round 4's reviewer found no blocking issue and reproduced every number:
1300 / 1318 / 231 / 45 tests, all four falsifiers firing, and — checking the
docs rather than the code — loading and running both new fixtures on ORT 1.28's
CPUExecutionProviderto confirm it really does runuse_sparse_mixer=1andsmooth_softmax=1. Two MINORs, both fixed in34584d70e:PackedMultiHeadAttentionrow of the pure-Rust test passed for thewrong reason — its claim-time check inspects inputs, shapes and dtypes
before reaching the attribute mirror, so an attribute-only node declined with
"Q, K, V, token_offset, and cumulative_sequence_length are required" and the
row still passed with the mirror deleted. It now builds a fully-formed node
and asserts the decline reason; deleting the mirror fails it.
wired op; wiring an op is still discipline. The reviewer audited the rest of
the EP and found ten other factories rejecting things
supports_opdoes notpre-check (
MatMulNBitsbits/block_size,Resize,Pad,GridSample,LpNormalization,Unique,BitShift,ConstantOfShape, twopkg.nxrtinternals). None is a live regression — each refuses only schema-invalid
values or configurations ORT's own CPU kernel also refuses. Both the claim and
that audit are now in the docs so it does not have to be redone.
The reviewer also independently reproduced every §23 number and ran each
falsifier — reverting the dtype rows, restoring
smooth_softmax != 0, forcingAttentionback toDeclined, and perturbing one baseline input to confirm thenumerics test is not vacuous (it failed at rel err 0.458). All fired as claimed.
Numerics — assignment is not the guarantee
A claimed node that computes the wrong answer is worse than the deferral it
replaced.
rope_and_gqa_execute_on_our_ep_and_match_ort_numericsruns eachfixture twice over the same model file with the same input bytes: once with our
EP appended and fallback disabled, once with our EP simply not appended so ORT
resolves the node to its own contrib CPU kernel.
rotary_assignment_f32outputgqa_assignment_f32outputgqa_assignment_f32present_keygqa_assignment_f32present_valueBoth attention outputs are at float32 rounding; the KV cache is bit-identical.
Limitations
This makes the losing ranges reachable. It does not make them fast. The
fused attention region is still 3–15x short of ORT at 8–16 threads (§15), the
Transpose and MoE gaps in §18–20 are unchanged, and #1078's own evidence has
float32 RoPE losing 12/12 cells at 1.53–17.21x. Under the rule above every one
of those is a kernel to fix with no exit; RoPE is first because it is the op that
was being given away. Tracked in §23.6 of the benchmark doc.
Two further caveats:
present_*shape is the non-shared-buffer form,so ORT logs a benign
MergeShapeInfo ... lenient mergewarning (it infers theshared-buffer form). Both sides still produce identical output; our static
rule is an upper bound on what the kernel writes, so it over-allocates by one
row rather than under-allocating.
QMoE's fixture needsblock_size = 32,because our kernel implements only the blocked form while an absent
block_sizeselects the column-wise one. ORT can run that form, so theright follow-up is to implement it here rather than to leave it declined.
Recorded in §23.6.
Docs
§10's assignment matrix and §15's "decline the whole region to ORT" conclusion
presented fallback as a solution. Both are marked withdrawn in place — the
ratios are kept because they are real evidence and they are the work queue, but
the conclusions drawn from them are void. New §23 records the rule, the audit,
both matrices and what is left.
CPU_ACTIVATION_GAPS.mdgains the two failuremodes this found (a dtype union that is too narrow, and the post-assignment
kernel factory).
Validation
cargo test -p onnx-runtime-ep-cpu --lib— 1300 passed;--features mlas— 1318 passedcargo test -p onnx-runtime-ep-plugin— 231 + 2 + 9 passed (one stale testasserting
com.microsoft::AttentionisDeclinedupdated to assert its newrule, and renamed to say what it checks)
cargo test -p onnx-runtime-ep-cpu-pluginwithNXRT_REQUIRE_ORT_TESTS=1against real ORT 1.27 — all suites green,
plugin_ort_e2e45 passed / 1ignored, with all 38 fixtures
cargo fmt --all -- --checkclean;cargo clippy --all-targetsclean on thethree touched crates
The
CLI ORT (Linux/Windows)CI jobs — the ones that setNXRT_REQUIRE_ORT_TESTS=1— fail onmainalready, so the ORT suites above wererun locally against
libonnxruntime.so1.27 rather than relied on from CI.