Skip to content

fix(cpu-ep): gate x86_64-only prefill fan-out symbols for non-x86 targets - #1443

Merged
justinchuby merged 2 commits into
mainfrom
fix/aarch64-clippy-dead-prefill
Aug 19, 2026
Merged

justinchuby merged 2 commits into
mainfrom
fix/aarch64-clippy-dead-prefill

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Fixes #1415 — the aarch64 clippy gate is red on main.

WIDE_PREFILL_MACS, prefill_fan_out and prefill_column_grain are only reachable from the x86_64 prefill route, so on other architectures they are genuinely dead and -D warnings fails the build. This adds #[cfg(target_arch = "x86_64")] to the three symbols and to the six tests that exercise them, matching the treatment their neighbours already have.

Verified both sides, because gating dead code is only half the job — the risk is silently removing something x86 needs:

  • aarch64: reproduced the three never used diagnostics on unmodified main first, then confirmed cargo clippy -p onnx-runtime-ep-cpu --target aarch64-pc-windows-msvc --all-targets -- -D warnings exits 0 with this change.
    Caveat, stated honestly: aarch64 clippy gate is red on main: three prefill fan-out symbols are x86-only but not cfg-gated #1415 cites aarch64-unknown-linux-gnu, whose std is not installed on this box (can't find crate for core). I used aarch64-pc-windows-msvc instead. The symbols are dead on any non-x86_64 target, and that target reproduces the identical three diagnostics and clears them, so it is a valid stand-in — but the exact triple in the issue was not the one I ran.
  • x86_64: cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings exits 0, and the six gated tests still run: 26 passed / 0 failed / 2 ignored.

No logic changes — attributes only.

Copilot AI added 2 commits August 19, 2026 02:04
…ing time)

The `GLOBAL_VRAM_FREE_NS` / `vram_free_ns` counter previously documented
itself as measuring "only unmap/release". The 2026-08-19 reconciliation
(docs/benchmarks/2026-08-19-vram-free-attribution-reconciliation.md) showed
that is misleading: with the stream drain already split into
`vram_free_sync_ns`, the counter still wraps the whole free code path —
the `cuMemUnmap`/`cuMemRelease` driver calls PLUS the per-granule Rust
bookkeeping in `decommit_allocation_range`/`deallocate_span` — and the
bookkeeping dominates. On an RTX 4060 weight-lending run the driver
`cuMemUnmap` was a stable ~16.9 ms/call (~101 ms over 6 releases) while
`vram_free_ms` swung 82 ms <-> 2450 ms (~30x) on byte-identical
deterministic work. A metric that varies 30x while the work is constant is
measuring variable non-driver time, not freeing.

Documentation-only: updates the comment at the static definition and the
`GlobalOffloadStats::vram_free_ns` field doc so the next reader does not
mistake this counter for a driver-freeing figure. No behaviour, symbol, or
output-schema change (the `vram_free_ms` CSV column is consumed by
benchmark tooling and is intentionally left unrenamed).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…gets

The aarch64 clippy gate is red on main (#1415): WIDE_PREFILL_MACS,
prefill_fan_out and prefill_column_grain are only reachable from the
x86_64 prefill route, so on other architectures they are genuinely dead
and `-D warnings` fails the build.

Adds `#[cfg(target_arch = "x86_64")]` to the three symbols and to the six
tests that exercise them, matching the treatment their neighbours already
have.

Verified both sides, because gating dead code is only half the job:
- aarch64-pc-windows-msvc: reproduced the three "never used" diagnostics
  on unmodified main, then `cargo clippy --target aarch64-pc-windows-msvc
  --all-targets -- -D warnings` exits 0 with this change. (The issue cites
  aarch64-unknown-linux-gnu; that target's std is not installed here, but
  the symbols are dead on any non-x86_64 target so the aarch64 Windows
  target reproduces and clears the same diagnostics.)
- x86_64: `cargo clippy --all-targets -- -D warnings` exits 0, and the six
  gated tests still run — 26 passed / 0 failed / 2 ignored.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
@justinchuby
justinchuby merged commit beb15c2 into main Aug 19, 2026
7 checks passed
@justinchuby
justinchuby deleted the fix/aarch64-clippy-dead-prefill branch August 19, 2026 09:20
justinchuby pushed a commit that referenced this pull request Aug 19, 2026
#1443 stopped the aarch64 dead-code errors by `#[cfg]`-ing the three
prefill fan-out symbols to `target_arch = "x86_64"`. That removes the
items outright, and they have two callers gated on *different* things:

  matmul_nbits.rs:2457  run_mlas_shards                    #[cfg(feature = "mlas")]
  matmul_nbits.rs:6914  borrowed_affine_int4_matmul_prefill #[cfg(target_arch = "x86_64")]

So on `aarch64 + feature = "mlas"` -- Apple Silicon, the primary aarch64
target -- the MLAS caller is still compiled while its callee is not:

  error[E0425]: cannot find function `prefill_fan_out` in this scope
    --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:2457:27
  error[E0425]: cannot find function `prefill_column_grain` in this scope
    --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:6915:48

The cross-compile gate cannot see this: its aarch64 pass builds default
features, and MLAS is off by default.

Gating the items also forced gating the six unit tests that reference
them, so the #1363 fan-out policy stopped being checked on aarch64 at
all. Those tests are pure functions of explicit arguments -- nothing in
them is architecture-specific.

`cfg_attr(.., allow(dead_code))` fixes both: the item always exists, so
whichever caller survives can reach it, and the lint is silenced only in
the configuration where neither caller exists. `prefill_tile_grain`
already used exactly this idiom, which is what #1363 dropped.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby pushed a commit that referenced this pull request Aug 19, 2026
#1443 stopped the aarch64 dead-code errors by `#[cfg]`-ing the three
prefill fan-out symbols to `target_arch = "x86_64"`. That removes the
items outright, and they have two callers gated on *different* things:

  matmul_nbits.rs:2457  run_mlas_shards                    #[cfg(feature = "mlas")]
  matmul_nbits.rs:6914  borrowed_affine_int4_matmul_prefill #[cfg(target_arch = "x86_64")]

So on `aarch64 + feature = "mlas"` -- Apple Silicon, the primary aarch64
target -- the MLAS caller is still compiled while its callee is not:

  error[E0425]: cannot find function `prefill_fan_out` in this scope
    --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:2457:27
  error[E0425]: cannot find function `prefill_column_grain` in this scope
    --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:6915:48

The cross-compile gate cannot see this: its aarch64 pass builds default
features, and MLAS is off by default.

Gating the items also forced gating the six unit tests that reference
them, so the #1363 fan-out policy stopped being checked on aarch64 at
all. Those tests are pure functions of explicit arguments -- nothing in
them is architecture-specific.

`cfg_attr(.., allow(dead_code))` fixes both: the item always exists, so
whichever caller survives can reach it, and the lint is silenced only in
the configuration where neither caller exists. `prefill_tile_grain`
already used exactly this idiom, which is what #1363 dropped.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Heads-up from validation (Pris, Tester): this fix removes the three symbols rather than silencing the lint, and that breaks a configuration the aarch64 cross gate cannot see.

prefill_fan_out / WIDE_PREFILL_MACS have two callers gated on different things:

call site fn gate
matmul_nbits.rs:2457 run_mlas_shards #[cfg(feature = "mlas")]
matmul_nbits.rs:6914 borrowed_affine_int4_matmul_prefill #[cfg(target_arch = "x86_64")]

Off x86 the second disappears but the first does not, so on aarch64 + mlas the caller is compiled and the callee is not:

error[E0425]: cannot find function `prefill_fan_out` in this scope
  --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:2457:27

Reproduced by mirroring the cfg resolution on the host: retarget the three item gates off x86_64 on main, leave callers alone, cargo check -p onnx-runtime-ep-cpu --features mlas --lib → exit 101. check_cross_compile.sh's aarch64 pass builds default features and mlas is off by default, which is why it stayed green. The rust-coverage macOS-arm64 leg does build it (-D warnings, --features mlas, ci.yml L434/439/501-503) — that is the shipped-wheel path.

Second effect: gating the items forced #[cfg(target_arch = "x86_64")] onto six unit tests, all of which are pure functions of literal arguments. The #1363 fan-out policy now has no aarch64 coverage.

#1382 replaces the gates with per-symbol cfg_attr(.., allow(dead_code)) (the idiom prefill_tile_grain already uses) and ungates the tests. Predicates differ per symbol — prefill_column_grain's only caller is the x86 one, so it gets not(target_arch = "x86_64") while the other two get the union; using the union for all three just relocates the failure to never used on aarch64+mlas. Verified across all four (mlas) x (arch) configs.

No action needed from you — the repair is in #1382, waiting on required CI.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 126.39 µs 223.74 µs +77.0%
🔴 tokenization/encode_tokens_per_second 376.52 µs 611.63 µs +62.4%
🔴 sampling_latency/top_k_per_token 52.31 µs 76.79 µs +46.8%
🔴 sampling_latency/top_p_per_token 387.92 µs 559.88 µs +44.3%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 4.42 ms 6.31 ms +42.8%
🔴 matmul/small_generic_f32_threads=8/1x256x256 47.38 µs 64.08 µs +35.2%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.36 ms 1.82 ms +33.8%
🔴 matmul/large_generic_bf16_threads=1/32x1024x1024 2.23 ms 2.97 ms +32.8%
🔴 add/medium_f32_threads=1-internal/262144 25.45 µs 33.58 µs +32.0%
🔴 tokenization/decode_tokens_per_second 6.29 ms 8.22 ms +30.6%
⚠️ matmul/medium_generic_bf16_threads=8/32x512x512 578.17 µs 727.07 µs +25.8%
⚠️ gather/large_bf16_threads=1-internal/131072 16.02 µs 19.47 µs +21.5%
⚠️ sampling_latency/greedy_per_token 3.21 µs 3.89 µs +21.1%
⚠️ matmul/large_generic_f32_threads=1/32x1024x1024 10.03 ms 12.01 ms +19.7%
⚠️ add/small_bf16_threads=1-internal/1024 498.8 ns 584.9 ns +17.3%
⚠️ matmul/medium_generic_f16_threads=8/32x512x512 39.42 µs 45.88 µs +16.4%
⚠️ add/small_f16_threads=1-internal/1024 510.1 ns 592.8 ns +16.2%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 667.13 µs 766.69 µs +14.9%
✅ gather/small_f16_threads=1-internal/4096 515.8 ns 583.4 ns +13.1%
✅ matmul/small_generic_f16_threads=1/1x256x256 39.85 µs 44.42 µs +11.5%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.10 µs 17.87 µs +11.0%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.27 ms 4.72 ms +10.5%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 87.73 µs 96.18 µs +9.6%
✅ qwen3_sampling_processors/top_k_partial_selection 180.20 µs 194.36 µs +7.9%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 638.00 µs 687.31 µs +7.7%
✅ reduce_mean/medium_f32_threads=1-internal/65536 252.09 µs 268.20 µs +6.4%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.68 ms 2.79 ms +4.2%
✅ matmul/small_generic_bf16_threads=8/1x256x256 39.24 µs 40.86 µs +4.1%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 584.33 µs 601.72 µs +3.0%
✅ gather/small_f32_threads=1-internal/4096 742.2 ns 760.1 ns +2.4%
✅ add/small_f32_threads=1-internal/1024 223.1 ns 228.4 ns +2.4%
✅ gather/small_bf16_threads=1-internal/4096 537.4 ns 545.0 ns +1.4%
✅ logit_processing/seven_processor_chain_per_step 355.82 µs 357.29 µs +0.4%
✅ matmul/small_generic_f32_threads=1/1x256x256 51.71 µs 51.27 µs -0.8%
✅ matmul/small_generic_bf16_threads=1/1x256x256 36.82 µs 36.41 µs -1.1%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 605.53 µs 596.42 µs -1.5%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.08 ms 1.06 ms -1.9%
✅ matmul/small_generic_f16_threads=8/1x256x256 40.67 µs 39.67 µs -2.5%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.67 ms 6.49 ms -2.8%
✅ qwen3_sampling_processors/top_k_top_p_fast 758.57 µs 734.53 µs -3.2%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 140.23 µs 134.64 µs -4.0%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 106.95 µs 102.58 µs -4.1%
✅ gather/medium_f16_threads=1-internal/32768 2.98 µs 2.85 µs -4.1%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.62 ms 2.47 ms -5.6%
✅ add/large_bf16_threads=1-internal/4194304 1.92 ms 1.77 ms -7.6%
✅ gather/medium_f32_threads=1-internal/32768 4.50 µs 4.12 µs -8.4%
✅ gather/medium_bf16_threads=1-internal/32768 2.86 µs 2.55 µs -11.1%
✅ gather/large_f16_threads=1-internal/131072 15.36 µs 13.48 µs -12.2%
✅ add/large_f16_threads=1-internal/4194304 2.07 ms 1.80 ms -12.8%
✅ add/medium_f16_threads=1-internal/262144 126.38 µs 109.28 µs -13.5%
✅ sampling_latency/min_p_per_token 291.19 µs 251.58 µs -13.6%
✅ grammar_masking/llguidance_compute_mask/32 97.82 µs 83.95 µs -14.2%
🟢 kv_cache/alloc_dealloc_pages 48.64 µs 41.14 µs -15.4%
🟢 add/medium_bf16_threads=1-internal/262144 126.62 µs 106.11 µs -16.2%
🟢 matmul/medium_generic_f16_threads=1/32x512x512 43.39 µs 36.01 µs -17.0%
🟢 gather/large_f32_threads=1-internal/131072 33.12 µs 27.11 µs -18.1%
🟢 add/large_f32_threads=1-internal/4194304 907.47 µs 736.71 µs -18.8%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.62 ms 1.25 ms -22.8%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 114.05 µs 68.00 µs -40.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.45 4.47 5.57 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 19, 2026
…1382)

## What this fixes

Validating the merged scheduler/fmt wave, I found `#1363` had dropped
the `allow(dead_code)` guards on three prefill fan-out symbols, breaking
the aarch64 cross-compile gate. **#1443 has since fixed that** — but by
`#[cfg]`-ing the three items to `target_arch = "x86_64"`, which removes
them outright. That trades one break for two others.

### 1. 🔴 `aarch64 + feature = "mlas"` no longer compiles

The three symbols have two callers, gated on **different** things:

| call site | enclosing fn | its gate |
| --- | --- | --- |
| `matmul_nbits.rs:2457` | `run_mlas_shards` | `#[cfg(feature =
"mlas")]` |
| `matmul_nbits.rs:6914-6915` | `borrowed_affine_int4_matmul_prefill` |
`#[cfg(target_arch = "x86_64")]` |

Off x86 the second caller disappears, but the **first does not** — it is
arch-independent. So on `aarch64 + mlas`, which is Apple Silicon, the
caller is compiled and its callee is not:

```
error[E0425]: cannot find function `prefill_fan_out` in this scope
  --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:2457:27
error[E0425]: cannot find function `prefill_fan_out` in this scope
  --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:6914:8
error[E0425]: cannot find function `prefill_column_grain` in this scope
  --> crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:6915:48
```

**How that was produced.** MLAS's vendored sources do not cross-build to
aarch64 in this container (`arm_neon.h: inlining failed in call to
always_inline vaddq_f16 — target specific option mismatch`, an
mlas-sys/toolchain issue unrelated to this PR), so instead I reproduced
the *exact cfg resolution* on the host: on `main`, retarget the three
item gates from `x86_64` to a third arch so they are absent, leave every
caller alone, and build the lib with MLAS on —

```
sed -i '395s/x86_64/s390x/; 403s/x86_64/s390x/; 543s/x86_64/s390x/' \
  crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs
cargo check --locked -p onnx-runtime-ep-cpu --features mlas --lib     # exit 101
```

That is precisely the configuration `aarch64 + mlas` produces. The
`:2457` error is the load-bearing one — that call site is
`feature`-gated only, so it is present on aarch64 for real.

**Why no gate caught it.** `check_cross_compile.sh`'s aarch64 pass
builds **default features**, and `mlas` is off by default — the same
blind spot #1443 was written under. The configuration *is* built
elsewhere: the `rust-coverage` job's macOS-arm64 leg runs with
`RUSTFLAGS: -D warnings` (`ci.yml` L434, L439) and builds `cargo build
-p onnx-runtime-ep-cpu-plugin --features mlas` (L501-503), whose `mlas`
feature forwards to `onnx-runtime-ep-cpu/mlas`. That is the
shipped-wheel build path, so `main` as it stands also breaks the
macOS-arm64 MLAS wheel at release time.

### 2. 🔴 The #1363 fan-out policy stopped being tested on aarch64

Gating the items forced gating their tests, so #1443 also put
`#[cfg(target_arch = "x86_64")]` on six unit tests. All six are pure
functions of explicit literal arguments —
`prefill_fan_out(WIDE_PREFILL_MACS - 1, 16, 32)`,
`prefill_column_grain(8, 1024, 3072)` — with nothing
architecture-specific in them. They are now simply not compiled off x86,
so the policy that #1363 rewrote has no aarch64 coverage.

## The fix

`cfg_attr(.., allow(dead_code))` instead of `cfg`. The item always
exists, so whichever caller survives can reach it; the lint is silenced
only where **no** caller exists. The six tests are ungated and run
everywhere again. `prefill_tile_grain` in this same file already uses
this idiom (`not(feature = "mlas")`) — that is the shape #1363 deleted.

The predicate is **per-symbol**, because the caller sets differ:

| symbol | callers | predicate |
| --- | --- | --- |
| `WIDE_PREFILL_MACS`, `prefill_fan_out` | `run_mlas_shards` **and**
`borrowed_affine_int4_matmul_prefill` | `not(any(feature = "mlas",
target_arch = "x86_64"))` |
| `prefill_column_grain` | `borrowed_affine_int4_matmul_prefill`
**only** — `run_mlas_shards` takes `prefill_tile_grain` instead |
`not(target_arch = "x86_64")` |

Giving `prefill_column_grain` the union predicate would leave the lint
live on `aarch64 + mlas`, where it has no caller — converting #1443's
`E0425` into a `never used` error in the same configuration. Review
caught exactly that in the first draft of this branch; the four-way
probe below is the regression check for it.

### Four-config probe of the predicates

`.validation-worktrees/cfgprobe/probe.rs` reproduces the two items, the
two callers and their gates with `mlas`/`x86` standing in for the real
cfgs, compiled under `-D warnings`:

```
=== union predicate on prefill_column_grain (wrong) ===
  PASS  [aarch64 default]
  FAIL  [--cfg mlas] <- error: function `prefill_column_grain` is never used
  PASS  [--cfg x86]
  PASS  [--cfg mlas --cfg x86]

=== per-symbol predicates (this PR) ===
  PASS  [aarch64 default]
  PASS  [--cfg mlas]
  PASS  [--cfg x86]
  PASS  [--cfg mlas --cfg x86]
```

The probe is sharp, not vacuous: it fails on exactly the configuration
that is wrong, and only that one.

## Second commit: unbreaking `main`'s required lane

`main` currently fails **both** required checks, from merges landed past
queued checks:

| defect | source | breaks |
| --- | --- | --- |
| `map_or(true, ..)` in `executor/dispatch.rs` — clippy `this map_or can
be simplified` under `-D warnings` | #1427 | `Rust quality` — and it
aborts the cross-compile gate *before* its aarch64 pass, which is why
the gate never reported defect 1 |
| `dispatch.rs`, `gather_block_quantized.rs`,
`gpt_oss_20b_decode_lock.rs` unformatted | #1427, #1418 | `Fast (Linux
x86_64)` **and** `Rust quality` |

`cargo fmt --all -- --check` runs in *both* required jobs (`ci.yml`
L162, L276) while `check_cross_compile.sh` runs only in `Rust quality`
(L401), and PR checks run against `merge(base, head)`. So while `main`
is broken this way a fmt-only PR still fails the cross-compile step and
this PR alone still fails fmt — only a branch carrying both can go
green. It is mechanical (`cargo fmt --all`, plus `map_or(true, f)` →
`is_none_or(f)`, identical on `Option`) and `git rebase` drops it once
fixed upstream.

## Verification at `9fb04f5b5` (base `main` `81f99ff42`)

| check | step | `main` | this branch |
| --- | --- | --- | --- |
| `Fast` + `Rust quality` | `cargo fmt --all -- --check` | **FAIL** (4
diffs / 3 files) | **pass** |
| `Rust quality` | `bash scripts/check_cross_compile.sh` | **FAIL** exit
1 | **pass** exit 0, `scope: full offline set (aarch64 cross toolchain
present)` |
| `Rust quality` | 30-crate `cargo clippy --locked --all-targets … -- -D
warnings` | **FAIL** exit 1 | **pass** exit 0 |
| `Rust quality` | 9 guard scripts | pass | **9/9 pass** |
| aarch64 | `cargo clippy --target aarch64-unknown-linux-gnu
--all-targets -p onnx-runtime-ep-cpu -- -D warnings` | pass | **pass**
(now *with* the 6 tests compiled) |
| `aarch64 + mlas` cfg resolution | `cargo check -p onnx-runtime-ep-cpu
--features mlas --lib`, items absent | **FAIL** exit 101, 3 × E0425 |
**pass** — `cfg_attr` never removes the item, so E0425 cannot occur |
| all 4 `(mlas on/off) x (x86 / non-x86)` | `rustc -D warnings` cfg
probe | — | **4/4 pass** |
| tests | `cargo test -p onnx-runtime-ep-cpu --lib` | — | **1447 passed
/ 0 failed**; the 8 prefill policy tests pass |

## A note on the gate that found this

`scripts/check_cross_compile.sh` **false-passes locally** without an
aarch64 cross toolchain: at L191-194 it silently swaps `CRATES_FULL` →
`CRATES_NO_FFI`, dropping `onnx-runtime-ep-cpu` — the crate the gate
exists for — and still exits 0 with a ✓. The "REDUCED SCOPE" note prints
*below* the checkmark. **Read the scope note, not the exit code**; only
`scope: full offline set (aarch64 cross toolchain present)` means
anything. On Actions it `exit 2`s instead (L178-190), and `ci.yml`
L396-399 installs `gcc-aarch64-linux-gnu` + `libc6-dev-arm64-cross`
before invoking it, so the fail-loud coverage is intact — this is a
local-only trap. All results above were produced with the toolchain
installed, at full scope.

Two of the three defects in this PR would have been caught by the
required checks had they been allowed to run.

## Process

No admin bypass, no ruleset bypass, no merge with checks queued or
failing. Auto-merge has been armed since 2026-08-19T04:55:37Z and merges
only once `Fast (Linux x86_64)` and `Rust quality` are green. Every CI
run in this repo is currently `queued` with zero in progress, so the
required contexts have not been created yet. Waiting.

---------

Co-authored-by: Pris <pris@squad.local>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 19, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.10%. Comparing base (e19697f) to head (e211f03).
⚠️ Report is 104 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@             Coverage Diff             @@
##             main    #1443       +/-   ##
===========================================
+ Coverage   80.14%   82.10%    +1.96%     
===========================================
  Files         364       12      -352     
  Lines      160978     5471   -155507     
  Branches   160978     5471   -155507     
===========================================
- Hits       129012     4492   -124520     
+ Misses      27310      780    -26530     
+ Partials     4656      199     -4457     
Flag Coverage Δ
cli-ort-windows 82.10% <ø> (?)
mlas ?
offline ?

Flags with carried forward coverage won't be shown. Click here to find out more.
see 376 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

aarch64 clippy gate is red on main: three prefill fan-out symbols are x86-only but not cfg-gated

2 participants