Skip to content

fix(cpu): keep the int4 prefill fan-out policy off the aarch64 dead-code lint - #1429

Closed
justinchuby wants to merge 2 commits into
mainfrom
squad/leon-p30-aarch64-int4-fanout
Closed

justinchuby wants to merge 2 commits into
mainfrom
squad/leon-p30-aarch64-int4-fanout

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 19, 2026 •

Copy link
Copy Markdown
Owner

Rust quality is red on main right now. Its cross-arch step fails:

$ cargo clippy --locked --target aarch64-unknown-linux-gnu \
    -p onnx-runtime-ep-cpu --lib -- -D warnings
error: constant `WIDE_PREFILL_MACS` is never used
error: function `prefill_fan_out` is never used
error: function `prefill_column_grain` is never used

Rust quality is a required check, so this blocks every open PR, not just this one.

Cause

In a default-feature build all three are reachable only from
borrowed_affine_int4_matmul_prefill, which is #[cfg(target_arch = "x86_64")].
(prefill_fan_out has one further caller in the mlas-gated run_mlas_shards, but
mlas is not a default feature, so the lint's build compiles it out too.) On any
non-x86 target the lib-only build therefore has no consumer left, and the step runs with
-D warnings.

This is exactly the failure mode scripts/check_cross_compile.sh documents in its own
error text (the #1037 case).

Bisected to bf722725a (#1363, "route the int4 flat output-row fan-out through the task
runtime"). Precisely, that commit removed the #[cfg_attr(not(feature = "mlas"), allow(dead_code))] guards that had been protecting the pre-existing WIDE_PREFILL_MACS
and prefill_fan_out, added prefill_column_grain, and introduced the x86-only
consumer -- so the items lost their only dead-code cover in the same change that made
them x86-only:

commit dead-code errors on aarch64
dcea14c74 (parent) 0
bf722725a (#1363) 3

Fix

#[cfg_attr(not(target_arch = "x86_64"), allow(dead_code))] on the three items, which is
the second remedy the check prescribes. cfg-gating them outright is the other option but
it is wrong here: the unit tests
(small_prefill_work_stays_on_the_task_runtime, prefill_column_grain_*, and friends)
reference all three and are portable, so gating the items would break the aarch64 test
build to fix the aarch64 lib build.

No behaviour change on any target: an allow attribute is lint-only.

Verification

lane before after
cargo clippy --target aarch64-unknown-linux-gnu -p onnx-runtime-ep-cpu --lib -- -D warnings 3 errors exit 0
bash scripts/check_cross_compile.sh (full offline set, both dimensions) fail exit 0
cargo clippy --locked --all-targets -p onnx-runtime-ep-cpu (-D warnings) exit 0 exit 0
cargo test --locked -p onnx-runtime-ep-cpu 1447 passed 1447 passed, 0 failed
cargo fmt --all --check clean clean

Found while replicating the CI lanes locally on main @ f8f3878ba, because Actions has
not concluded a run on this repo in some time. Every other lane of Fast (Linux x86_64)
and Rust quality is green on that commit (3942 tests over 188 targets); this was the only
red.

justinchuby and others added 2 commits August 19, 2026 07:04
…ode lint

`Rust quality`'s cross-arch step fails on main: `WIDE_PREFILL_MACS`,
`prefill_fan_out` and `prefill_column_grain` are consumed only by
`borrowed_affine_int4_matmul_prefill`, which is `#[cfg(target_arch =
"x86_64")]`, so a lib-only build for aarch64 sees three unused items under
`-D warnings`. Bisected to bf72272 (#1363); its parent dcea14c is clean.

Mark them `#[cfg_attr(not(target_arch = "x86_64"), allow(dead_code))]` rather
than `cfg`-gating them, because the unit tests that pin the policy are portable
and reference all three -- gating would fix the aarch64 lib build by breaking
the aarch64 test build. Lint-only, so no behaviour changes anywhere.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Review found two inaccuracies in the prose. `prefill_fan_out` also has a
caller in the `mlas`-gated `run_mlas_shards`, so "only the x86_64 prefill
entry point" was only true because `mlas` is off by default; say that. And
bf72272 did not introduce the two older items -- it removed their
`cfg_attr(not(feature = "mlas"), allow(dead_code))` guards while adding an
x86-only consumer, which is what actually lost them their cover.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Superseded — closing without merging, and not merging any overlapping aarch64 fix.

Two things changed since I opened this:

  1. main (a7bd79f72) now carries #[cfg(target_arch = "x86_64")] on all three items and on the tests that exercise them, so the cross-arch lint failure I bisected to bf722725a (perf(cpu): route the int4 flat output-row fan-out through the task runtime #1363) is already resolved on main. This PR is now DIRTY against it.
  2. fix(cpu-ep): keep the prefill fan-out policy present on aarch64+MLAS #1382 covers the same ground and is strictly better than what I wrote here. It uses not(any(feature = "mlas", target_arch = "x86_64")) for WIDE_PREFILL_MACS and prefill_fan_out — which is the correct predicate, because run_mlas_shards keeps them live on aarch64 + mlas — and the narrower not(target_arch = "x86_64") only for prefill_column_grain, whose sole non-test caller really is x86-only (run_mlas_shards takes prefill_tile_grain instead). My version used the single broad predicate for all three, which suppresses the lint on aarch64 + mlas where two of the items are genuinely live. fix(cpu-ep): keep the prefill fan-out policy present on aarch64+MLAS #1382 also restores the policy tests to aarch64 coverage, which main's current gating drops.

Deferring to #1382. Branch squad/leon-p30-aarch64-int4-fanout is abandoned.

auto-merge was automatically disabled August 19, 2026 14:14

Pull request was closed

@justinchuby

Copy link
Copy Markdown
Owner Author

Overlap audit: #1382 / #1393 / #1420 / #1429, and why only two of them should merge

Asked to find the duplication across these four and recommend a single minimal repair path. Short version: the aarch64 problem all of this was chasing is already fixed on main, by #1443, which merged at 09:20Z today. Two of the four PRs are now obsolete or actively harmful, and they are obsolete for different reasons.

The aarch64 lint is already fixed

beb15c202 ("gate x86_64-only prefill fan-out symbols for non-x86 targets", #1443) is on main. On current main all three symbols carry #[cfg(target_arch = "x86_64")]:

symbol line state on main
WIDE_PREFILL_MACS 395-396 #[cfg(target_arch = "x86_64")]
prefill_fan_out 403-404 #[cfg(target_arch = "x86_64")]
prefill_column_grain 543-544 #[cfg(target_arch = "x86_64")]

and so do the three tests that reference them (17303, 17316, 17332). The items and their only consumers vanish together on aarch64, so there is no dead code and no lint to silence. Nothing further is needed for the aarch64 lint.

This is the fifth attempt at the same problem — 32834e758, ef23f4a89, #1382, #1429, and finally #1443. That is the actual finding worth acting on, and I have suggested a process fix at the end.

#1429 — supersededr, and would rebase into a contradiction

#1429 adds #[cfg_attr(not(target_arch = "x86_64"), allow(dead_code))] to the same three symbols. It was branched from blob 817c05383, a state in which the item-level #[cfg] was absent. That state is no longer main.

Rebased onto current main the result is:

#[cfg(target_arch = "x86_64")]
#[cfg_attr(not(target_arch = "x86_64"), allow(dead_code))]
const WIDE_PREFILL_MACS: usize = 1 << 29;

The allow fires only when not x86_64, and when not x86_64 the cfg has already deleted the item. The attribute can never apply. Recommend closing #1429 as superseded by #1443 — it is a no-op that leaves a misleading attribute behind.

#1382 — not a duplicate repair, a competing design; must not merge as-is

#1382 is aimed at something genuinely different and arguably better: keep the prefill policy present and tested on aarch64 instead of compiling it away. It deletes #[cfg(target_arch = "x86_64")] from the three tests to restore aarch64 coverage, and then adds allow(dead_code) with a predicate that is more careful than #1429's — not(any(feature = "mlas", target_arch = "x86_64")) for the two symbols that have an MLAS-gated caller, but the narrower not(target_arch = "x86_64") for prefill_column_grain, whose only caller is x86-only because run_mlas_shards takes prefill_tile_grain instead. That distinction is correct and #1429's uniform predicate is over-broad: on aarch64 + mlas it would suppress a lint for symbols that genuinely are live.

But #1443 resolved the same question the opposite way. Merging #1382 now would re-delete the gating #1443 just added, so this is a design disagreement to settle deliberately, not a repair to land. Recommend either closing it, or re-scoping it to only the "restore aarch64 test coverage" argument on top of #1443 — with the ~25 lines of rationale it carries, because that rationale is the most accurate description of the caller structure anyone has written so far and should not be lost.

#1393 and #1420 — keep, no overlap between them

#1393 is the fmt repair, and it is still needed: cargo fmt --check fails on a clean checkout of current main at five sites across four files (gpt_oss_20b_decode_lock.rs:41, qwen35_0_8b_text_decode_lock.rs:70, gather_block_quantized.rs:132, dispatch.rs:23, dispatch.rs:725). Rust quality is a required check, so main is red and every open PR that merges main inherits the failure. I have refreshed it onto latest main; cargo fmt --check is clean on the result. This should merge first — it is the unblocker for the entire queue.

Note the overlap that did exist: #1382 also carries rustfmt fixes for three of those files. If #1382 is closed or re-scoped as recommended, #1393 is the single fmt path and there is no duplicate.

#1420 only shares a filename with the others. It fixes a different test, parallel_output_rows_dispatches_to_the_task_runtime, which asserted a task-runtime dispatch on hosts narrower than MIN_ROUTED_FAN_OUT_WIDTH where policy deliberately keeps the fan-out on Rayon — i.e. it was testing the host, not the policy. It installs a routing-width Rayon pool so the decision under test is the same on a 4-vCPU runner as on a 32-thread workstation. Refreshed onto latest main: 1447 passed, 0 failed. No interaction with the aarch64 gating.

Recommended path

  1. Merge ci: unbreak the Rust quality lane on main, a fourth time #1393 first (unblocks the required check for everything else).
  2. Close fix(cpu): keep the int4 prefill fan-out policy off the aarch64 dead-code lint #1429 as superseded by fix(cpu-ep): gate x86_64-only prefill fan-out symbols for non-x86 targets #1443.
  3. Close or re-scope fix(cpu-ep): keep the prefill fan-out policy present on aarch64+MLAS #1382; do not merge it as a repair.
  4. Merge fix(cpu): make the flat fan-out dispatch test host-independent #1420 on its own merits; it is unrelated to the aarch64 cluster.

The process point

Five PRs from three people attacked one lint, and the reason is visible in the history: the fmt/clippy gates only run on PR branches, so main can go red from an interaction between two independently-green PRs, and whoever notices opens a fix. That is why #1393 is titled "a fourth time" and is now on its fifth site. A push-triggered cargo fmt --check and cross-clippy job on main would attribute the breakage to the commit that caused it instead of to whoever merges next, and would have made four of these five PRs unnecessary.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 527.03 µs 1.49 ms +182.5%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.28 ms 2.62 ms +105.4%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 44.56 µs 73.49 µs +64.9%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 900.08 µs 1.48 ms +64.8%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 46.28 µs 66.53 µs +43.8%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 373.52 µs 499.51 µs +33.7%
⚠️ matmul/large_generic_bf16_threads=1/32x1024x1024 2.02 ms 2.55 ms +26.6%
⚠️ gather/medium_f32_threads=1-internal/32768 5.42 µs 6.83 µs +26.1%
⚠️ matmul/small_generic_f32_threads=8/1x256x256 37.96 µs 44.21 µs +16.5%
⚠️ matmul/medium_generic_f32_threads=1/32x512x512 2.22 ms 2.57 ms +15.7%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 568.48 µs 631.14 µs +11.0%
✅ add/large_f16_threads=1-internal/4194304 1.60 ms 1.75 ms +9.7%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.25 ms 2.46 ms +9.3%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.22 ms 10.06 ms +9.1%
✅ matmul/medium_generic_f16_threads=1/32x512x512 31.13 µs 33.73 µs +8.3%
✅ gather/large_bf16_threads=1-internal/131072 13.22 µs 14.28 µs +8.0%
✅ matmul/small_generic_bf16_threads=8/1x256x256 29.36 µs 31.61 µs +7.7%
✅ matmul/small_generic_f16_threads=8/1x256x256 35.22 µs 37.80 µs +7.3%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.79 ms 4.07 ms +7.1%
✅ qwen3_sampling_processors/top_k_top_p_fast 644.85 µs 684.78 µs +6.2%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 525.69 µs 554.45 µs +5.5%
✅ sampling_latency/greedy_per_token 3.05 µs 3.22 µs +5.4%
✅ add/large_bf16_threads=1-internal/4194304 1.64 ms 1.72 ms +4.8%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.27 ms 6.47 ms +3.2%
✅ tokenization/encode_tokens_per_second 397.82 µs 402.76 µs +1.2%
✅ qwen3_sampling_processors/top_k_partial_selection 154.44 µs 156.20 µs +1.1%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 61.23 µs 61.59 µs +0.6%
✅ kv_cache/alloc_dealloc_pages 41.55 µs 41.66 µs +0.3%
✅ add/small_f32_threads=1-internal/1024 213.3 ns 213.4 ns +0.1%
✅ gather/medium_f16_threads=1-internal/32768 2.57 µs 2.55 µs -0.9%
✅ matmul/small_generic_f16_threads=1/1x256x256 35.14 µs 34.76 µs -1.1%
✅ grammar_masking/llguidance_compute_mask/32 80.84 µs 79.17 µs -2.1%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 4.59 ms 4.46 ms -2.9%
✅ logit_processing/seven_processor_chain_per_step 335.45 µs 323.73 µs -3.5%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 95.24 µs 90.97 µs -4.5%
✅ gather/large_f32_threads=1-internal/131072 33.39 µs 31.87 µs -4.5%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.06 ms 1.00 ms -5.5%
✅ sampling_latency/min_p_per_token 252.03 µs 236.89 µs -6.0%
✅ reduce_mean/medium_f32_threads=1-internal/65536 263.47 µs 245.12 µs -7.0%
✅ matmul/small_generic_f32_threads=1/1x256x256 39.95 µs 36.90 µs -7.7%
✅ sampling_latency/top_k_per_token 59.23 µs 54.46 µs -8.1%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.66 µs 15.25 µs -8.4%
✅ gather/small_f32_threads=1-internal/4096 725.9 ns 660.9 ns -9.0%
✅ gather/medium_bf16_threads=1-internal/32768 2.72 µs 2.45 µs -10.0%
✅ add/small_f16_threads=1-internal/1024 527.4 ns 474.3 ns -10.1%
✅ sampling_latency/top_p_per_token 456.54 µs 399.66 µs -12.5%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 587.37 µs 511.49 µs -12.9%
✅ gather/small_bf16_threads=1-internal/4096 737.6 ns 640.3 ns -13.2%
✅ matmul/small_generic_bf16_threads=1/1x256x256 36.75 µs 31.59 µs -14.1%
🟢 add/medium_bf16_threads=1-internal/262144 123.03 µs 104.49 µs -15.1%
🟢 add/medium_f16_threads=1-internal/262144 122.00 µs 102.68 µs -15.8%
🟢 gather/large_f16_threads=1-internal/131072 18.16 µs 15.07 µs -17.0%
🟢 add/small_bf16_threads=1-internal/1024 567.9 ns 465.0 ns -18.1%
🟢 gather/small_f16_threads=1-internal/4096 647.9 ns 529.6 ns -18.3%
🟢 tokenization/decode_tokens_per_second 7.70 ms 6.01 ms -22.0%
🟢 add/medium_f32_threads=1-internal/262144 36.33 µs 26.76 µs -26.4%
🟢 add/large_f32_threads=1-internal/4194304 1.03 ms 697.17 µs -32.6%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 51.04 µs 33.19 µs -35.0%
🟢 matmul/large_generic_f16_threads=1/32x1024x1024 130.86 µs 83.22 µs -36.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.99 3.28 4.71 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

Confirmed superseded: #1382 merged as fbf519e87, which keeps WIDE_PREFILL_MACS/prefill_fan_out/prefill_column_grain compiled via cfg_attr(..., allow(dead_code)) and restores the six prefill-policy unit tests on aarch64.

Independent evidence that this was the correct fix rather than a cosmetic one: the exact Rust (Windows ARM64) step cargo build --target aarch64-* -p onnx-runtime-ep-cpu-plugin --features mlas failed on main 7de4bb1dc with error[E0425]: cannot find function \prefill_fan_out`(exit 101) and passes onfbf519e(exit 0).run_mlas_shardscalls it underfeature = "mlas"` on any arch, so the x86-only item gate left the caller without a callee.

No action needed here — closing stands. cc @justinchuby (and Leon for the duplicate-path audit): #1382 is the single surviving repair, no duplicate landed.

@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.64%. Comparing base (1c9f31d) to head (9333ca9).
⚠️ Report is 111 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@             Coverage Diff             @@
##             main    #1429       +/-   ##
===========================================
+ Coverage   80.26%   82.64%    +2.38%     
===========================================
  Files         362       12      -350     
  Lines      157159     5475   -151684     
  Branches   157159     5475   -151684     
===========================================
- Hits       126137     4525   -121612     
+ Misses      26406      757    -25649     
+ Partials     4616      193     -4423     
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (?)
cli-ort-windows 82.19% <ø> (?)
offline ?

Flags with carried forward coverage won't be shown. Click here to find out more.
see 374 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant