Skip to content

style: rustfmt device_argmax.rs (unblocks Fast CI after #1119) - #1120

Merged
justinchuby merged 1 commit into
mainfrom
squad/roy-fmt-1119
Aug 17, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/roy-fmt-1119

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

Pure rustfmt-only repair of test blocks in crates/onnx-runtime-ep-cuda/src/kernels/device_argmax.rs that #1119 left unformatted.

Details

Files changed

crates/onnx-runtime-ep-cuda/src/kernels/device_argmax.rs | 16 ++++++++++------
 1 file changed, 10 insertions(+), 6 deletions(-)

References #1119.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit b7fa5e1 into main Aug 17, 2026
9 of 14 checks passed
@justinchuby
justinchuby deleted the squad/roy-fmt-1119 branch August 17, 2026 09:01
@codecov

codecov Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 79.94%. Comparing base (e7a291c) to head (e551ed4).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1120      +/-   ##
==========================================
+ Coverage   79.84%   79.94%   +0.09%     
==========================================
  Files         367      369       +2     
  Lines      157368   160348    +2980     
  Branches   157368   160348    +2980     
==========================================
+ Hits       125656   128186    +2530     
- Misses      26991    27424     +433     
- Partials     4721     4738      +17     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.31% <ø> (ø)
mlas 84.54% <ø> (?)
offline 79.71% <ø> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 11 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 47.37 µs 135.54 µs +186.1%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 595.20 µs 1.17 ms +97.3%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 54.36 µs 106.99 µs +96.8%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 436.59 µs 583.65 µs +33.7%
⚠️ kv_cache/alloc_dealloc_pages 42.14 µs 52.38 µs +24.3%
⚠️ logit_processing/seven_processor_chain_per_step 337.35 µs 413.29 µs +22.5%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 77.61 µs 93.11 µs +20.0%
⚠️ matmul/small_generic_bf16_threads=1/1x256x256 31.08 µs 36.94 µs +18.9%
✅ gather/small_bf16_threads=1-internal/4096 441.8 ns 500.0 ns +13.2%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 156.32 µs 176.80 µs +13.1%
✅ reduce_mean/small_f32_threads=1-internal/4096 13.95 µs 15.70 µs +12.5%
✅ add/medium_bf16_threads=1-internal/262144 96.64 µs 108.10 µs +11.9%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 72.26 µs 80.16 µs +10.9%
✅ gather/large_f32_threads=1-internal/131072 38.13 µs 41.69 µs +9.3%
✅ matmul/medium_generic_f16_threads=1/32x512x512 30.39 µs 33.12 µs +9.0%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.82 ms 1.98 ms +8.8%
✅ sampling_latency/top_k_per_token 53.29 µs 57.96 µs +8.7%
✅ qwen3_sampling_processors/top_k_top_p_fast 662.30 µs 718.96 µs +8.6%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.65 ms 3.95 ms +8.3%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.43 ms 2.62 ms +7.8%
✅ matmul/medium_generic_f16_threads=8/32x512x512 37.66 µs 40.45 µs +7.4%
✅ add/large_bf16_threads=1-internal/4194304 1.51 ms 1.61 ms +6.5%
✅ add/large_f16_threads=1-internal/4194304 1.56 ms 1.66 ms +6.3%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 8.75 ms 9.23 ms +5.5%
✅ tokenization/encode_tokens_per_second 396.83 µs 418.35 µs +5.4%
✅ reduce_mean/large_f32_threads=1-internal/262144 925.22 µs 973.00 µs +5.2%
✅ reduce_mean/medium_f32_threads=1-internal/65536 230.52 µs 240.30 µs +4.2%
✅ sampling_latency/top_p_per_token 386.67 µs 400.35 µs +3.5%
✅ grammar_masking/llguidance_compute_mask/32 80.42 µs 82.27 µs +2.3%
✅ gather/small_f16_threads=1-internal/4096 498.5 ns 502.4 ns +0.8%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.30 ms 1.31 ms +0.3%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.13 ms 2.14 ms +0.2%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.57 ms 3.58 ms +0.2%
✅ gather/small_f32_threads=1-internal/4096 654.1 ns 653.6 ns -0.1%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 389.75 µs 389.38 µs -0.1%
✅ sampling_latency/greedy_per_token 3.31 µs 3.30 µs -0.2%
✅ gather/medium_f16_threads=1-internal/32768 2.35 µs 2.34 µs -0.2%
✅ tokenization/decode_tokens_per_second 6.49 ms 6.46 ms -0.5%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 521.48 µs 518.52 µs -0.6%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 525.36 µs 522.32 µs -0.6%
✅ sampling_latency/min_p_per_token 207.73 µs 206.32 µs -0.7%
✅ gather/large_bf16_threads=1-internal/131072 14.38 µs 14.07 µs -2.1%
✅ matmul/small_generic_f16_threads=1/1x256x256 36.13 µs 35.16 µs -2.7%
✅ add/small_bf16_threads=1-internal/1024 453.6 ns 437.7 ns -3.5%
✅ add/medium_f16_threads=1-internal/262144 109.62 µs 105.41 µs -3.8%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.86 ms 5.62 ms -4.0%
✅ qwen3_sampling_processors/top_k_partial_selection 141.11 µs 134.58 µs -4.6%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.15 ms 1.08 ms -5.9%
✅ add/medium_f32_threads=1-internal/262144 25.28 µs 23.73 µs -6.1%
✅ matmul/small_generic_f16_threads=8/1x256x256 34.37 µs 32.09 µs -6.6%
✅ matmul/small_generic_f32_threads=1/1x256x256 42.15 µs 39.19 µs -7.0%
✅ gather/medium_f32_threads=1-internal/32768 4.42 µs 4.07 µs -7.8%
✅ matmul/small_generic_bf16_threads=8/1x256x256 38.30 µs 34.97 µs -8.7%
✅ gather/medium_bf16_threads=1-internal/32768 2.78 µs 2.41 µs -13.3%
✅ add/large_f32_threads=1-internal/4194304 759.43 µs 654.88 µs -13.8%
🟢 matmul/small_generic_f32_threads=8/1x256x256 54.31 µs 42.42 µs -21.9%
🟢 add/small_f32_threads=1-internal/1024 276.4 ns 203.6 ns -26.4%
🟢 gather/large_f16_threads=1-internal/131072 19.35 µs 13.83 µs -28.5%
🟢 add/small_f16_threads=1-internal/1024 648.5 ns 456.9 ns -29.5%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.63 3.27 4.87 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 17, 2026
…ild shipped (#1115)

## The published wheel shipped the slow build

ORT's own CPU execution provider **is** MLAS. This repository vendors
MLAS (`crates/mlas-sys`, 833 files) and `onnx-runtime-ep-cpu` has an
opt-in `mlas` feature for it — but nothing in the packaging ever turned
it on. `python/nxrt-ep-cpu/setup.py` ran `cargo build --release -p
onnx-runtime-ep-cpu-plugin` with `CARGO_FEATURES: list[str] = []`, so
**every published `nxrt-ep-cpu` wheel contained the pure-Rust fallback
paths**.

That is not a small difference. Measured end-to-end through the plugin
path (the harness added in #1110), same host, same ORT, interleaved A/B:

| case | ours/ORT p50, no MLAS | ours/ORT p50, MLAS | ours p50 (ms) no
MLAS → MLAS |
|---|---|---|---|
| `MatMulNBits` int4 M=128 | **81.1x** | 7.3x | 115.9 → 8.80 |
| `MatMulNBits` int4 M=1 | 14.8x | 5.2x | 1.850 → 0.400 |
| `MatMulNBits` int4 f16-act M=1 | 20.1x | 4.7x | 1.418 → 0.437 |
| `MatMulNBits` int8 M=256 | 7.1x | 1.5x | 46.38 → 13.86 |
| `QLinearMatMul` u8 M=128 | **54.9x** | 9.3x | 47.62 → 10.36 |
| `QLinearMatMul` u8 M=1 | 122.8x | 2.1x | 6.414 → 0.092 |
| `QLinearMatMul` i8 M=1 | 10.1x | **0.038x** | 5.457 → 0.092 |
| `MatMul` f32 M=128 | 1.58x | **0.82x** | 12.56 → 3.82 |
| `MatMul` f32 M=1 | 30.9x | 1.32x | 5.246 → 0.118 |
| `MatMul` f16 M=1 | 0.58x | 0.85x | 2.321 → 2.084 |
| `MatMul` f16 M=128 | 2.73x | 4.08x | 24.76 → 20.36 |

Host: AMD EPYC 9V74 (32 vCPU / 16 physical cores, AVX2+FMA+F16C, no
AVX-512), ORT 1.27.0, release build, K=N=2048, 3 warmups + 41
interleaved iterations, p50. Ratios are **ours/ORT**, so below 1.0 means
we are faster. Cold session-creation time is reported separately by the
same harness and is not folded into these numbers. The two f16 rows move
in opposite directions because that path does not go through MLAS at all
— they are the control, and they bracket the host noise for this table
(±0.3x at p50 on a shared machine).

## What this PR changes

1. **`python/nxrt-ep-cpu/setup.py` enables the feature** on every
`(system, machine)` pair in `MLAS_TARGETS`. `NXRT_EP_CPU_NO_MLAS=1`
builds the pure-Rust cdylib anyway, for a toolchain with no C++
compiler.

2. **`onnx-runtime-ep-cpu-plugin` gets its own `mlas` feature** that
forwards to `onnx-runtime-ep-cpu/mlas`, so the cdylib crate can `cfg` on
it.

3. **The cdylib now exports `nxrt_ep_build_features()`.** A compiled
library says nothing about how it was built, and here that difference is
81x. `setup.py` refuses to bundle a cdylib whose report disagrees with
what it asked cargo for (`target/release` is shared with every other
build in the checkout, so the file that exists after `cargo build` is
not necessarily the file that build produced), and the wheel's
cibuildwheel smoke test re-checks the installed artifact against
`nxrt_ep_cpu._build.EXPECTED_FEATURES`, which `setup.py` generates.
`nxrt_ep_cpu.build_features()` exposes the same fact to users.

4. **A test-harness bug that made feature-specific testing
meaningless.** `onnx_runtime_ort_testkit::find_plugin_cdylib` rebuilds
the cdylib with `cargo build -p <pkg>` — *no features*. So `cargo test
-p onnx-runtime-ep-cpu-plugin --features mlas` compiled the test binary
with MLAS and then **overwrote the MLAS cdylib with a default-feature
build**, asserting against the wrong library, silently, in the direction
that hides problems. `find_plugin_cdylib_with_features` fixes it and
`cdylib_resolve.rs` passes the features it was compiled with.
Consequence worth stating: the ORT conformance suite has now run against
the MLAS cdylib for the first time, and passes on both feature sets.

5. **CI builds the MLAS cdylib on each lane that matches a wheel
target** — `Fast (Linux x86_64)`, `Rust coverage (Windows x86_64)`,
`Rust coverage (macOS arm64)`, `Rust (Windows ARM64)`. A target is only
listed in `MLAS_TARGETS` if a lane compiles it; if these lanes go red on
some platform I will remove that platform from the set rather than ship
a wheel that fails to build at release time.

## Tests

- `l1_build_features_match_the_compiled_feature_set` (new,
`plugin_export_abi`) — dlopens the cdylib the harness resolved and
asserts its reported features equal `cfg!(feature = "mlas")`. **This is
the falsifier for item 4:** with the testkit fix reverted it fails with
`the cdylib at …/libonnx_runtime_ep_cpu_plugin.so reports features ""
but this test binary was built with "mlas"`.
- `check_wheel.py` (new) replaces the wheel's one-line `test-command`.
Falsified by hand: editing the installed `_build.py` to claim
`"avx9000"` fails with `bundled cdylib reports build features 'mlas',
but this wheel was built asking for 'avx9000'`.
- `plugin_export_abi::l1_no_symbol_leakage` — the new export is added to
the allow-list; `nm -D` on the MLAS cdylib shows the same 8 exported
symbols as before plus this one, i.e. the vendored C++ does not leak
symbols.
- Local wheel build + install + smoke test on Linux x86_64: `OK
…/libonnx_runtime_ep_cpu_plugin.so features='mlas'`.

## Verification

- `cargo test -p onnx-runtime-ep-cpu-plugin` — 50 + 9 + 7 + 1 passed
(default features)
- `cargo test -p onnx-runtime-ep-cpu-plugin --features mlas` — same
counts, all passed (first time this actually tested the MLAS build)
- `cargo clippy --all-targets` clean for `onnx-runtime-ep-cpu-plugin`
and `onnx-runtime-ort-testkit`, with and without `--features mlas`
- `cargo fmt --all -- --check` clean
- `python -m build --wheel` + install + `check_wheel.py` green locally

## Still losing after this change

With MLAS, on this host: `MatMul` f16 M=128 4.08x, `MatMulNBits` int4
M=128 7.3x and M=1 5.2x, `QLinearMatMul` u8 M=128 9.3x. Those are kernel
gaps and stay open on my task list — this PR only stops us from shipping
the *much* slower build. Nothing here reduces precision or hides setup
cost.


---

## Post-review fix (commit 2): the smoke test pointed at the wrong path

Independent review found a release-blocking bug, and it was right.

`test-command = 'python {project}/check_wheel.py'` — but cibuildwheel is
invoked as `cibuildwheel python/nxrt-ep-cpu` **from the repository
root** (`publish-ep-plugins.yml:122`), so `{project}` is the repository
root and `{package}` is this directory. All four wheel lanes would have
failed with `python: can't open file '/project/check_wheel.py'`, and
only on an `nxrt-ep-v*` tag — the release-time failure this PR exists to
prevent.

Fixed to `{package}`, and pinned by a new test binary that runs in
ordinary CI:

- `wheel_test_command_names_a_file_that_exists` — rejects `{project}`
and asserts the referenced script exists. Falsified by restoring
`{project}`: *"test-command uses {project} (the repository root)"*.
- `every_mlas_wheel_target_is_built_by_a_ci_lane` — asserts each
operating system in `MLAS_TARGETS` has a lane in `ci.yml` that compiles
the MLAS cdylib. Falsified by adding `("freebsd", "x86_64")`: *"enables
MLAS for the operating systems {"darwin", "freebsd", "linux", "windows"}
but ci.yml builds the MLAS cdylib on only 3 lanes"*. It counts operating
systems rather than targets because one coverage-matrix step covers both
`windows/amd64` and `darwin/arm64`.

Also from the review: the `comma-separated` wording in the
`nxrt_ep_build_features` doc (only one token is ever emitted), the
undocumented `AttributeError` in `build_features()`, the stale
`_mlas_features()` reference in the pyproject comment, and a note on the
testkit cache key (feature sets share one `target/<profile>` path; no
caller resolves two in one process, and the key stops the two answers
being conflated).

Rebased onto `main` after #1110; the export allow-list now carries both
that PR's counters and this PR's build-identity symbol.

### Verification (re-run after rebase)

- `cargo test -p onnx-runtime-ep-cpu-plugin --features mlas` — 52 e2e (1
ignored) + 9 + 7 + 2 + 1 passed
- `cargo test -p onnx-runtime-ep-cpu-plugin` — same counts on default
features
- `cargo fmt --all -- --check` clean **for this branch**; note `main` is
currently fmt-broken at
`crates/onnx-runtime-ep-cuda/src/kernels/device_argmax.rs`, repaired by
#1118
- `cargo clippy --all-targets` clean for both crates, with and without
`--features mlas`


---

## Second review round (commit 3): the CI guard could pass while broken

Review returned APPROVE with two MINORs that were both real, and both
are fixed.

**1. `every_mlas_wheel_target_is_built_by_a_ci_lane` was vacuous under
the exact failure it guards.** It counted text occurrences of the MLAS
build command in `ci.yml` and compared against the number of operating
systems. darwin/arm64's only MLAS build is the `if: runner.os !=
'Linux'` step on the coverage matrix, so **deleting the `macos-latest`
matrix row removes that build while the count stays at 3** and the test
stays green.

It is now structural: it parses `ci.yml`, resolves each job's runner
operating systems (expanding `runs-on: ${{ matrix.os }}` over
`matrix.include[].os` and plain `matrix.os`), applies each step's
`runner.os` condition, and asserts the union covers every operating
system in `MLAS_TARGETS`. Both falsifiers fire:

- delete the `macos-latest` row → *"setup.py enables MLAS for {"darwin",
"linux", "windows"} but no ci.yml lane compiles the MLAS cdylib on
["darwin"] … (lanes cover {"linux", "windows"})"*
- add `("freebsd", "x86_64")` → the same, for freebsd.

**2. CI built the MLAS cdylib but never tested it.** With default
features the build-identity assertion is trivially satisfied (`"" ==
""`), the vendored C++ is not linked so the leakage check has nothing to
leak, and the testkit rebuild defect this PR fixes is only observable
when features are requested — so reverting that fix would have left CI
green. The ORT-gate lane (the only one with a real ONNX Runtime) now
runs the plugin suite a **second time** with `--features mlas`, which is
where the conformance and build-identity claims are actually enforced.

**NIT:** both ctypes identity reads now release the handle in a
`finally` (on Windows a retained handle locks the DLL for the life of
the process).

**NIT — thread counts, which the tables omitted.** Every measurement
below and above uses **default `SessionOptions`** on both sides: the
harness never sets `intra_op_num_threads`, so ORT uses its own default
(physical cores — 16 on this host) and our EP uses its own pools (sized
from available parallelism — 32 vCPUs). Both sides therefore run
multi-threaded, and neither is throttled. Ratios are ours/ORT p50, so >1
means we are slower.

### Verification (re-run after rebase onto `main` @ b7fa5e1)

- `NXRT_REQUIRE_ORT_TESTS=1 cargo test -p onnx-runtime-ep-cpu-plugin
--features mlas` — 52 e2e (1 ignored) + 9 + 7 + 2 + 1 passed
- `NXRT_REQUIRE_ORT_TESTS=1 cargo test -p onnx-runtime-ep-cpu-plugin` —
identical counts on default features
- `cargo clippy --all-targets` clean for both crates, with and without
`--features mlas`
- `cargo fmt --all -- --check` clean (`main`'s unrelated fmt breakage
was repaired by #1120; my #1118 was closed as superseded)

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant