Skip to content

style: rustfmt repair after #1100 - #1102

Merged
justinchuby merged 1 commit into
mainfrom
squad/roy-fmt-repair-1100
Aug 17, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/roy-fmt-repair-1100

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What

$ git checkout origin/main && cargo fmt --all -- --check
Diff in crates/onnx-runtime-ep-cpu/src/kernels/governed_weight_cache.rs:263
Diff in crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs:354, 2127, 2148, 2163, 3206
Diff in crates/onnx-runtime-ep-cpu/src/lib.rs:92

main is fmt-broken at a0ed99a (#1100). main is unprotected and the merge queue does not re-run CI on the merge result, so this fails the Check formatting step of both Fast (Linux x86_64) and Rust quality on every branch cut from it — including #1101, which is how I found it.

What this is

The literal output of cargo fmt --all, nothing else. No logic change, no test change. 17 insertions, 20 deletions, all whitespace and line-wrapping.

Verification

  • cargo fmt --all -- --check — clean
  • cargo test -p onnx-runtime-ep-cpu --lib — unchanged

`main` fails `cargo fmt --all -- --check` at a0ed99a, which fails every
branch cut from it. Pure rustfmt output, no logic change.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.40%. Comparing base (a0ed99a) to head (0f00684).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1102      +/-   ##
==========================================
+ Coverage   79.76%   80.40%   +0.64%     
==========================================
  Files         367      367              
  Lines      156416   156410       -6     
  Branches   156416   156410       -6     
==========================================
+ Hits       124765   125764     +999     
+ Misses      26947    25940    -1007     
- Partials     4704     4706       +2     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.31% <ø> (-0.10%) ⬇️
offline 80.28% <100.00%> (+0.66%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...untime-ep-cpu/src/kernels/governed_weight_cache.rs 89.55% <100.00%> (+0.07%) ⬆️
crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs 86.34% <100.00%> (+5.05%) ⬆️

... and 5 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@justinchuby
justinchuby merged commit 867f02d into main Aug 17, 2026
12 of 18 checks passed
@justinchuby
justinchuby deleted the squad/roy-fmt-repair-1100 branch August 17, 2026 03:13
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 40.56 µs 88.88 µs +119.1%
🔴 gather/large_f32_threads=1-internal/131072 28.94 µs 38.65 µs +33.6%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 828.73 µs 1.09 ms +32.1%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.65 ms 2.14 ms +30.0%
⚠️ matmul/large_generic_f32_threads=8/32x1024x1024 5.52 ms 6.81 ms +23.4%
⚠️ gather/medium_f32_threads=1-internal/32768 3.49 µs 4.21 µs +20.7%
✅ matmul/small_generic_f16_threads=8/1x256x256 32.88 µs 37.52 µs +14.1%
✅ gather/medium_f16_threads=1-internal/32768 2.28 µs 2.53 µs +11.3%
✅ matmul/small_generic_bf16_threads=8/1x256x256 33.72 µs 36.70 µs +8.8%
✅ reduce_mean/medium_f32_threads=1-internal/65536 234.09 µs 254.21 µs +8.6%
✅ add/large_f32_threads=1-internal/4194304 665.89 µs 721.49 µs +8.3%
✅ add/medium_f32_threads=1-internal/262144 24.61 µs 26.25 µs +6.6%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 598.75 µs 635.39 µs +6.1%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.03 ms 2.11 ms +3.8%
✅ matmul/small_generic_f16_threads=1/1x256x256 35.61 µs 36.77 µs +3.3%
✅ add/medium_f16_threads=1-internal/262144 103.35 µs 106.68 µs +3.2%
✅ add/small_f16_threads=1-internal/1024 459.6 ns 474.3 ns +3.2%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.50 µs 14.96 µs +3.2%
✅ gather/small_f16_threads=1-internal/4096 492.5 ns 501.5 ns +1.8%
✅ sampling_latency/top_k_per_token 54.51 µs 55.36 µs +1.6%
✅ add/small_bf16_threads=1-internal/1024 434.2 ns 440.9 ns +1.5%
✅ add/large_bf16_threads=1-internal/4194304 1.67 ms 1.67 ms +0.2%
✅ add/medium_bf16_threads=1-internal/262144 104.50 µs 104.50 µs +0.0%
✅ tokenization/decode_tokens_per_second 6.92 ms 6.84 ms -1.2%
✅ add/small_f32_threads=1-internal/1024 204.5 ns 200.5 ns -2.0%
✅ logit_processing/seven_processor_chain_per_step 335.30 µs 327.79 µs -2.2%
✅ gather/small_f32_threads=1-internal/4096 675.8 ns 658.5 ns -2.6%
✅ matmul/medium_generic_f16_threads=1/32x512x512 36.04 µs 35.09 µs -2.6%
✅ add/large_f16_threads=1-internal/4194304 1.73 ms 1.69 ms -2.7%
✅ matmul/small_generic_bf16_threads=1/1x256x256 38.04 µs 36.97 µs -2.8%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.02 ms 985.42 µs -3.1%
✅ grammar_masking/llguidance_compute_mask/32 82.07 µs 79.49 µs -3.1%
✅ matmul/small_generic_f32_threads=1/1x256x256 40.32 µs 38.94 µs -3.4%
✅ matmul/medium_generic_f16_threads=8/32x512x512 43.72 µs 41.81 µs -4.4%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.58 ms 2.46 ms -4.7%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 109.15 µs 103.28 µs -5.4%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 566.74 µs 525.73 µs -7.2%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 604.47 µs 555.38 µs -8.1%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.55 ms 9.68 ms -8.3%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 90.67 µs 82.61 µs -8.9%
✅ sampling_latency/greedy_per_token 3.63 µs 3.25 µs -10.3%
✅ kv_cache/alloc_dealloc_pages 45.12 µs 40.26 µs -10.8%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.13 ms 3.67 ms -11.1%
✅ qwen3_sampling_processors/top_k_partial_selection 168.54 µs 148.63 µs -11.8%
✅ sampling_latency/top_p_per_token 461.18 µs 398.33 µs -13.6%
🟢 sampling_latency/min_p_per_token 266.12 µs 226.03 µs -15.1%
🟢 gather/large_bf16_threads=1-internal/131072 16.64 µs 13.68 µs -17.8%
🟢 gather/small_bf16_threads=1-internal/4096 600.4 ns 491.8 ns -18.1%
🟢 tokenization/encode_tokens_per_second 480.29 µs 388.23 µs -19.2%
🟢 qwen3_sampling_processors/top_k_top_p_full_sort_baseline 7.25 ms 5.85 ms -19.3%
🟢 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 227.78 µs 176.27 µs -22.6%
🟢 matmul/small_generic_f32_threads=8/1x256x256 45.16 µs 34.59 µs -23.4%
🟢 qwen3_sampling_processors/top_k_full_sort_baseline 2.89 ms 2.19 ms -24.2%
🟢 gather/large_f16_threads=1-internal/131072 17.45 µs 13.02 µs -25.4%
🟢 qwen3_sampling_processors/top_k_top_p_fast 920.03 µs 665.82 µs -27.6%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.66 ms 1.15 ms -30.5%
🟢 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 785.41 µs 506.34 µs -35.5%
🟢 gather/medium_bf16_threads=1-internal/32768 4.36 µs 2.41 µs -44.6%
🟢 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 129.26 µs 68.39 µs -47.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.01 3.46 5.18 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 17, 2026
`cargo fmt --all -- --check` fails on `main` at 9b7a458 (#1104), in
`crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs:6239`. Every
open PR's `Rust quality` job therefore fails for a reason unrelated to
the PR.

This is the third time (see #1089, #1102): the merge queue does not
re-run formatting against the merge result, so an individually-green PR
can still land unformatted `main`.

Pure `cargo fmt -p onnx-runtime-ep-cpu` output, two lines, no behaviour
change.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant