Skip to content

style: rustfmt the four files main is currently unformatted in - #1168

Closed
justinchuby wants to merge 1 commit into
mainfrom
squad/resch-fmt-main
Closed

justinchuby wants to merge 1 commit into
mainfrom
squad/resch-fmt-main

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What

cargo fmt --all --check fails on main (84a2765). Four files are unformatted:

file why
crates/mlas-sys/src/lib.rs QGEMM_PACKED_LIVE_BYTES.fetch_sub call wrapping (#1151)
crates/onnx-runtime-ep-cpu/src/kernels/governed_accumulator_budget.rs assert! wrapping (#1151)
crates/onnx-runtime-ep-cpu/src/kernels/qlinear_matmul.rs two call sites (#1133/#1151)
crates/onnx-runtime-ep-cpu/src/kernels/simd_activations.rs trailing blank line (#1136/#1143)

Whitespace and call-site wrapping only — no token changes, no behaviour change.

Evidence

$ cargo fmt --all --check   # before: 4 files reported
$ cargo fmt --all
$ cargo fmt --all --check   # after: clean, exit 0

Split out of the no-defer PR so that one carries no unrelated churn. This is the
same shape as #1131 and #1144.

`cargo fmt --all --check` fails on `main` at 84a2765: #1133/#1151's
accumulator work left `mlas-sys/src/lib.rs`,
`governed_accumulator_budget.rs` and `qlinear_matmul.rs` unformatted, and
`simd_activations.rs` carries a trailing blank line. No behaviour change --
whitespace and call-site wrapping only.

Verified: `cargo fmt --all --check` is clean afterwards.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby enabled auto-merge (squash) August 18, 2026 01:47
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 5.26 ms 6.93 ms +31.7%
⚠️ matmul/large_generic_bf16_threads=1/32x1024x1024 2.04 ms 2.54 ms +24.3%
⚠️ matmul/large_generic_bf16_threads=8/32x1024x1024 1.90 ms 2.34 ms +23.1%
⚠️ gather/large_bf16_threads=1-internal/131072 14.95 µs 17.97 µs +20.2%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 95.87 µs 109.34 µs +14.1%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 102.05 µs 111.54 µs +9.3%
✅ gather/large_f32_threads=1-internal/131072 36.58 µs 39.42 µs +7.8%
✅ gather/large_f16_threads=1-internal/131072 13.01 µs 14.02 µs +7.8%
✅ tokenization/encode_tokens_per_second 386.71 µs 415.83 µs +7.5%
✅ qwen3_sampling_processors/top_k_partial_selection 143.30 µs 153.77 µs +7.3%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.66 ms 1.78 ms +7.2%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.94 µs 15.87 µs +6.2%
✅ matmul/medium_generic_f16_threads=1/32x512x512 36.11 µs 38.13 µs +5.6%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.55 ms 10.03 ms +5.0%
✅ gather/small_f32_threads=1-internal/4096 653.6 ns 686.4 ns +5.0%
✅ add/medium_bf16_threads=1-internal/262144 102.58 µs 106.30 µs +3.6%
✅ add/medium_f16_threads=1-internal/262144 103.41 µs 107.12 µs +3.6%
✅ tokenization/decode_tokens_per_second 6.35 ms 6.56 ms +3.2%
✅ gather/small_f16_threads=1-internal/4096 482.4 ns 498.0 ns +3.2%
✅ matmul/small_generic_f16_threads=1/1x256x256 36.87 µs 37.90 µs +2.8%
✅ matmul/small_generic_f16_threads=8/1x256x256 36.78 µs 37.69 µs +2.5%
✅ sampling_latency/min_p_per_token 208.29 µs 212.41 µs +2.0%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 633.03 µs 644.98 µs +1.9%
✅ matmul/small_generic_f32_threads=8/1x256x256 49.37 µs 50.04 µs +1.4%
✅ sampling_latency/top_k_per_token 52.98 µs 53.57 µs +1.1%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.57 ms 3.61 ms +1.1%
✅ matmul/small_generic_bf16_threads=1/1x256x256 37.16 µs 37.50 µs +0.9%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.15 ms 2.16 ms +0.7%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.35 ms 2.36 ms +0.5%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 174.73 µs 175.56 µs +0.5%
✅ add/large_bf16_threads=1-internal/4194304 1.72 ms 1.73 ms +0.4%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 675.90 µs 677.34 µs +0.2%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.84 ms 5.85 ms +0.2%
✅ add/small_f16_threads=1-internal/1024 461.0 ns 461.5 ns +0.1%
✅ sampling_latency/top_p_per_token 399.01 µs 399.42 µs +0.1%
✅ add/medium_f32_threads=1-internal/262144 25.03 µs 25.00 µs -0.2%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.00 ms 1.00 ms -0.3%
✅ kv_cache/alloc_dealloc_pages 40.24 µs 40.08 µs -0.4%
✅ add/large_f32_threads=1-internal/4194304 889.50 µs 884.58 µs -0.6%
✅ matmul/small_generic_bf16_threads=8/1x256x256 38.85 µs 38.59 µs -0.7%
✅ gather/medium_bf16_threads=1-internal/32768 2.40 µs 2.38 µs -0.7%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 541.37 µs 537.39 µs -0.7%
✅ gather/medium_f16_threads=1-internal/32768 2.43 µs 2.41 µs -0.8%
✅ add/small_bf16_threads=1-internal/1024 449.8 ns 445.1 ns -1.0%
✅ sampling_latency/greedy_per_token 3.31 µs 3.27 µs -1.2%
✅ grammar_masking/llguidance_compute_mask/32 79.65 µs 78.51 µs -1.4%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.16 ms 1.14 ms -2.0%
✅ gather/medium_f32_threads=1-internal/32768 4.08 µs 4.00 µs -2.0%
✅ logit_processing/seven_processor_chain_per_step 332.34 µs 325.21 µs -2.1%
✅ gather/small_bf16_threads=1-internal/4096 495.2 ns 484.5 ns -2.2%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 534.69 µs 521.89 µs -2.4%
✅ qwen3_sampling_processors/top_k_top_p_fast 682.83 µs 663.99 µs -2.8%
✅ reduce_mean/medium_f32_threads=1-internal/65536 257.22 µs 248.51 µs -3.4%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 129.13 µs 124.16 µs -3.8%
✅ matmul/small_generic_f32_threads=1/1x256x256 47.81 µs 45.72 µs -4.4%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 91.15 µs 86.99 µs -4.6%
✅ matmul/medium_generic_f16_threads=8/32x512x512 43.76 µs 41.49 µs -5.2%
✅ add/large_f16_threads=1-internal/4194304 1.83 ms 1.71 ms -6.7%
🟢 add/small_f32_threads=1-internal/1024 252.7 ns 194.4 ns -23.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.31 3.08 3.67 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

Superseded by #1172, which landed the same four files. Closing.

@codecov

codecov Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 88.88889% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 79.84%. Comparing base (84a2765) to head (06b6417).
⚠️ Report is 13 commits behind head on main.

Files with missing lines Patch % Lines
.../onnx-runtime-ep-cpu/src/kernels/qlinear_matmul.rs 75.00% 0 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main    #1168   +/-   ##
=======================================
  Coverage   79.84%   79.84%           
=======================================
  Files         359      359           
  Lines      157243   157245    +2     
  Branches   157243   157245    +2     
=======================================
+ Hits       125549   125552    +3     
+ Misses      27136    27135    -1     
  Partials     4558     4558           
Flag Coverage Δ
mlas 85.39% <100.00%> (+0.03%) ⬆️
offline 79.73% <83.33%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/mlas-sys/src/lib.rs 84.78% <100.00%> (+<0.01%) ⬆️
...-ep-cpu/src/kernels/governed_accumulator_budget.rs 88.67% <100.00%> (+0.10%) ⬆️
...nnx-runtime-ep-cpu/src/kernels/simd_activations.rs 88.58% <ø> (ø)
.../onnx-runtime-ep-cpu/src/kernels/qlinear_matmul.rs 83.60% <75.00%> (ø)

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant