Skip to content

style: rustfmt the two files #1116 left unformatted - #1131

Merged
justinchuby merged 1 commit into
mainfrom
deckard/fmt-1116
Aug 17, 2026
Merged

justinchuby merged 1 commit into
mainfrom
deckard/fmt-1116

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Fast (Linux x86_64) is red on main for every open PR: #1116 landed with rustfmt drift.

Diff in crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs:4950
Diff in crates/onnx-runtime-ep-cpu/src/kernels/x86_sgemm.rs:349

Verified on a pristine detached checkout of origin/main @ 1f1ce4b74, so it is not introduced by any open branch. This is pure cargo fmt --all output — no semantic change, no test change.

`Fast (Linux x86_64)` is red on main for every open PR because #1116
landed with rustfmt drift in `matmul.rs` and `x86_sgemm.rs`. Pure
`cargo fmt --all` output, no semantic change.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 2b4eaeb into main Aug 17, 2026
10 of 16 checks passed
@justinchuby
justinchuby deleted the deckard/fmt-1116 branch August 17, 2026 16:19
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/small_generic_f32_threads=8/1x256x256 39.19 µs 67.80 µs +73.0%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 55.78 µs 86.71 µs +55.5%
🔴 gather/large_bf16_threads=1-internal/131072 15.97 µs 21.04 µs +31.8%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 67.02 µs 88.26 µs +31.7%
⚠️ matmul/small_generic_f32_threads=1/1x256x256 42.19 µs 53.63 µs +27.1%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 32.47 µs 39.65 µs +22.1%
⚠️ qwen3_sampling_processors/top_k_partial_selection 141.37 µs 168.64 µs +19.3%
⚠️ qwen3_sampling_processors/top_k_full_sort_baseline 2.15 ms 2.56 ms +19.1%
⚠️ tokenization/encode_tokens_per_second 377.16 µs 440.32 µs +16.7%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 594.71 µs 670.31 µs +12.7%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 561.35 µs 630.68 µs +12.4%
✅ grammar_masking/llguidance_compute_mask/32 74.73 µs 83.25 µs +11.4%
✅ logit_processing/seven_processor_chain_per_step 323.45 µs 357.50 µs +10.5%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.49 ms 3.83 ms +9.6%
✅ matmul/medium_generic_f16_threads=8/32x512x512 33.20 µs 36.37 µs +9.5%
✅ kv_cache/alloc_dealloc_pages 39.42 µs 43.05 µs +9.2%
✅ matmul/small_generic_f16_threads=8/1x256x256 36.34 µs 39.05 µs +7.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 669.66 µs 717.96 µs +7.2%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 520.69 µs 556.95 µs +7.0%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 81.81 µs 86.43 µs +5.7%
✅ tokenization/decode_tokens_per_second 6.10 ms 6.32 ms +3.6%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.00 ms 6.22 ms +3.6%
✅ matmul/small_generic_bf16_threads=8/1x256x256 40.23 µs 40.82 µs +1.5%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 190.93 µs 191.61 µs +0.4%
✅ matmul/small_generic_bf16_threads=1/1x256x256 37.91 µs 37.59 µs -0.9%
✅ sampling_latency/greedy_per_token 3.25 µs 3.22 µs -1.1%
✅ sampling_latency/top_k_per_token 55.31 µs 53.29 µs -3.7%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.16 ms 1.09 ms -5.3%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.49 ms 2.33 ms -6.4%
✅ gather/medium_bf16_threads=1-internal/32768 2.91 µs 2.67 µs -8.2%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 499.85 µs 457.81 µs -8.4%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 11.12 ms 10.16 ms -8.6%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.69 ms 2.45 ms -8.9%
✅ sampling_latency/min_p_per_token 237.50 µs 214.55 µs -9.7%
✅ gather/small_f32_threads=1-internal/4096 829.1 ns 746.3 ns -10.0%
✅ matmul/small_generic_f16_threads=1/1x256x256 39.19 µs 35.20 µs -10.2%
✅ sampling_latency/top_p_per_token 449.62 µs 402.40 µs -10.5%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 5.69 ms 5.06 ms -11.0%
✅ reduce_mean/medium_f32_threads=1-internal/65536 341.20 µs 303.15 µs -11.2%
✅ gather/small_bf16_threads=1-internal/4096 601.6 ns 525.8 ns -12.6%
✅ gather/small_f16_threads=1-internal/4096 626.5 ns 532.8 ns -14.9%
🟢 gather/large_f16_threads=1-internal/131072 17.17 µs 14.30 µs -16.7%
🟢 reduce_mean/small_f32_threads=1-internal/4096 20.20 µs 16.73 µs -17.2%
🟢 matmul/large_generic_f16_threads=8/32x1024x1024 117.03 µs 94.00 µs -19.7%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.26 ms 1.00 ms -20.6%
🟢 gather/large_f32_threads=1-internal/131072 37.64 µs 28.85 µs -23.3%
🟢 add/large_bf16_threads=1-internal/4194304 2.38 ms 1.78 ms -25.1%
🟢 add/large_f16_threads=1-internal/4194304 2.26 ms 1.64 ms -27.4%
🟢 gather/medium_f16_threads=1-internal/32768 3.78 µs 2.70 µs -28.5%
🟢 add/medium_bf16_threads=1-internal/262144 146.82 µs 103.64 µs -29.4%
🟢 add/large_f32_threads=1-internal/4194304 945.90 µs 657.65 µs -30.5%
🟢 add/medium_f16_threads=1-internal/262144 149.66 µs 103.51 µs -30.8%
🟢 add/medium_f32_threads=1-internal/262144 38.81 µs 24.94 µs -35.7%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.61 ms 1.55 ms -40.4%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.74 ms 1.03 ms -41.1%
🟢 add/small_bf16_threads=1-internal/1024 771.9 ns 451.8 ns -41.5%
🟢 add/small_f32_threads=1-internal/1024 350.1 ns 193.7 ns -44.7%
🟢 add/small_f16_threads=1-internal/1024 814.2 ns 448.1 ns -45.0%
🟢 gather/medium_f32_threads=1-internal/32768 7.48 µs 3.91 µs -47.7%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 5.54 4.15 6.98 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants