Skip to content

Split MatMulNBits CPU oracle self-checks out of the GPU numerics target (#1177) - #1477

Merged
justinchuby merged 1 commit into
mainfrom
squad/1177-split-marlin-oracle-cpu-tests
Aug 19, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/1177-split-marlin-oracle-cpu-tests

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

The CUDA test-honesty gate (CUDA compile (Linux x86_64) → Verify CUDA test inventory and skip honesty) has been red on main independent of any PR (reproduced at 84a27653, observed on a rustfmt-only PR). Root cause per #1177: crates/onnx-runtime-ep-cuda/tests/matmul_nbits_marlin_numerics.rs is a mixed target — five pure-CPU oracle self-checks pass on the CPU lane alongside three CUDA tests that correctly ignore without gpu-tests. The honesty checker sees CUDA-target tests passing without a GPU and correctly objects.

Diagnosis verified before acting (per the issue's request): the five CPU tests (oracle_matches_independent_reference_symmetric/_asymmetric, oracle_is_exact_on_a_hand_checkable_case, envelope_scales_with_output_magnitude_and_has_a_floor, parity_flags_a_perturbed_candidate) issue no CUDA calls — they only touch the device-free oracle/envelope helpers. The three GPU tests (current_path_matches_f64_oracle_group_size_sweep, ..._projection_shapes, fp16_mixed_gemv_matches_f64_oracle_glm_decode) exclusively drive run_matmul_nbits_f16/maybe_cuda. The split is clean.

Fix — split, don't weaken the checker

  • New shared non-target module tests/marlin_numerics/mod.rs holds the device-free machinery (Int4Problem, the f64 oracle, Envelope/ParityReport, f32_dequant_reference, GROUP_SIZES).
  • matmul_nbits_marlin_numerics.rs is now purely-CUDA: the 3 GPU tests + the run_matmul_nbits_f16/maybe_cuda driver, all ignored without gpu-tests.
  • New matmul_nbits_marlin_oracle.rs holds the 5 CPU self-checks.
  • matmul_nbits_marlin_oracle added to the checker's ALWAYS_RUN set (documented as a genuine CPU-only probe), and run explicitly on the CPU lane via a new CI step so the oracle math stays exercised.
  • verify_cuda_test_honesty.py pass/fail logic is untouched.

Verification (both directions, CPU/script-only — GPU left to the agent using it)

  • Honesty base-config phase passes on the fixed tree for every target: matmul_nbits_marlin_numerics = 0 passed / 0 failed / 3 ignored (now policed & clean); matmul_nbits_marlin_oracle correctly exempt. (The checker's GPU-execution phase is by design a no-CUDA-host check; the base phase is the CI: CUDA test-honesty check fails on main (matmul_nbits_marlin_numerics mixes CPU oracle tests with GPU tests) #1177-relevant half and was validated in isolation to avoid competing for the busy GPU.)
  • Oracle target runs & passes in its new home: cargo test -p onnx-runtime-ep-cuda --features cuda --test matmul_nbits_marlin_oracle → 5 passed / 0 failed / 0 ignored.
  • CUDA tests stay ignored: --test matmul_nbits_marlin_numerics → 0 passed / 0 failed / 3 ignored.
  • Both targets compile under --features cuda,gpu-tests (inventory reconciliation).
  • Checker still has teeth: it rejects a simulated pre-split shape (numerics target passing 5 tests without gpu-tests → "CUDA tests must be ignored, not pass") and accepts the post-split shape. Self-test passes.
  • My three files are rustfmt-clean and clippy-clean (pre-existing optimizer.rs/lib lints from a newer local clippy are unrelated and out of scope).

Hardware: i7-13800H / RTX 4060 Laptop, CUDA 13.1. No GPU tests were run.

Closes #1177

…et (#1177)

The CUDA test-honesty gate (`verify_cuda_test_honesty.py`) was red on `main`
because `matmul_nbits_marlin_numerics.rs` was a mixed target: five pure-CPU
oracle self-checks (validating the f64 dequant->GEMM ground truth and the
justified tolerance envelope the GPU gate depends on) passed on the CPU lane
alongside three CUDA tests that correctly ignore without `gpu-tests`. The
checker saw CUDA-target tests passing without a GPU and rightly objected.

Fix by splitting rather than weakening the checker:

- Move the device-free machinery (`Int4Problem`, the f64 oracle, the
  `Envelope`/`ParityReport` tolerance model, `f32_dequant_reference`,
  `GROUP_SIZES`) into a shared non-target module `tests/marlin_numerics/mod.rs`.
- Keep `matmul_nbits_marlin_numerics.rs` purely-CUDA: the three GPU tests plus
  the `run_matmul_nbits_f16`/`maybe_cuda` driver, all ignored without
  `gpu-tests`.
- Add `tests/matmul_nbits_marlin_oracle.rs` holding the five CPU self-checks.
- List `matmul_nbits_marlin_oracle` in the checker's `ALWAYS_RUN` set as a
  genuine CPU-only probe (documented), and run it explicitly on the CPU lane via
  a new CI step so the oracle math stays exercised.

The checker's pass/fail logic is untouched. Verified: the base-config honesty
phase passes with `matmul_nbits_marlin_numerics` at 0 passed / 3 ignored and
`matmul_nbits_marlin_oracle` exempt; the oracle target runs 5 passed / 0 failed
/ 0 ignored; both targets still compile under `cuda,gpu-tests`. The checker
still rejects the pre-split shape (a numerics target passing 5 tests without
gpu-tests).

Closes #1177

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 1d5ef75 into main Aug 19, 2026
4 checks passed
@justinchuby
justinchuby deleted the squad/1177-split-marlin-oracle-cpu-tests branch August 19, 2026 15:40
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 781.81 µs 1.14 ms +45.7%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.61 ms 2.34 ms +45.3%
🔴 matmul/large_generic_f16_threads=8/32x1024x1024 84.84 µs 121.84 µs +43.6%
🔴 grammar_masking/llguidance_compute_mask/32 68.55 µs 97.40 µs +42.1%
🔴 gather/large_f32_threads=1-internal/131072 21.53 µs 29.83 µs +38.5%
⚠️ matmul/large_generic_bf16_threads=1/32x1024x1024 1.94 ms 2.51 ms +29.4%
⚠️ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 447.70 µs 573.47 µs +28.1%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 79.17 µs 98.95 µs +25.0%
⚠️ matmul/large_generic_f32_threads=8/32x1024x1024 5.23 ms 6.45 ms +23.3%
⚠️ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.83 ms 7.04 ms +20.8%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 30.31 µs 36.50 µs +20.4%
⚠️ matmul/medium_generic_bf16_threads=8/32x512x512 497.83 µs 599.14 µs +20.4%
⚠️ kv_cache/alloc_dealloc_pages 39.07 µs 45.14 µs +15.6%
✅ logit_processing/seven_processor_chain_per_step 330.04 µs 376.45 µs +14.1%
✅ gather/medium_f16_threads=1-internal/32768 2.33 µs 2.65 µs +13.6%
✅ gather/medium_bf16_threads=1-internal/32768 2.27 µs 2.58 µs +13.5%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 547.98 µs 619.45 µs +13.0%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.54 ms 3.99 ms +12.8%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.10 ms 1.24 ms +12.5%
✅ qwen3_sampling_processors/top_k_partial_selection 149.26 µs 166.34 µs +11.4%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 87.81 µs 97.54 µs +11.1%
✅ matmul/small_generic_bf16_threads=8/1x256x256 36.90 µs 40.76 µs +10.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 666.40 µs 704.33 µs +5.7%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 114.86 µs 121.27 µs +5.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.21 ms 2.32 ms +5.2%
✅ sampling_latency/min_p_per_token 217.44 µs 225.28 µs +3.6%
✅ matmul/medium_generic_f16_threads=8/32x512x512 39.73 µs 41.03 µs +3.3%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 551.90 µs 562.06 µs +1.8%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.86 ms 9.92 ms +0.6%
✅ matmul/small_generic_bf16_threads=1/1x256x256 35.63 µs 35.84 µs +0.6%
✅ matmul/small_generic_f16_threads=8/1x256x256 36.44 µs 36.49 µs +0.1%
✅ tokenization/encode_tokens_per_second 417.48 µs 416.98 µs -0.1%
✅ sampling_latency/top_p_per_token 417.29 µs 415.74 µs -0.4%
✅ sampling_latency/greedy_per_token 3.91 µs 3.90 µs -0.4%
✅ gather/small_f32_threads=1-internal/4096 632.2 ns 628.2 ns -0.6%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.43 ms 2.42 ms -0.8%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 79.23 µs 78.39 µs -1.1%
✅ reduce_mean/large_f32_threads=1-internal/262144 977.68 µs 934.97 µs -4.4%
✅ gather/small_bf16_threads=1-internal/4096 465.8 ns 440.0 ns -5.5%
✅ gather/small_f16_threads=1-internal/4096 472.0 ns 441.0 ns -6.6%
✅ add/small_bf16_threads=1-internal/1024 447.1 ns 412.6 ns -7.7%
✅ gather/large_f16_threads=1-internal/131072 10.98 µs 10.06 µs -8.4%
✅ reduce_mean/medium_f32_threads=1-internal/65536 247.74 µs 225.65 µs -8.9%
✅ gather/medium_f32_threads=1-internal/32768 3.90 µs 3.53 µs -9.4%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.96 µs 14.33 µs -10.2%
✅ add/large_bf16_threads=1-internal/4194304 1.73 ms 1.52 ms -11.7%
✅ sampling_latency/top_k_per_token 77.28 µs 68.21 µs -11.7%
✅ add/medium_bf16_threads=1-internal/262144 109.23 µs 94.91 µs -13.1%
✅ matmul/small_generic_f16_threads=1/1x256x256 42.98 µs 37.10 µs -13.7%
✅ add/medium_f32_threads=1-internal/262144 26.57 µs 22.81 µs -14.1%
✅ add/medium_f16_threads=1-internal/262144 113.95 µs 97.21 µs -14.7%
🟢 add/large_f16_threads=1-internal/4194304 1.79 ms 1.52 ms -15.4%
🟢 matmul/small_generic_f32_threads=1/1x256x256 49.27 µs 40.98 µs -16.8%
🟢 tokenization/decode_tokens_per_second 8.06 ms 6.57 ms -18.5%
🟢 add/large_f32_threads=1-internal/4194304 734.58 µs 578.24 µs -21.3%
🟢 add/small_f16_threads=1-internal/1024 566.0 ns 434.4 ns -23.3%
🟢 add/small_f32_threads=1-internal/1024 246.6 ns 185.0 ns -25.0%
🟢 gather/large_bf16_threads=1-internal/131072 12.65 µs 9.44 µs -25.4%
🟢 matmul/small_generic_f32_threads=8/1x256x256 61.71 µs 38.18 µs -38.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 16.08 7.19 6.93 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@codecov

codecov Bot commented Aug 20, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.64%. Comparing base (a91557c) to head (bb96f3a).
⚠️ Report is 67 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1477      +/-   ##
==========================================
+ Coverage   82.10%   82.64%   +0.54%     
==========================================
  Files          12       12              
  Lines        5471     5475       +4     
  Branches     5471     5475       +4     
==========================================
+ Hits         4492     4525      +33     
+ Misses        780      757      -23     
+ Partials      199      193       -6     
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (?)
cli-ort-windows 82.10% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

justinchuby added a commit that referenced this pull request Aug 20, 2026
CI's "Verify CUDA test inventory and skip honesty" step was red for two
reasons, both of them the guard doing its job across a merge boundary.

The eager-allocator allowlist expected one malloc_sync and one free_sync
in runtime.rs. main has two of each: alloc_raw drains the raw pool and
retries rather than reporting out-of-memory while still holding device
memory back, and the frees are that drain plus free_raw. That retry
landed on main while this test was being written on the stack, and the
merge took main's runtime.rs verbatim -- byte-identical to main -- while
keeping the stack's allowlist. The counts are textual, not per seam; the
seam count is still the two disclosed ones, so the allowlist is updated
and the reason recorded next to it as the test's own message demands.

The guard also requires every integration target under the two CUDA
crates to be all-ignored without gpu-tests unless it is registered in
ALWAYS_RUN with an argument. deferred_release_queue,
vmm_release_quarantine and no_built_in_eager_allocator are the stack's
CPU-side probes and belong in that list on the same grounds as
dummy_fill_and_crossover: the first two drive state machines through
fakes and issue no CUDA calls (#636 is why the rules were moved out of
*_gpu.rs at all), and the third is a static source audit whose entire
value is proving a negative the GPU tests structurally cannot. The
guard's allowlist is on main, the probes are new here, and nothing
reconciled them.

NOT fixed here, and pre-existing on main rather than introduced by this
PR: matmul_nbits_marlin_numerics passes on a no-CUDA host with gpu-tests
because its three tests silently return instead of failing loud. That
file and the guard arrived on main in the same commit (#1477), whose own
CI reported exactly these two errors before it merged. Untouched by this
PR; filed separately rather than fixed inside a 105-file merge.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CI: CUDA test-honesty check fails on main (matmul_nbits_marlin_numerics mixes CPU oracle tests with GPU tests)

2 participants