Repository navigation
Split MatMulNBits CPU oracle self-checks out of the GPU numerics target (#1177) - #1477
Merged
Merged
Conversation
…et (#1177) The CUDA test-honesty gate (`verify_cuda_test_honesty.py`) was red on `main` because `matmul_nbits_marlin_numerics.rs` was a mixed target: five pure-CPU oracle self-checks (validating the f64 dequant->GEMM ground truth and the justified tolerance envelope the GPU gate depends on) passed on the CPU lane alongside three CUDA tests that correctly ignore without `gpu-tests`. The checker saw CUDA-target tests passing without a GPU and rightly objected. Fix by splitting rather than weakening the checker: - Move the device-free machinery (`Int4Problem`, the f64 oracle, the `Envelope`/`ParityReport` tolerance model, `f32_dequant_reference`, `GROUP_SIZES`) into a shared non-target module `tests/marlin_numerics/mod.rs`. - Keep `matmul_nbits_marlin_numerics.rs` purely-CUDA: the three GPU tests plus the `run_matmul_nbits_f16`/`maybe_cuda` driver, all ignored without `gpu-tests`. - Add `tests/matmul_nbits_marlin_oracle.rs` holding the five CPU self-checks. - List `matmul_nbits_marlin_oracle` in the checker's `ALWAYS_RUN` set as a genuine CPU-only probe (documented), and run it explicitly on the CPU lane via a new CI step so the oracle math stays exercised. The checker's pass/fail logic is untouched. Verified: the base-config honesty phase passes with `matmul_nbits_marlin_numerics` at 0 passed / 3 ignored and `matmul_nbits_marlin_oracle` exempt; the oracle target runs 5 passed / 0 failed / 0 ignored; both targets still compile under `cuda,gpu-tests`. The checker still rejects the pre-split shape (a numerics target passing 5 tests without gpu-tests). Closes #1177 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1477 +/- ##
==========================================
+ Coverage 82.10% 82.64% +0.54%
==========================================
Files 12 12
Lines 5471 5475 +4
Branches 5471 5475 +4
==========================================
+ Hits 4492 4525 +33
+ Misses 780 757 -23
+ Partials 199 193 -6
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
justinchuby
added a commit
that referenced
this pull request
Aug 20, 2026
CI's "Verify CUDA test inventory and skip honesty" step was red for two reasons, both of them the guard doing its job across a merge boundary. The eager-allocator allowlist expected one malloc_sync and one free_sync in runtime.rs. main has two of each: alloc_raw drains the raw pool and retries rather than reporting out-of-memory while still holding device memory back, and the frees are that drain plus free_raw. That retry landed on main while this test was being written on the stack, and the merge took main's runtime.rs verbatim -- byte-identical to main -- while keeping the stack's allowlist. The counts are textual, not per seam; the seam count is still the two disclosed ones, so the allowlist is updated and the reason recorded next to it as the test's own message demands. The guard also requires every integration target under the two CUDA crates to be all-ignored without gpu-tests unless it is registered in ALWAYS_RUN with an argument. deferred_release_queue, vmm_release_quarantine and no_built_in_eager_allocator are the stack's CPU-side probes and belong in that list on the same grounds as dummy_fill_and_crossover: the first two drive state machines through fakes and issue no CUDA calls (#636 is why the rules were moved out of *_gpu.rs at all), and the third is a static source audit whose entire value is proving a negative the GPU tests structurally cannot. The guard's allowlist is on main, the probes are new here, and nothing reconciled them. NOT fixed here, and pre-existing on main rather than introduced by this PR: matmul_nbits_marlin_numerics passes on a no-CUDA host with gpu-tests because its three tests silently return instead of failing loud. That file and the guard arrived on main in the same commit (#1477), whose own CI reported exactly these two errors before it merged. Untouched by this PR; filed separately rather than fixed inside a 105-file merge. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
This was referenced Aug 20, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The CUDA test-honesty gate (
CUDA compile (Linux x86_64)→ Verify CUDA test inventory and skip honesty) has been red onmainindependent of any PR (reproduced at84a27653, observed on a rustfmt-only PR). Root cause per #1177:crates/onnx-runtime-ep-cuda/tests/matmul_nbits_marlin_numerics.rsis a mixed target — five pure-CPU oracle self-checks pass on the CPU lane alongside three CUDA tests that correctlyignorewithoutgpu-tests. The honesty checker sees CUDA-target tests passing without a GPU and correctly objects.Diagnosis verified before acting (per the issue's request): the five CPU tests (
oracle_matches_independent_reference_symmetric/_asymmetric,oracle_is_exact_on_a_hand_checkable_case,envelope_scales_with_output_magnitude_and_has_a_floor,parity_flags_a_perturbed_candidate) issue no CUDA calls — they only touch the device-free oracle/envelope helpers. The three GPU tests (current_path_matches_f64_oracle_group_size_sweep,..._projection_shapes,fp16_mixed_gemv_matches_f64_oracle_glm_decode) exclusively driverun_matmul_nbits_f16/maybe_cuda. The split is clean.Fix — split, don't weaken the checker
tests/marlin_numerics/mod.rsholds the device-free machinery (Int4Problem, the f64 oracle,Envelope/ParityReport,f32_dequant_reference,GROUP_SIZES).matmul_nbits_marlin_numerics.rsis now purely-CUDA: the 3 GPU tests + therun_matmul_nbits_f16/maybe_cudadriver, all ignored withoutgpu-tests.matmul_nbits_marlin_oracle.rsholds the 5 CPU self-checks.matmul_nbits_marlin_oracleadded to the checker'sALWAYS_RUNset (documented as a genuine CPU-only probe), and run explicitly on the CPU lane via a new CI step so the oracle math stays exercised.verify_cuda_test_honesty.pypass/fail logic is untouched.Verification (both directions, CPU/script-only — GPU left to the agent using it)
matmul_nbits_marlin_numerics= 0 passed / 0 failed / 3 ignored (now policed & clean);matmul_nbits_marlin_oraclecorrectly exempt. (The checker's GPU-execution phase is by design a no-CUDA-host check; the base phase is the CI: CUDA test-honesty check fails on main (matmul_nbits_marlin_numerics mixes CPU oracle tests with GPU tests) #1177-relevant half and was validated in isolation to avoid competing for the busy GPU.)cargo test -p onnx-runtime-ep-cuda --features cuda --test matmul_nbits_marlin_oracle→ 5 passed / 0 failed / 0 ignored.--test matmul_nbits_marlin_numerics→ 0 passed / 0 failed / 3 ignored.--features cuda,gpu-tests(inventory reconciliation).optimizer.rs/lib lints from a newer local clippy are unrelated and out of scope).Hardware: i7-13800H / RTX 4060 Laptop, CUDA 13.1. No GPU tests were run.
Closes #1177