Skip to content

perf(cuda): derive RMSNorm fold hidden floor from SM count (#1421) - #1582

Merged
justinchuby merged 3 commits into
mainfrom
squad/1421-rmsnorm-device-aware
Aug 20, 2026
Merged

justinchuby merged 3 commits into
mainfrom
squad/1421-rmsnorm-device-aware

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 20, 2026 •

Copy link
Copy Markdown
Owner

perf(cuda): derive the RMSNorm-fold hidden floor from SM count (#1421)

Replaces the global RMSNORM_FUSION_MIN_HIDDEN = 1280 with a per-device floor
derived from the device's SM count, anchored on the single H200 calibration point
rather than a per-device lookup table.

Closes #1421.

What the change does

derived_min_hidden(sm) = round_nearest(10 * sm / 132) chunks * 128, clamped to >= 1 chunk
  • Derived from multiprocessor_count() — the property that sets how much
    parallel slack the device has to keep the standalone norm "almost free" at M=1.
  • Not a device table. A single H200 anchor (132 SM → 1280) fixes the slope;
    every other device gets its floor from its own SM count. A big datacenter part
    inherits the protective H200-class floor; a small edge part folds aggressively.
    Avoids the §40 "regime-tuned threshold" trap.
  • Reproduces both known data points: H200 (132 SM) → 1280 (exactly);
    RTX 4060 Laptop (24 SM) → 256.
  • M-invisibility finding: the optimizer cannot see M. This pass runs once at
    model load on a decode graph whose batch dim is symbolic and shared across all
    batch sizes, so M ≥ 2 cannot gate the fold per step. The residual uncovered
    case (small model batched on a large GPU) is documented and left to the
    ONNX_GENAI_RMSNORM_MIN_HIDDEN escape hatch.

Tests (measured)

cargo test -p onnx-runtime-ep-cuda --features "cuda,gpu-tests" --lib -- optimizer::tests
→ 85 passed, 0 failed, 0 ignored (10x default-parallel runs, 0 failures — the new tests route env access through EnvVarGuard per #1588, closing the process-wide env race; the +1 vs the earlier count is #1588's guard test). (cuda,gpu-tests, not cuda alone, so the
suite is not silently all-ignored.) Four new derived_min_hidden tests cover the
anchor self-reproduction (132→1280), 24→256, .max(1) clamp at sm 0/1,
monotonicity, whole-128-chunk output, the device-derived gate, and env-override
precedence.

The fold is NOT bit-identical — and this PR corrects that long-unverified claim

The SkipSimplifiedLayerNormalization → GEMV fold was documented in the code as
"byte-identical". That was never verified, and it is false. Folding the RMS
normalization into the following GEMV's prologue changes the fp16 reduction
order
, so at a greedy-argmax near-tie it can flip a token. The residual epilogue
(fp16(fp16(acc) + residual) == __hadd2) remains bit-identical; the reordering
is in the normalization reduction only. The code comments are updated to say this
accurately (see the second commit).

Measured (binary SHA256-pinned across runs, machine quiet)

Switch efficacy proven with ONNX_GENAI_PROFILE_OPS=1:
SkipSimplifiedLayerNormalization op-count 23 (fold ON) vs 48 (fold OFF) — the
two sides genuinely run different graphs.

profile_native (native CUDA, greedy, --tokens 128 --decode-skip 8, no
--steady), qwen0.5B (hidden 896), fold ON (ONNX_GENAI_RMSNORM_MIN_HIDDEN=256)
vs OFF (=2048):

Prompt ON vs OFF
default, "Once upon a time…", "Explain quantum computing…", "def fibonacci(n):", "The capital of France is" identical
"The quick brown fox jumps over the lazy dog and then" diverges at token 49 (ON=448, OFF=304)

granite-1b-a400m (hidden 1024): identical on the fox prompt. So the divergence is
real but rare and input-dependent (1 of 6 sampled prompts), consistent with a
reduction-order change tipping only near-ties.

Which side is correct? The folded side.

Divergence alone doesn't say which token is right. Adjudicated with a
high-precision oracle: take prompt (11 tok) + generated[0..48] (49 tok) = a
60-token prefix and run a full prefill on the CPU EP (a completely different
implementation and precision the CUDA fold never touches), dumping top-k:

n_prompt_tokens: 60
selected_token:  448
top: [[448, -1.58743], [304, -1.60305], [438, -1.82180], ...]

The reference top-1 is 448 — the token fold ON selects. Fold OFF (today's
default: keep the standalone norm) selects 304, the runner-up, only 0.0156
nats
behind. So the fp16 reduction reorder rounded toward the more accurate
result here, not away from it. This is a reduction-order effect, not a fusion
bug
— there is no evidence of a numerical defect in the N=896 fused GEMV.

Repro

# fold ON vs OFF, compare generated_token_ids (differ at index 49):
$env:ONNX_GENAI_RMSNORM_MIN_HIDDEN='256'
profile_native --model <qwen05b> --ep cuda --backend native --tokens 128 --decode-skip 8 --prompt "The quick brown fox jumps over the lazy dog and then"
$env:ONNX_GENAI_RMSNORM_MIN_HIDDEN='2048'
# (same command)

# oracle: --prompt-ids takes a FILE PATH to a JSON array, not inline JSON.
# Write the 60-token prefix [prompt..+generated[0..48]] to a file, then:
profile_native --model <qwen05b> --ep cpu --backend ort --prompt-ids <prefix.json> --tokens 1   # dump top-k

Not verified / out of scope

  • Performance: not measured (box not guaranteed quiet across the session).
  • Generality of the "folded side is more accurate" result: established at the
    single measured flip point only; not proven for every possible near-tie.
  • H200 behavior (132 SM → floor 1280, so 896/1024 stay unfused) verified via the
    derived value in unit tests, not on H200 hardware.

@justinchuby justinchuby changed the title perf(cuda): derive RMSNorm fold hidden floor from SM count (#1421) [DRAFT: byte-identity fails on qwen0.5B/896] perf(cuda): derive RMSNorm fold hidden floor from SM count (#1421) [DRAFT: fold not strictly byte-identical on qwen0.5B/896, rare/input-dependent] Aug 20, 2026
@justinchuby justinchuby changed the title perf(cuda): derive RMSNorm fold hidden floor from SM count (#1421) [DRAFT: fold not strictly byte-identical on qwen0.5B/896, rare/input-dependent] perf(cuda): derive RMSNorm fold hidden floor from SM count (#1421) Aug 20, 2026
@justinchuby
justinchuby marked this pull request as ready for review August 20, 2026 15:58
@justinchuby
justinchuby force-pushed the squad/1421-rmsnorm-device-aware branch from 72da204 to 85c9e35 Compare August 20, 2026 16:06
@codecov

codecov Bot commented Aug 20, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.78%. Comparing base (4c1b594) to head (bf20e41).
⚠️ Report is 8 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@             Coverage Diff             @@
##             main    #1582       +/-   ##
===========================================
- Coverage   82.19%   80.78%    -1.42%     
===========================================
  Files          12      381      +369     
  Lines        5471   173786   +168315     
  Branches     5471   173786   +168315     
===========================================
+ Hits         4497   140392   +135895     
- Misses        775    28500    +27725     
- Partials      199     4894     +4695     
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (?)
cli-ort-windows 82.19% <ø> (ø)
mlas 85.19% <ø> (?)
offline 80.64% <ø> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 372 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

justinchuby added a commit that referenced this pull request Aug 20, 2026
…1588)

## Problem

`cargo test -p onnx-runtime-ep-cuda --features "cuda,gpu-tests" --lib --
optimizer::tests` was flaky in default parallel mode (~1 failure in 5
runs); `--test-threads=1` never failed.

`cargo test` runs tests as parallel threads **inside one process**, so
`std::env` is process-wide shared mutable state. Several optimizer tests
mutated env vars under a `// SAFETY: single-threaded test` comment that
was **never true** under the default harness. The proven failure:

- `opt_out_env_preserves_exported_gate_chains` does
`set_var(LINEAR_ATTENTION_GATING_DISABLE_ENV, "1")` -> run ->
`remove_var`.
- `precomputed_neg_exp_a_initializer_does_not_set_neg_exp_marker` sets
**no** env but asserts the fusion is enabled by default. If it runs
inside the other test's set/remove window, it observes the disable flag
and panics.

## Fix — structural, not a workaround

New reusable `EnvVarGuard` in `test_support` (RAII over a
**process-global mutex**):

- Serialises **every** env-touching test — writers *and* default-readers
— on one lock. The default-reader half matters: a plain writer-only
mutex would leave readers exposed, so both sides take the same lock and
the race becomes impossible, not merely unlikely.
- Restores each touched variable to its prior value on drop, even on
panic.
- `acquire()` (lock only), `with_var`/`without_var`, and `set`/`unset`
for read-then-toggle tests.

All env-touching `optimizer.rs` tests converted, covering **all five**
variables set there: linear-attention gating (the proven-flaky one),
skip-rmsnorm, const-cast-fold, identity-cast-fold, qkv-fusion. The
ad-hoc `IDENTITY_CAST_TEST_LOCK` (which only ever covered one var) is
replaced by the shared guard, and every false `// SAFETY:
single-threaded test` comment is deleted/replaced with an accurate one.

## Behaviour change you should know about — tests are now immune to
external env (intentional)

This is a real, deliberate behaviour change, called out explicitly
rather than buried:

The default-reader tests now `unset` their variable inside the guard, so
they assert the **declared default**, not whatever the invoking shell
happens to export. Measured on the previously-flaky target with
`ONNX_GENAI_CUDA_DISABLE_LINATTN_GATING_FUSION=1` injected
**externally**:

| | external `DISABLE_LINATTN=1` set before the run |
|---|---|
| before this PR (measured) | `exit=101` `FAILED. 81 passed; 3 failed` |
| after this PR (measured, this branch) | `exit=0` `ok. 81 passed; 0
failed; 0 ignored` |

Why this is correct: a test that asserts *default behaviour* must run
under a deterministic default. Letting the outcome depend on the
runner's ambient environment is just another form of the same
non-determinism this PR removes — the pre-fix "3 failed" were the same
bug's other face (the tests were reading the **process environment**
instead of their **declared premise**).

**The cost, stated plainly:** if someone *wants* to force a whole "this
fusion disabled" test pass by exporting `ONNX_GENAI_CUDA_DISABLE_*` and
running the suite, that path **no longer works for guard-covered tests**
— they pin the variable themselves. That is a genuine trade-off; the
per-test opt-out assertions (e.g. `opt_out_env_preserves_*`) remain the
supported way to exercise the disabled path.

## Regression guard (falsifiable)

Added a source-level test,
`optimizer_source_routes_all_env_mutation_through_guard`, asserting
`optimizer.rs` contains no bare `std::env::set_var`/`remove_var`
(needles assembled from fragments so it never matches itself). If a
future edit re-introduces a direct env mutation here, this test fails
immediately instead of the flake resurfacing under the parallel harness.
Production code in this file only *reads* the environment, which is
race-free, so a zero-mutation source is the correct invariant. Scope:
`optimizer.rs` only; the follow-up issue extends the same guard per file
as they are swept.

## Test-only, no runtime footprint (measured)

`test_support` is `#[cfg(test)]`-gated. Verified against the non-test
dev rlib (`cargo build -p onnx-runtime-ep-cuda --features cuda`):

```
occurrences of 'EnvVarGuard' in non-test rlib = 0
occurrences of 'env_lock'    in non-test rlib = 0
occurrences of 'ENV_LOCK'    in non-test rlib = 0
```

(The one `test_support` string hit is a debuginfo file path, not
compiled code.) No guard code enters a non-test/release binary.

## Evidence (measured)

`optimizer::tests` run **25x in default parallel mode** via the built
`cuda,gpu-tests` binary (20 in the first pass + 5 after adding the guard
test): zero failures. Representative line:

```
test result: ok. 81 passed; 0 failed; 0 ignored; 0 measured; 398 filtered out
```

(`81` includes the new regression-guard test; the pre-guard passes
reported `80`.)

`cargo fmt -p onnx-runtime-ep-cuda -- --check` exits 0.

## Scope / deferred — tracked in #1591

- Fully fixes `optimizer.rs` (the proven-flaky file) plus its four
latent siblings, with a reusable guard and regression check.
- `EnvVarGuard` lives in `test_support` so other files
(`matmul_nbits.rs` 33 sites, `provider.rs`, etc.) can adopt the same
pattern. **#1591** tracks the sweep and records the decision criterion:
*any variable that some test sets must have its default-readers hold the
same lock*.
- Passes reading vars that **no test sets** (e.g. L2-norm fusion,
rmsnorm min-hidden) are left unlocked — no writer means no race.

## Coordination

Rebased onto latest `main` (`1c97c9693`, includes #1491). Diff is
confined to `optimizer.rs` tests + `test_support.rs` — no overlap with
#1491's `standard_attention.rs`/`geometry.rs`. #1582 (RMSNorm) is not
yet merged; it edits `optimizer.rs` production code + adds tests,
whereas this PR only changes existing tests' env handling, so the
conflict surface is small.

---------

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Copilot AI added 3 commits August 20, 2026 09:40
Replace the global RMSNORM_FUSION_MIN_HIDDEN=1280 with a per-device floor
derived from the device's SM count, anchored on the single H200 (132 SM)
calibration point rather than a per-device lookup table. A small GPU (few
SMs) folds aggressively; a large GPU keeps the protective H200-class floor
at M=1. The optimizer cannot see M (the decode graph's batch dim is
symbolic and shared across batch sizes), so the SM-derived floor covers the
common small-GPU batch case and ONNX_GENAI_RMSNORM_MIN_HIDDEN remains the
escape hatch for a small model batched on a large GPU. The fold stays
byte-identical.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
The SkipSimplifiedLayerNormalization -> GEMV fold was documented as
bit-for-bit identical to the standalone norm. That was never verified and
is false: folding the RMS normalization into the following GEMV's prologue
changes the fp16 reduction order, so at a greedy-argmax near-tie it can flip
a token. Measured on qwen0.5B (hidden 896): 1 of 6 sampled prompts flipped
one token at a 0.0156-nats top-2 gap. A CPU-EP full-prefill oracle put the
folded token (448) closer to the high-precision reference than the unfused
one (304), so the reorder rounds toward the more accurate result -- a
reduction-order effect, not a fusion bug. Update the load-bearing comments
accordingly; the residual epilogue remains bit-identical.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Rebasing onto #1588 replaced the ad-hoc RMSNORM_FLOOR_ENV_LOCK and bare
std::env::set_var/remove_var in the new floor tests with the crate-wide
EnvVarGuard, which serialises env access on a process-global lock (writers
and default-value readers share it) and restores touched vars on drop. The
two device-derived gate tests read the real ONNX_GENAI_RMSNORM_MIN_HIDDEN
default, so they now hold the guard too. Satisfies the include_str! guard
test that bans bare env mutation in optimizer.rs.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
@justinchuby
justinchuby force-pushed the squad/1421-rmsnorm-device-aware branch from 85c9e35 to bf20e41 Compare August 20, 2026 16:51
@justinchuby
justinchuby merged commit ccf1de5 into main Aug 20, 2026
7 of 14 checks passed
@justinchuby
justinchuby deleted the squad/1421-rmsnorm-device-aware branch August 20, 2026 17:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Batch-decode SwiGLU capture-safe fix (#1404) is gated OFF by default for small resident models: RMSNorm-fold hidden floor (1280) > qwen05b hidden (896)

2 participants