Skip to content

perf(cuda): split-KV FlashDecoding for attention_row decode (+10% V2-Lite wide-ctx) - #1340

Merged
justinchuby merged 1 commit into
mainfrom
squad/attention-splitkv-flashdecode
Aug 18, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/attention-splitkv-flashdecode

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

attention_row launches one block per (batch, q_head, query) row, so at decode it uses only ~16 of 132 SMs and is memory-latency bound. Profiling V2-Lite eager decode at wide context (GPU3, ~390-word prompt, 256 tok) showed attention_row is the #1 decode kernel at 36% of GPU decode — ncu: grid=16, waves/SM=0.02 (~116 idle SMs), DRAM 0.79%, achieved occupancy 9.8% → a definitive under-utilization signal.

This PR adds a two-kernel FlashDecoding split-KV path that spreads each row's key reduction across many blocks to fill the idle SMs.

Design

Behind ONNX_GENAI_ATTN_SPLITKV (default ON; =0 reproduces the monolithic baseline for A/B):

  • attention_split — grid (total_rows, num_splits). Each block reduces one KV slice with a chunk-local online softmax (running max + rescale), writing an unnormalized partial P·V plus per-split max/sum meta. A single-split fast path (total_seq <= chunk) reproduces attention_row exactly and writes the final normalized output with sentinel meta (max=0, sum=1).
  • attention_combine — grid (total_rows). Uniform log-sum-exp merge of the per-split partials. The single-split sentinel makes it a bit-exact pass-through, so any context <= chunk (default 256) stays byte-identical to the monolithic kernel.

Capture-safety: num_splits and the engage decision derive only from the fixed KV capacity (cap), never live seqlen, so eager and capture make identical launch decisions → the eager==capture long-context lock holds. Split scratch is drawn from the internal capture-warmed ws pool (WS_SPLIT slot), leaving the governed workspace layout (and its exact-byte unit tests) untouched.

Multi-split reorders fp32 partials, so wide context is validated to the #1150 f64 tolerance (2e-4.max(·*2e-5)), not byte-identity. Greedy token output is unchanged in practice.

Numerics gates (H200, GPU3, pinned, single-threaded)

Gate Result
V2-Lite golden 24-tok lock (matches_golden) ✅ byte-identical (single-split path)
V2-Lite 340-tok long-context lock (eager==capture + golden prefix) ✅ pass (multi-split engaged, capture-safe)
standard_attention_capture_gpu / _fp16_gpu / _bf16_gpu ✅ pass
standard_attention_gpu ✅ 23/24 (the 1 failure ..._requires_homogeneous_floating_input_dtypes is pre-existing on clean origin/main, unrelated)
layout unit tests (--lib standard_attention) ✅ 14/14
Dense qwen2.5-0.5b token md5 ✅ byte-identical (no regression)

Perf A/B (H200 GPU3, pinned, medians-of-5)

Wide-context kernel (nsys, same binary, split ON vs =0):

decode attention kernel (median) DRAM throughput
baseline (attention_row) 147.6 µs 0.79%
split (attention_split+combine) ~62.5 + 2.5 µs (~2.3× faster) 9.26%

E2E throughput (tok/s):

context split OFF (=0) split ON (default) Δ
wide (~390-word prompt, 128 tok) 51.64 56.90 +10.2%
short "Hello" eager (128 tok) 62.25 63.04 +1.3% (neutral)
short "Hello" capture (128 tok) 128.48 128.17 −0.2% (neutral)
dense qwen2.5-0.5b capture 610.36 610.25 neutral

num_splits = cap.div_ceil(chunk).clamp(1, 64) with chunk default 256 (tunable via ONNX_GENAI_ATTN_SPLIT_CHUNK). At the true long-context regime where attention_row reaches ~40% of decode (cap≈2048), this yields ~8 splits → grid ~128, filling the machine; smaller chunks extract a bit more at moderate context at the cost of engaging split earlier.

Verdict: GO, default-ON

Short context stays byte-identical (comfortable margin below chunk); wide context gets +10.2% E2E / ~2.3× kernel; capture-safe; no dense regression. All hard gates green.

Do not self-merge — reporting to coordinator for admin-merge after review.

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

…Lite wide-ctx)

attention_row launches one block per (batch, q_head, query) row, so decode
uses only ~16 of 132 SMs and is memory-latency bound (~36% of GPU decode at
wide context; ncu: grid=16, waves/SM=0.02, DRAM 0.79%, occupancy 9.8%).

Add a two-kernel FlashDecoding split-KV path behind ONNX_GENAI_ATTN_SPLITKV
(default ON):
- attention_split: grid (total_rows, num_splits). Each block reduces one KV
  slice with chunk-local online softmax (running max + rescale), writing an
  unnormalized partial P.V plus per-split max/sum meta. A single-split fast
  path (total_seq <= chunk) reproduces attention_row exactly and writes the
  final normalized output with sentinel meta (max=0, sum=1).
- attention_combine: grid (total_rows). Uniform log-sum-exp merge of the per
  -split partials. The single-split sentinel makes it a bit-exact pass-through,
  so contexts <= chunk stay byte-identical to the monolithic kernel.

Capture-safe: num_splits and the engage decision derive only from the fixed KV
capacity (cap), never live seqlen, so eager and capture make identical launch
decisions. Split scratch is drawn from the internal capture-warmed ws pool
(WS_SPLIT slot), leaving the governed workspace layout untouched.

Multi-split reorders fp32 partials so wide context is validated to the #1150
f64 tolerance rather than byte-identity; greedy token output is unchanged.

Gates (H200, GPU3, pinned single-threaded):
- V2-Lite golden 24-tok lock: byte-identical (single-split path).
- V2-Lite 340-tok long-context lock: eager==capture holds with multi-split
  engaged; golden prefix matches.
- standard_attention capture/fp16/bf16 gpu-tests + layout unit tests pass.
- Dense qwen2.5-0.5b: byte-identical, tok/s neutral.
- Wide-ctx kernel: attention_row 147.6us -> attention_split+combine ~65us
  (~2.3x); DRAM throughput 0.79% -> 9.26%.
- E2E wide-ctx: 51.64 -> 56.90 tok/s (+10.2%); short-ctx neutral.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 763d81f into main Aug 18, 2026
6 checks passed
@justinchuby
justinchuby deleted the squad/attention-splitkv-flashdecode branch August 18, 2026 23:11
justinchuby added a commit that referenced this pull request Aug 18, 2026
Records #1340 (split-KV/FlashDecoding for attention_row: ~2.3× kernel /
+10.2% E2E wide-ctx, merged @ 763d81f) in the campaign brain. Docs-only
(.squad/identity/now.md).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 18, 2026
…log (PR #1337, #1340) (#1343)

## Summary

Scribe consolidation of two Deckard decision inbox notes into the
canonical `.squad/decisions.md` ledger.

### Decisions logged

**PR #1337 — attention_row block-width 128→256 (GO, byte-identical,
merged 37bdefe)**
- Profile evidence: grid=16, achieved occupancy 6.23%, waves/SM=0.01,
62.7% barrier-stall cycles — memory-latency bound, not roofline-limited.
- attention_row context scaling: 12.4% GPU decode share at short ctx →
40% at deep ctx → **#1 decode kernel at realistic/wide context** (QMoE
levers exhausted).
- Numerics gate: BYTE-IDENTICAL (same accumulation order, more warps).
- H200 A/B: −14% attention_row kernel latency, **+5.5% E2E** on V2-Lite
~500-tok prompt; neutral on short/dense.

**PR #1340 — split-KV FlashDecoding for attention_row (GO, default-ON,
merged 763d81f)**
- Profile evidence: grid=16, waves/SM=0.02, ~116 idle SMs, DRAM 0.79%,
occupancy 9.8% — machine starved.
- Single-split sentinel fast path: contexts ≤ chunk (default 256) →
**byte-identical pass-through**.
- Wide-ctx (multi-split): reordered fp32 partials → **f64-tol** (#1150
oracle); capture-safe (num_splits derived from fixed KV cap, not live
seqlen).
- H200 A/B: ~2.3× kernel speedup, DRAM 0.79%→9.26%, **+10.2% E2E** wide
ctx; neutral short/dense.

### Archive gate
All entries in decisions.md are dated 2026-08-18 (active campaign arc) —
nothing >30 days old. No archiving performed per the size-not-age rule.

### File scope
- Modified: `.squad/decisions.md` (+86 lines)
- Deleted (untracked, rm only):
`.squad/decisions/inbox/deckard-attention-row.md`,
`.squad/decisions/inbox/deckard-attention-splitkv.md`

Do NOT self-merge — awaiting coordinator review.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 18, 2026
Two independent failures on `main` (c55a3fa), both in files the
failing PRs do not touch, both landed while the Actions queue was backed
up so nothing caught them at merge time. Between them they make "is this
PR's CI green?" unanswerable for every open PR.

1. `cargo fmt --all --check` fails at `standard_attention.rs:2125` —
   the split-KV workspace sizing from #1340 is unformatted. Whitespace
   only, and it is rustfmt's own output on rustc 1.97.1, the toolchain CI
   resolves `stable` to.
2. `verify_documented_env_vars.py` fails on
   `ONNX_GENAI_DECODE_GEMV_PROBE_ROWS`, documented by
   `2026-08-18-multirow-gemv-ceiling-probe.md`. That doc's method section
   already says the probe was "reverted after measuring" — it is
   deliberately not shipped, which is exactly what `KNOWN_UNIMPLEMENTED`
   is for, so it goes there with that reason rather than being wired up.

After this commit every step of the `Rust quality` lane passes locally:
`cargo fmt --all --check`, `check_publish_order`,
`workspace_test_packages verify` (49 tested, 5 denied),
`check_profile_table`, `check_platform_naming`,
`check_dispatch_reachability`, `check_dispatch_manifest` (+ self-test),
`check_feature_gate_coverage`, and `verify_documented_env_vars`
(106 documented, 13 known-unimplemented, all accounted for).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 18, 2026
…over fixed-256; +70% deep-ctx) (#1350)

## Summary

Follow-up to #1340 (split-KV FlashDecoding for `attention_row`). That PR
fixed keys-per-split at `chunk=256`, so `num_splits = cap/256`.
Re-profiling at **deep context** (~2600-tok prompt, the regime where
`attention_row` reaches ~40% of decode) showed two things:

1. **The split-KV win is far bigger at depth than the shallow A/B in
#1340 measured:** deep-ctx E2E is **28.40 → 46.37 tok/s (+63%)** with
fixed-256, vs the +10% originally reported at ~500-ctx.
2. **Fixed-256 under-fills the machine at depth.** At `cap≈4096` the
grid is `(16 rows × 16 splits) = 256` blocks, waves/SM 0.48, warps
active 11%. The decode split kernel is memory-latency bound, so it wants
**~a full occupancy wave of resident CTAs, not one block per SM** — an
ncu + E2E chunk sweep found the optimum at `chunk=128` (grid 512), not
`chunk=512` (grid 132, which is ~5% slower).

This PR sizes `num_splits` **adaptively** from the fixed `cap` **and**
the fixed row count to target ~one occupancy wave, floored so no split
is starved of work:

```
target_splits = ATTN_SPLIT_TARGET_BLOCKS.div_ceil(total_rows)   // 512 blocks ≈ one wave on 132 SMs
num_splits    = min(target_splits, cap.div_ceil(ATTN_SPLIT_MIN_CHUNK/*128*/), ATTN_SPLIT_MAX_SPLITS)
chunk         = cap.div_ceil(num_splits)
```

For V2-Lite's 16 decode rows this lands `num_splits=32` (grid 512) at
deep context and scales down cleanly at shallow context. The pure sizing
arithmetic is factored into `attention_split_geometry` with a unit test.

## Capture-safety (unchanged invariant)

`num_splits` still derives **only from `cap` + row count**, never live
seqlen → eager and capture make identical launch decisions. `chunk`
stays `>= MIN_CHUNK (128)`, so any live context `<= 128` remains on the
**single-split bit-exact fast path** (the golden 24-tok lock sits well
under this). `ONNX_GENAI_ATTN_SPLIT_CHUNK` still pins a fixed chunk
(reproduces pre-adaptive behaviour, for A/B).

## Occupancy A/B — ncu `attention_split`, V2-Lite deep ctx (~2600 tok)

| | grid | waves/SM | warps active | DRAM throughput |
|---|---|---|---|---|
| monolithic (OFF) | 16 | 0.02 | 9.8% | 0.8% |
| fixed-256 (#1340) | 256 | 0.48 | 11% | 34% |
| **adaptive (this PR)** | **512** | **0.97** | **24%** | **48%** |

Adaptive ~doubles the grid, doubles occupancy, and lifts effective DRAM
throughput — filling the previously idle SMs.

## Numerics gates (H200, GPU4, pinned, single-threaded)

| Gate | Result |
|---|---|
| V2-Lite golden **24-tok** lock | ✅ **byte-identical** (single-split
path) |
| V2-Lite **340-tok** long-ctx lock (`eager==capture` + golden prefix) |
✅ pass (deep multi-split engaged, capture-safe) |
| `standard_attention_capture_gpu` / `_fp16_gpu` / `_bf16_gpu` | ✅ pass
|
| lib unit tests (`--lib standard_attention`) | ✅ 15/15 (incl. new
`adaptive_split_geometry_*`) |
| `standard_attention_gpu` | ✅ 23/24 — the 1 fail
`..._requires_homogeneous_floating_input_dtypes` is **pre-existing on
clean origin/main**, unrelated |
| Dense qwen2.5-0.5b token md5 | ✅ byte-identical (split does not
engage) |

## Perf A/B (H200 GPU4, eager, medians-of-5)

| context | OFF (monolithic) | fixed-256 (#1340) | **adaptive (this
PR)** | Δ vs fixed-256 |
|---|---|---|---|---|
| **deep** (~2600-tok prompt) | 28.40 | 46.37 | **48.32** | **+4.2%** |
| **shallow** (~520-tok prompt) | 51.69 | 58.08 | **60.82** | **+4.7%**
|
| short "Hello" (capture) | 127.77 | — | 130.17 | neutral (single-split)
|

## Verdict: **GO**, default-ON (adaptive)

Adaptive beats the fixed-256 default at both depths (+4-5%), is neutral
at short context, capture-safe, byte-identical short / f64-tol wide, no
dense regression. It also documents that split-KV's true deep-context
win is **+70% vs monolithic**, far larger than the +10% shallow number
in #1340.

Do **not** self-merge — reporting to coordinator for admin-merge after
review.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
Two independent failures on current `main` (c55a3fa), in files no open
PR touches. Both landed while the Actions queue was backed up, so
nothing caught them at merge time — and between them they make *"is this
PR's CI green?"* unanswerable for every open PR, because `Rust quality`
gates on both and the fmt step also runs inside `Fast`, `CUDA compile`,
`CLI ORT` and the three coverage lanes.

## 1. `cargo fmt --all --check` fails at `standard_attention.rs:2125`

The split-KV workspace sizing added in #1340 is unformatted. Whitespace
only, and it is rustfmt's own output on **rustc 1.97.1** — the exact
toolchain CI resolves `stable` to (`stable-x86_64-unknown-linux-gnu
unchanged - rustc 1.97.1 (8bab26f4f 2026-07-14)` in the last completed
run).

## 2. `verify_documented_env_vars.py` fails on
`ONNX_GENAI_DECODE_GEMV_PROBE_ROWS`

```
ONNX_GENAI_DECODE_GEMV_PROBE_ROWS: documented in
2026-08-18-multirow-gemv-ceiling-probe.md but no crate reads it.
```

The gate is right that nothing reads it, and the fix is *not* to wire it
up. That document's own method section says the probe was **"reverted
after measuring"** — a throwaway control-arm instrument, deliberately
not shipped. That is precisely the case `KNOWN_UNIMPLEMENTED` exists for
(`ONNX_GENAI_WEIGHT_FOLD`, `ONNX_GENAI_GEMV_KSPLIT` and
`ONNX_GENAI_GEMV_CPASYNC` are the same shape), and the script's own rule
— the allowlist entry must stay accompanied by the caveat in the prose —
is already satisfied by the existing text, so the two cannot drift apart
silently.

## Verification

Every step of the `Rust quality` lane, run locally on this branch:

| step | result |
|---|---|
| `cargo fmt --all -- --check` | clean |
| `check_publish_order.py` | ok |
| `workspace_test_packages.py verify` | 49 tested, 5 denied |
| `check_profile_table.py` | ok |
| `check_platform_naming.py` | ok |
| `check_dispatch_reachability.py` | ok |
| `check_dispatch_manifest.py --self-test` + manifest | ok |
| `check_feature_gate_coverage.py` | ok |
| `verify_documented_env_vars.py` | 106 documented, 13
known-unimplemented, all accounted for |

On `main` at the same commit, steps 1 and 9 fail; every other step
already passed. No behaviour change: one whitespace hunk and one
allowlist entry.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 43.18 µs 81.64 µs +89.1%
🔴 gather/large_f32_threads=1-internal/131072 24.81 µs 43.58 µs +75.7%
🔴 gather/large_f16_threads=1-internal/131072 12.10 µs 20.04 µs +65.6%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 29.78 µs 48.83 µs +64.0%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.29 ms 2.00 ms +55.3%
🔴 matmul/medium_generic_bf16_threads=1/32x512x512 502.91 µs 763.77 µs +51.9%
🔴 gather/large_bf16_threads=1-internal/131072 10.55 µs 16.02 µs +51.8%
🔴 matmul/medium_generic_f16_threads=1/32x512x512 27.45 µs 40.99 µs +49.4%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 29.27 µs 42.19 µs +44.2%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 359.10 µs 516.23 µs +43.8%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 539.46 µs 757.26 µs +40.4%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 43.11 µs 58.82 µs +36.4%
🔴 matmul/small_generic_f16_threads=1/1x256x256 30.99 µs 42.17 µs +36.1%
🔴 gather/medium_f32_threads=1-internal/32768 3.54 µs 4.78 µs +35.1%
🔴 matmul/small_generic_bf16_threads=1/1x256x256 30.93 µs 41.02 µs +32.6%
⚠️ gather/small_bf16_threads=1-internal/4096 457.3 ns 585.7 ns +28.1%
⚠️ matmul/medium_generic_bf16_threads=8/32x512x512 355.71 µs 448.84 µs +26.2%
⚠️ add/large_bf16_threads=1-internal/4194304 1.98 ms 2.45 ms +24.0%
⚠️ gather/medium_f16_threads=1-internal/32768 2.34 µs 2.89 µs +23.6%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 78.95 µs 96.75 µs +22.5%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 73.40 µs 86.80 µs +18.3%
⚠️ matmul/small_generic_f32_threads=1/1x256x256 34.35 µs 40.55 µs +18.0%
⚠️ matmul/medium_generic_f32_threads=1/32x512x512 2.19 ms 2.58 ms +17.7%
⚠️ reduce_mean/large_f32_threads=1-internal/262144 958.06 µs 1.12 ms +16.8%
⚠️ matmul/medium_generic_f32_threads=8/32x512x512 896.67 µs 1.05 ms +16.7%
✅ gather/small_f32_threads=1-internal/4096 697.5 ns 791.7 ns +13.5%
✅ gather/small_f16_threads=1-internal/4096 565.0 ns 630.9 ns +11.7%
✅ gather/medium_bf16_threads=1-internal/32768 2.34 µs 2.59 µs +10.6%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.71 ms 4.06 ms +9.5%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.10 ms 9.87 ms +8.5%
✅ tokenization/encode_tokens_per_second 347.77 µs 372.20 µs +7.0%
✅ matmul/small_generic_f32_threads=8/1x256x256 36.05 µs 37.53 µs +4.1%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 482.94 µs 499.28 µs +3.4%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.87 ms 1.92 ms +2.4%
✅ sampling_latency/top_k_per_token 48.65 µs 49.75 µs +2.3%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.50 µs 16.69 µs +1.1%
✅ sampling_latency/top_p_per_token 352.84 µs 356.35 µs +1.0%
✅ sampling_latency/min_p_per_token 193.25 µs 194.84 µs +0.8%
✅ sampling_latency/greedy_per_token 3.05 µs 3.07 µs +0.8%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.29 ms 3.24 ms -1.5%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 1.99 ms 1.96 ms -1.6%
✅ add/medium_f16_threads=1-internal/262144 127.69 µs 125.49 µs -1.7%
✅ add/medium_f32_threads=1-internal/262144 26.15 µs 25.70 µs -1.7%
✅ qwen3_sampling_processors/top_k_partial_selection 133.16 µs 130.34 µs -2.1%
✅ tokenization/decode_tokens_per_second 5.94 ms 5.70 ms -4.1%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.81 ms 5.48 ms -5.7%
✅ kv_cache/alloc_dealloc_pages 40.75 µs 38.32 µs -6.0%
✅ reduce_mean/medium_f32_threads=1-internal/65536 281.44 µs 258.14 µs -8.3%
✅ logit_processing/seven_processor_chain_per_step 333.32 µs 296.82 µs -11.0%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 55.54 µs 49.44 µs -11.0%
✅ add/small_bf16_threads=1-internal/1024 490.9 ns 420.6 ns -14.3%
✅ matmul/small_generic_f16_threads=8/1x256x256 42.68 µs 36.55 µs -14.4%
✅ add/large_f16_threads=1-internal/4194304 2.50 ms 2.14 ms -14.4%
🟢 qwen3_sampling_processors/top_k_top_p_fast 727.25 µs 608.95 µs -16.3%
🟢 grammar_masking/llguidance_compute_mask/32 82.67 µs 68.79 µs -16.8%
🟢 add/large_f32_threads=1-internal/4194304 926.13 µs 738.11 µs -20.3%
🟢 add/medium_bf16_threads=1-internal/262144 137.57 µs 104.68 µs -23.9%
🟢 add/small_f32_threads=1-internal/1024 292.6 ns 196.8 ns -32.7%
🟢 add/small_f16_threads=1-internal/1024 666.8 ns 437.1 ns -34.5%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.09 3.45 7.46 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.80%. Comparing base (06c62e0) to head (975123c).
⚠️ Report is 76 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1340      +/-   ##
==========================================
- Coverage   80.88%   80.80%   -0.09%     
==========================================
  Files         364      362       -2     
  Lines      160729   157677    -3052     
  Branches   160729   157677    -3052     
==========================================
- Hits       130005   127405    -2600     
+ Misses      26069    25632     -437     
+ Partials     4655     4640      -15     
Flag Coverage Δ
mlas ?
offline 80.80% <ø> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

justinchuby added a commit to justinchuby/onnxruntime that referenced this pull request Aug 24, 2026
## Summary

Plan contiguous GQA FlashDecode split-KV launches from fixed KV-cache capacity during CUDA graph capture and replay, while retaining live-sequence-length planning for ordinary eager execution.

## Why

CUDA graphs freeze launch geometry and workspace addresses at capture time. Planning NumSplits from the current live sequence length can become stale as the cache grows, preventing adaptive split-KV behavior from remaining valid across replay. Graph-enabled warmup now reserves capacity-sized workspace, capture uses the fixed-capacity plan, and eager decode avoids redundant capacity heuristic work.

## Behavior

- Capture/replay uses fixed cache capacity for stable NumSplits and workspace sizing
- Graph warmup reserves replay-sized workspace before capture
- Eager execution continues to tune from the live sequence length
- Active memset size remains limited to the launch plan
- Debug output reports the resolved NumSplits

## Validation

Focused host tests cover head sizes 64, 128, and 256; local-window and sequence-tail behavior; non-decode inputs; capture planning; and ordinary eager routing. For an SM108 configuration with live length 129 and capacity 4097, capture selected 17/17/22 splits for head sizes 64/128/256, while eager retained live-length plans. Independent review found one redundant eager heuristic computation, which is fixed in this commit.

No CUDA kernel timing is claimed because this Windows host did not have nvcc. The change preserves eager routing and targets graph planning correctness and replay-stable adaptive split-KV behavior.

Based on the mechanisms validated in justinchuby/onnx-genai#1340 and microsoft#1350.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant