Repository navigation
perf(cuda): split-KV FlashDecoding for attention_row decode (+10% V2-Lite wide-ctx) - #1340
Merged
Merged
Conversation
…Lite wide-ctx) attention_row launches one block per (batch, q_head, query) row, so decode uses only ~16 of 132 SMs and is memory-latency bound (~36% of GPU decode at wide context; ncu: grid=16, waves/SM=0.02, DRAM 0.79%, occupancy 9.8%). Add a two-kernel FlashDecoding split-KV path behind ONNX_GENAI_ATTN_SPLITKV (default ON): - attention_split: grid (total_rows, num_splits). Each block reduces one KV slice with chunk-local online softmax (running max + rescale), writing an unnormalized partial P.V plus per-split max/sum meta. A single-split fast path (total_seq <= chunk) reproduces attention_row exactly and writes the final normalized output with sentinel meta (max=0, sum=1). - attention_combine: grid (total_rows). Uniform log-sum-exp merge of the per -split partials. The single-split sentinel makes it a bit-exact pass-through, so contexts <= chunk stay byte-identical to the monolithic kernel. Capture-safe: num_splits and the engage decision derive only from the fixed KV capacity (cap), never live seqlen, so eager and capture make identical launch decisions. Split scratch is drawn from the internal capture-warmed ws pool (WS_SPLIT slot), leaving the governed workspace layout untouched. Multi-split reorders fp32 partials so wide context is validated to the #1150 f64 tolerance rather than byte-identity; greedy token output is unchanged. Gates (H200, GPU3, pinned single-threaded): - V2-Lite golden 24-tok lock: byte-identical (single-split path). - V2-Lite 340-tok long-context lock: eager==capture holds with multi-split engaged; golden prefix matches. - standard_attention capture/fp16/bf16 gpu-tests + layout unit tests pass. - Dense qwen2.5-0.5b: byte-identical, tok/s neutral. - Wide-ctx kernel: attention_row 147.6us -> attention_split+combine ~65us (~2.3x); DRAM throughput 0.79% -> 9.26%. - E2E wide-ctx: 51.64 -> 56.90 tok/s (+10.2%); short-ctx neutral. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 18, 2026
…log (PR #1337, #1340) (#1343) ## Summary Scribe consolidation of two Deckard decision inbox notes into the canonical `.squad/decisions.md` ledger. ### Decisions logged **PR #1337 — attention_row block-width 128→256 (GO, byte-identical, merged 37bdefe)** - Profile evidence: grid=16, achieved occupancy 6.23%, waves/SM=0.01, 62.7% barrier-stall cycles — memory-latency bound, not roofline-limited. - attention_row context scaling: 12.4% GPU decode share at short ctx → 40% at deep ctx → **#1 decode kernel at realistic/wide context** (QMoE levers exhausted). - Numerics gate: BYTE-IDENTICAL (same accumulation order, more warps). - H200 A/B: −14% attention_row kernel latency, **+5.5% E2E** on V2-Lite ~500-tok prompt; neutral on short/dense. **PR #1340 — split-KV FlashDecoding for attention_row (GO, default-ON, merged 763d81f)** - Profile evidence: grid=16, waves/SM=0.02, ~116 idle SMs, DRAM 0.79%, occupancy 9.8% — machine starved. - Single-split sentinel fast path: contexts ≤ chunk (default 256) → **byte-identical pass-through**. - Wide-ctx (multi-split): reordered fp32 partials → **f64-tol** (#1150 oracle); capture-safe (num_splits derived from fixed KV cap, not live seqlen). - H200 A/B: ~2.3× kernel speedup, DRAM 0.79%→9.26%, **+10.2% E2E** wide ctx; neutral short/dense. ### Archive gate All entries in decisions.md are dated 2026-08-18 (active campaign arc) — nothing >30 days old. No archiving performed per the size-not-age rule. ### File scope - Modified: `.squad/decisions.md` (+86 lines) - Deleted (untracked, rm only): `.squad/decisions/inbox/deckard-attention-row.md`, `.squad/decisions/inbox/deckard-attention-splitkv.md` Do NOT self-merge — awaiting coordinator review. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This was referenced Aug 18, 2026
justinchuby
added a commit
that referenced
this pull request
Aug 18, 2026
Two independent failures on `main` (c55a3fa), both in files the failing PRs do not touch, both landed while the Actions queue was backed up so nothing caught them at merge time. Between them they make "is this PR's CI green?" unanswerable for every open PR. 1. `cargo fmt --all --check` fails at `standard_attention.rs:2125` — the split-KV workspace sizing from #1340 is unformatted. Whitespace only, and it is rustfmt's own output on rustc 1.97.1, the toolchain CI resolves `stable` to. 2. `verify_documented_env_vars.py` fails on `ONNX_GENAI_DECODE_GEMV_PROBE_ROWS`, documented by `2026-08-18-multirow-gemv-ceiling-probe.md`. That doc's method section already says the probe was "reverted after measuring" — it is deliberately not shipped, which is exactly what `KNOWN_UNIMPLEMENTED` is for, so it goes there with that reason rather than being wired up. After this commit every step of the `Rust quality` lane passes locally: `cargo fmt --all --check`, `check_publish_order`, `workspace_test_packages verify` (49 tested, 5 denied), `check_profile_table`, `check_platform_naming`, `check_dispatch_reachability`, `check_dispatch_manifest` (+ self-test), `check_feature_gate_coverage`, and `verify_documented_env_vars` (106 documented, 13 known-unimplemented, all accounted for). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 18, 2026
…over fixed-256; +70% deep-ctx) (#1350) ## Summary Follow-up to #1340 (split-KV FlashDecoding for `attention_row`). That PR fixed keys-per-split at `chunk=256`, so `num_splits = cap/256`. Re-profiling at **deep context** (~2600-tok prompt, the regime where `attention_row` reaches ~40% of decode) showed two things: 1. **The split-KV win is far bigger at depth than the shallow A/B in #1340 measured:** deep-ctx E2E is **28.40 → 46.37 tok/s (+63%)** with fixed-256, vs the +10% originally reported at ~500-ctx. 2. **Fixed-256 under-fills the machine at depth.** At `cap≈4096` the grid is `(16 rows × 16 splits) = 256` blocks, waves/SM 0.48, warps active 11%. The decode split kernel is memory-latency bound, so it wants **~a full occupancy wave of resident CTAs, not one block per SM** — an ncu + E2E chunk sweep found the optimum at `chunk=128` (grid 512), not `chunk=512` (grid 132, which is ~5% slower). This PR sizes `num_splits` **adaptively** from the fixed `cap` **and** the fixed row count to target ~one occupancy wave, floored so no split is starved of work: ``` target_splits = ATTN_SPLIT_TARGET_BLOCKS.div_ceil(total_rows) // 512 blocks ≈ one wave on 132 SMs num_splits = min(target_splits, cap.div_ceil(ATTN_SPLIT_MIN_CHUNK/*128*/), ATTN_SPLIT_MAX_SPLITS) chunk = cap.div_ceil(num_splits) ``` For V2-Lite's 16 decode rows this lands `num_splits=32` (grid 512) at deep context and scales down cleanly at shallow context. The pure sizing arithmetic is factored into `attention_split_geometry` with a unit test. ## Capture-safety (unchanged invariant) `num_splits` still derives **only from `cap` + row count**, never live seqlen → eager and capture make identical launch decisions. `chunk` stays `>= MIN_CHUNK (128)`, so any live context `<= 128` remains on the **single-split bit-exact fast path** (the golden 24-tok lock sits well under this). `ONNX_GENAI_ATTN_SPLIT_CHUNK` still pins a fixed chunk (reproduces pre-adaptive behaviour, for A/B). ## Occupancy A/B — ncu `attention_split`, V2-Lite deep ctx (~2600 tok) | | grid | waves/SM | warps active | DRAM throughput | |---|---|---|---|---| | monolithic (OFF) | 16 | 0.02 | 9.8% | 0.8% | | fixed-256 (#1340) | 256 | 0.48 | 11% | 34% | | **adaptive (this PR)** | **512** | **0.97** | **24%** | **48%** | Adaptive ~doubles the grid, doubles occupancy, and lifts effective DRAM throughput — filling the previously idle SMs. ## Numerics gates (H200, GPU4, pinned, single-threaded) | Gate | Result | |---|---| | V2-Lite golden **24-tok** lock | ✅ **byte-identical** (single-split path) | | V2-Lite **340-tok** long-ctx lock (`eager==capture` + golden prefix) | ✅ pass (deep multi-split engaged, capture-safe) | | `standard_attention_capture_gpu` / `_fp16_gpu` / `_bf16_gpu` | ✅ pass | | lib unit tests (`--lib standard_attention`) | ✅ 15/15 (incl. new `adaptive_split_geometry_*`) | | `standard_attention_gpu` | ✅ 23/24 — the 1 fail `..._requires_homogeneous_floating_input_dtypes` is **pre-existing on clean origin/main**, unrelated | | Dense qwen2.5-0.5b token md5 | ✅ byte-identical (split does not engage) | ## Perf A/B (H200 GPU4, eager, medians-of-5) | context | OFF (monolithic) | fixed-256 (#1340) | **adaptive (this PR)** | Δ vs fixed-256 | |---|---|---|---|---| | **deep** (~2600-tok prompt) | 28.40 | 46.37 | **48.32** | **+4.2%** | | **shallow** (~520-tok prompt) | 51.69 | 58.08 | **60.82** | **+4.7%** | | short "Hello" (capture) | 127.77 | — | 130.17 | neutral (single-split) | ## Verdict: **GO**, default-ON (adaptive) Adaptive beats the fixed-256 default at both depths (+4-5%), is neutral at short context, capture-safe, byte-identical short / f64-tol wide, no dense regression. It also documents that split-KV's true deep-context win is **+70% vs monolithic**, far larger than the +10% shallow number in #1340. Do **not** self-merge — reporting to coordinator for admin-merge after review. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
Two independent failures on current `main` (c55a3fa), in files no open PR touches. Both landed while the Actions queue was backed up, so nothing caught them at merge time — and between them they make *"is this PR's CI green?"* unanswerable for every open PR, because `Rust quality` gates on both and the fmt step also runs inside `Fast`, `CUDA compile`, `CLI ORT` and the three coverage lanes. ## 1. `cargo fmt --all --check` fails at `standard_attention.rs:2125` The split-KV workspace sizing added in #1340 is unformatted. Whitespace only, and it is rustfmt's own output on **rustc 1.97.1** — the exact toolchain CI resolves `stable` to (`stable-x86_64-unknown-linux-gnu unchanged - rustc 1.97.1 (8bab26f4f 2026-07-14)` in the last completed run). ## 2. `verify_documented_env_vars.py` fails on `ONNX_GENAI_DECODE_GEMV_PROBE_ROWS` ``` ONNX_GENAI_DECODE_GEMV_PROBE_ROWS: documented in 2026-08-18-multirow-gemv-ceiling-probe.md but no crate reads it. ``` The gate is right that nothing reads it, and the fix is *not* to wire it up. That document's own method section says the probe was **"reverted after measuring"** — a throwaway control-arm instrument, deliberately not shipped. That is precisely the case `KNOWN_UNIMPLEMENTED` exists for (`ONNX_GENAI_WEIGHT_FOLD`, `ONNX_GENAI_GEMV_KSPLIT` and `ONNX_GENAI_GEMV_CPASYNC` are the same shape), and the script's own rule — the allowlist entry must stay accompanied by the caveat in the prose — is already satisfied by the existing text, so the two cannot drift apart silently. ## Verification Every step of the `Rust quality` lane, run locally on this branch: | step | result | |---|---| | `cargo fmt --all -- --check` | clean | | `check_publish_order.py` | ok | | `workspace_test_packages.py verify` | 49 tested, 5 denied | | `check_profile_table.py` | ok | | `check_platform_naming.py` | ok | | `check_dispatch_reachability.py` | ok | | `check_dispatch_manifest.py --self-test` + manifest | ok | | `check_feature_gate_coverage.py` | ok | | `verify_documented_env_vars.py` | 106 documented, 13 known-unimplemented, all accounted for | On `main` at the same commit, steps 1 and 9 fail; every other step already passed. No behaviour change: one whitespace hunk and one allowlist entry. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1340 +/- ##
==========================================
- Coverage 80.88% 80.80% -0.09%
==========================================
Files 364 362 -2
Lines 160729 157677 -3052
Branches 160729 157677 -3052
==========================================
- Hits 130005 127405 -2600
+ Misses 26069 25632 -437
+ Partials 4655 4640 -15
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
justinchuby
added a commit
to justinchuby/onnxruntime
that referenced
this pull request
Aug 24, 2026
## Summary Plan contiguous GQA FlashDecode split-KV launches from fixed KV-cache capacity during CUDA graph capture and replay, while retaining live-sequence-length planning for ordinary eager execution. ## Why CUDA graphs freeze launch geometry and workspace addresses at capture time. Planning NumSplits from the current live sequence length can become stale as the cache grows, preventing adaptive split-KV behavior from remaining valid across replay. Graph-enabled warmup now reserves capacity-sized workspace, capture uses the fixed-capacity plan, and eager decode avoids redundant capacity heuristic work. ## Behavior - Capture/replay uses fixed cache capacity for stable NumSplits and workspace sizing - Graph warmup reserves replay-sized workspace before capture - Eager execution continues to tune from the live sequence length - Active memset size remains limited to the launch plan - Debug output reports the resolved NumSplits ## Validation Focused host tests cover head sizes 64, 128, and 256; local-window and sequence-tail behavior; non-decode inputs; capture planning; and ordinary eager routing. For an SM108 configuration with live length 129 and capacity 4097, capture selected 17/17/22 splits for head sizes 64/128/256, while eager retained live-length plans. Independent review found one redundant eager heuristic computation, which is fixed in this commit. No CUDA kernel timing is claimed because this Windows host did not have nvcc. The change preserves eager routing and targets graph planning correctness and replay-stable adaptive split-KV behavior. Based on the mechanisms validated in justinchuby/onnx-genai#1340 and microsoft#1350. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
attention_rowlaunches one block per (batch, q_head, query) row, so at decode it uses only ~16 of 132 SMs and is memory-latency bound. Profiling V2-Lite eager decode at wide context (GPU3, ~390-word prompt, 256 tok) showedattention_rowis the #1 decode kernel at 36% of GPU decode — ncu: grid=16, waves/SM=0.02 (~116 idle SMs), DRAM 0.79%, achieved occupancy 9.8% → a definitive under-utilization signal.This PR adds a two-kernel FlashDecoding split-KV path that spreads each row's key reduction across many blocks to fill the idle SMs.
Design
Behind
ONNX_GENAI_ATTN_SPLITKV(default ON;=0reproduces the monolithic baseline for A/B):attention_split— grid(total_rows, num_splits). Each block reduces one KV slice with a chunk-local online softmax (running max + rescale), writing an unnormalized partial P·V plus per-splitmax/summeta. A single-split fast path (total_seq <= chunk) reproducesattention_rowexactly and writes the final normalized output with sentinel meta (max=0, sum=1).attention_combine— grid(total_rows). Uniform log-sum-exp merge of the per-split partials. The single-split sentinel makes it a bit-exact pass-through, so any context<= chunk(default 256) stays byte-identical to the monolithic kernel.Capture-safety:
num_splitsand the engage decision derive only from the fixed KV capacity (cap), never live seqlen, so eager and capture make identical launch decisions → theeager==capturelong-context lock holds. Split scratch is drawn from the internal capture-warmed ws pool (WS_SPLIT slot), leaving the governed workspace layout (and its exact-byte unit tests) untouched.Multi-split reorders fp32 partials, so wide context is validated to the #1150 f64 tolerance (
2e-4.max(·*2e-5)), not byte-identity. Greedy token output is unchanged in practice.Numerics gates (H200, GPU3, pinned, single-threaded)
matches_golden)eager==capture+ golden prefix)standard_attention_capture_gpu/_fp16_gpu/_bf16_gpustandard_attention_gpu..._requires_homogeneous_floating_input_dtypesis pre-existing on clean origin/main, unrelated)--lib standard_attention)Perf A/B (H200 GPU3, pinned, medians-of-5)
Wide-context kernel (nsys, same binary, split ON vs
=0):attention_row)attention_split+combine)E2E throughput (tok/s):
=0)num_splits = cap.div_ceil(chunk).clamp(1, 64)withchunkdefault 256 (tunable viaONNX_GENAI_ATTN_SPLIT_CHUNK). At the true long-context regime whereattention_rowreaches ~40% of decode (cap≈2048), this yields ~8 splits → grid ~128, filling the machine; smaller chunks extract a bit more at moderate context at the cost of engaging split earlier.Verdict: GO, default-ON
Short context stays byte-identical (comfortable margin below
chunk); wide context gets +10.2% E2E / ~2.3× kernel; capture-safe; no dense regression. All hard gates green.Do not self-merge — reporting to coordinator for admin-merge after review.
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com