Skip to content

feat(kv): token-major/LATENT PagedAttention index emission (slice 3A.1) - #1955

Merged
justinchuby merged 3 commits into
mainfrom
squad/paged-attention-3a-cuda-latent
Aug 24, 2026
Merged

justinchuby merged 3 commits into
mainfrom
squad/paged-attention-3a-cuda-latent

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Slice 3A.1 — KV-authority half of native PagedAttention integration

First independently-reviewable increment of slice 3A. Follows #1940 (merged
corrected audit + typed validator/oracle crate onnx-genai-paged-attention).

Slice 3A as briefed (native CUDA com.microsoft::PagedAttention LATENT kernel +
KV APIs + full parity/capture matrix + A100 measurement) is larger than one
safe, verifiable increment. Per the "land in order / no dead code / each slice
independently reviewed" directive, this PR lands Requirement 1 only: additive
token-major / LATENT paged-index emission in onnx-genai-kv. It is a strict
prerequisite for the CUDA kernel and is fully validated on host (no GPU).

What it adds (strictly additive, read-only)

crates/onnx-genai-kv/src/paged_index.rs:

  • PagedIndexPlan::build(&PageTable, &[PagedRequest]) emits the exact ORT v1
    int32 index tensors from the existing page authority:
    • block_table[num_seqs, max_num_blocks_per_seq] — physical PageId per
      logical block, padded with PAGED_BLOCK_TABLE_PAD;
    • slot_mapping[token_count] — page_id*block_size + offset per query token;
      PAGED_SLOT_EMPTY == -1 documented skip sentinel;
    • cumulative_sequence_length[num_seqs+1] (cu_seqlens_q), past_seqlens,
      derived context_lens.
  • LatentCacheGeometry + validate(), and canonical token_major_element_offset
    / latent_element_offset so a CUDA kernel and the CPU oracle index the cache
    through one formula. block_size = power-of-two ≥ 16 (matches check_kv_cache).
  • PagedKvCache::emit_paged_index_plan convenience.

One-authority invariant (enforced in code)

onnx-genai-kv stays the sole owner of page allocation/lifetime. paged_index
is a read-only view — it allocates/frees/mutates nothing, no second manager, no
op-side allocation. The physical block id is the PageId
(KvViewKind::VirtuallyContiguous); a kernel binds these host-emitted indices as
stable device inputs and updates caller-owned cache tensors in place.

Byte-identity

No change to Page/PageTable storage; existing head-major layout untouched. A
read-only/leak-free test asserts materialize_sequence and pool usage()/
stats() are identical before and after emission.

Typed rejections (never silent miscompute) — all tested

non-pow2/<16 block size, windowed/attention-sink (non-contiguous) sequences,
query>context, missing backing pages, i32 block-id/slot overflow.

Tests

3 unit + 12 integration (prefill/decode slot math, exact pow2 boundaries
16/32/64, multi-request row-major+padding, page reuse after free, read-only/
leak-free, every typed rejection). Full crate suite (147+…) and
cargo clippy -p onnx-genai-kv --tests clean.

"No dead code" note

First production consumer is the native CUDA kernel (slice 3A.2). This mirrors
the crate's existing kv_capacity_bucket/ensure_kv_capacity/
KvCapacityGrowthBackend authority seam (exercised by tests, consumed by
backends). Its consumer lands in the next, independently-reviewed slice.

Remaining gates

  • 3A.2 native CUDA LATENT kernel + claim wiring (design in the decision note);
    quantized cache → typed NotImplemented; then A100 measurement (n≥3).
  • 3B Mobius --paged-attention export, only after 3A approval.

Reviewer: independent, excluding Leon (author). Final approval: Gaff or Roy.
Do not merge without explicit approval. Draft.

justinchuby and others added 3 commits August 24, 2026 06:38
… 3A.1)

Add an additive, read-only `paged_index` view over the existing page
authority so `onnx-genai-kv` can emit the exact `com.microsoft::PagedAttention`
v1 index tensors (block_table, slot_mapping, cumulative_sequence_length,
past_seqlens) without a second manager or any op-side allocation. This is the
KV-authority half of the one-authority integration and a prerequisite for the
native CUDA LATENT kernel (slice 3A.2); it is fully host-validated.

- `PagedIndexPlan::build(&PageTable, &[PagedRequest])` derives int32 index
  tensors from ordered page ids + lengths; slot = page_id*block_size + offset,
  slot -1 is the documented skip sentinel, block_table padded with a sentinel.
- `LatentCacheGeometry` + `validate()` and canonical `token_major_element_offset`
  / `latent_element_offset` addressing so a kernel and the CPU oracle index the
  token-major/LATENT cache identically. block_size = power-of-two >= 16.
- `PagedKvCache::emit_paged_index_plan` convenience.
- Existing head-major page storage is untouched (byte-identical); a
  read-only/leak-free test asserts materialization + pool usage/stats are
  unchanged across emission.
- Typed rejections (distinct KvError variants, all tested) for non-pow2/<16
  block size, windowed/attention-sink sequences, query>context, missing pages,
  and i32 block-id/slot overflow.

Tests: 3 unit + 12 integration (prefill/decode slot math, exact pow2 block
boundaries 16/32/64, multi-request row-major+padding, page reuse after free,
read-only/leak-free, and every typed rejection). Full crate suite + clippy clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address the independent review's non-blocking observation: the op's
rotary-cache check (`check_rotary_caches`) requires `cos_cache.dims[1] % 8 == 0`
and `rotary_dim = dims[1] * 2`, so a non-zero rotary suffix must be a multiple
of 16. `LatentCacheGeometry::validate` now rejects an even-but-not-16-aligned
rotary_dim with a typed error, so an invalid geometry is refused at the KV seam
instead of failing downstream in the op. rotary_dim == 0 (no rotary) stays
valid. Adds coverage for accepted (16, 32, 0) and rejected (8, 63) widths.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.03690% with 27 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.40%. Comparing base (d3bf2eb) to head (32abb66).
⚠️ Report is 17 commits behind head on main.

Files with missing lines Patch % Lines
crates/onnx-genai-kv/src/paged_index.rs 89.81% 16 Missing and 11 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1955      +/-   ##
==========================================
+ Coverage   80.19%   80.40%   +0.20%     
==========================================
  Files         400      423      +23     
  Lines      186547   207353   +20806     
  Branches   186547   207353   +20806     
==========================================
+ Hits       149606   166719   +17113     
- Misses      31567    34995    +3428     
- Partials     5374     5639     +265     
Flag Coverage Δ
cli-ort-linux 72.51% <ø> (?)
cli-ort-windows 72.10% <ø> (?)
mlas 85.20% <ø> (?)
offline 80.53% <90.03%> (+0.33%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-genai-kv/src/lib.rs 95.48% <ø> (ø)
crates/onnx-genai-kv/src/paged_cache.rs 91.38% <100.00%> (+0.02%) ⬆️
crates/onnx-genai-kv/src/paged_index.rs 89.81% <89.81%> (ø)

... and 70 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@justinchuby

Copy link
Copy Markdown
Owner Author

Independent review (reviewer ≠ Leon) — APPROVE 3A.1 FOR MERGE

Reviewed at HEAD 32abb66e4 vs merge-base 387f840b0 (3 commits, +995/−0, 5 files: paged_index.rs new, paged_index.rs tests new, lib.rs/paged_cache.rs additive, one .squad/ decision record). Read-only; host-only slice — no GPU work reviewed.

Gates (run independently)

  • cargo test -p onnx-genai-kv: all pass — 3 unit + 12 integration paged_index tests, plus the full crate suite (page-table, telemetry, send/sync, doc-tests).
  • cargo clippy -p onnx-genai-kv --all-targets -- -D warnings: clean.
  • cargo fmt -p onnx-genai-kv -- --check: clean.

Correctness — index tensors vs ORT v1.29.0 com.microsoft::PagedAttention

Verified each emitted tensor against the tagged bert_defs.cc schema and paged_attention_helper.h:

  • block_table[num_seqs, max_num_blocks_per_seq] i32, physical PageId per logical block, sentinel-padded — matches input(7) + CheckBlockTable.
  • slot_mapping[token_count] i32, slot = page_id*block_size + offset_in_block, -1 = skip — exactly input(10)'s documented (block_id*block_size + offset_in_block) / -1 semantics.
  • cumulative_sequence_length[num_seqs+1] (cu_seqlens_q) — matches input(5) + CheckSequenceLengthTensors.
  • past_seqlens[num_seqs] = context_len − query_len — matches input(6) "past lengths of cached sequence in the KV cache" (past-only, not total context). This is the subtle one and it is right.
  • block_size power-of-two ≥16 — matches check_kv_cache (block_size < 16 || (block_size & (block_size-1))).

One-authority / isolation / read-only

  • PagedIndexPlan::build(&PageTable, …) and emit_paged_index_plan(&self, …) take shared refs — cannot allocate/free/mutate; no second manager introduced.
  • "physical block id == PageId" holds: the pool pre-populates ids 0..num_gpu_pages and reallocates from the free list, so ids stay dense within [0, num_blocks) across free/reuse — the reused_pages_after_free test confirms valid slots after reuse.
  • Multi-request rows/slots use each sequence's own page list; emission_is_read_only_and_leak_free asserts usage(), stats(), and byte-identical materialize_sequence() before/after 8× emission (head-major layout untouched).
  • i32 overflow guarded for both block-id and slot (try_from → typed PagedBlockIdOverflow/PagedSlotOverflow).
  • Windowed/attention-sink sequences (start≠0/sink≠0) are correctly rejected (PagedNonContiguousSequence) since token-major slots assume contiguous positions — appropriate for this subset.
  • LatentCacheGeometry offset math (token_major_element_offset / latent_element_offset, kv_num_heads==1) is the single shared addressing formula; GLM-5.2 dims (latent 192 / v 128 / rotary 64 @ offset 128) validate.
  • rotary_dim multiple-of-16 review fix (commit 69e2512) is correct and covered (accepts 16/32/0, rejects 8/63), mirroring rotary_dim = cos_cache.dims[1]*2, dims[1] % 8 == 0.

Scope / rebase / no-overclaim

  • All changes confined to crates/onnx-genai-kv/ + one .squad/ record. No dispatch_manifest.toml, no kernels/, no CUDA/native code, no PagedAttention dispatch wiring. The decision record is explicit: "NOT merged," CUDA kernel "not yet implemented" — no premature production claims.
  • Cleanly rebaseable: none of the 10 main commits since the merge-base touch crates/onnx-genai-kv/.

Non-blocking observation (not a merge condition)

LatentCacheGeometry::validate() adopts the stated principle "reject a geometry the op would reject downstream," and correctly added rotary_dim % 16. For full parity with paged_attention_helper.h it could also mirror the two sibling constraints the op enforces at the same site: rotary_offset % 8 == 0 (helper line ~550) and head_size % 8 == 0 → latent_dim % 8 == 0 (helper lines ~27/96/129). Non-blocking because validate() is not on the index-emission path (build() never calls it), and all real MLA configs (DeepSeek kv_lora_rank=512/head_size=576, GLM latent 192/offset 128) already satisfy %8; the op also rejects a bad geometry downstream. Worth a one-line follow-up for contract completeness.

Recommendation

APPROVE. This host-only KV-authority half of the one-authority integration is a clean, additive, fail-closed, exactly-ORT-conformant slice, worth landing as an independently-reviewed increment before the CUDA LATENT kernel (3A.2) and Mobius export (3B). Note the PR is still a draft — mark ready before merge; per the decision record, Gaff/Roy final sign-off applies.

@justinchuby
justinchuby marked this pull request as ready for review August 24, 2026 08:18
@justinchuby
justinchuby merged commit 011fbb2 into main Aug 24, 2026
11 of 17 checks passed
@justinchuby
justinchuby deleted the squad/paged-attention-3a-cuda-latent branch August 24, 2026 08:18
justinchuby added a commit that referenced this pull request Aug 24, 2026
…MLA (slice 3A.2) (#1978)

## Slice 3A.2 — native onnx-genai CUDA support for
`com.microsoft::PagedAttention` LATENT (GLM-5.2 dense MLA)

Builds on the merged audit/validator/oracle (#1940) and the
token-major/LATENT KV index emission (#1955). This slice adds the
**native CUDA kernel** for the exact ORT v1
`com.microsoft::PagedAttention` **LATENT** (absorbed-MLA) subset that
GLM-5.2 dense MLA needs, with `onnx-genai-kv` remaining the **sole**
page/cache authority.

**Depends on / stacks after #1955** (already merged into `main`).

### What this adds
- `crates/onnx-runtime-ep-cuda/src/kernels/paged_attention.rs` — NVRTC
f16/bf16 write + attention kernels implemented **from the oracle
equations** (not copied upstream source): partial-RoPE suffix write into
the paged latent cache, online softmax over the latent cache honoring
`local_window_size`/softcap, V taken from the leading `v_head_size`
channels of the same latent row. Plus `PagedAttentionFactory`,
`PagedAttentionLatentKernel`, and `unsupported_reason()`.
- Five-place op registration in `kernels/mod.rs` + the
`unsupported_reason` arm in `provider.rs::supports_op`.

### Invariants held
- **Default-off / typed subset only.** The op is claimed **only** when
the typed geometry+dtype validator proves the exact supported subset
(fp16/bf16 LATENT, single latent KV head, GLM qk=192/v=128/partial RoPE,
block pow2 ≥ 16). Every unsupported optional mode returns a **typed
NotImplemented** reason rather than silently miscomputing: non-LATENT
layout, quantized cache (k/v quant type + int4/float4 cache dtype),
`head_sink`, q/k-norm, k/v scales, present `value`/`value_cache`,
non-f16/bf16.
- **One-authority.** No op-side allocation and no second KV manager —
the kernel consumes/mutates the caller's page buffers in place.
**In-place `key_cache_out` alias contract enforced** (non-aliased cache
output is rejected).
- **Capture/replay safe.** Warmed kernel signature; no host
sync/allocation during capture. Proven by a capture + 3×replay test that
is bit-equal to eager.
- **No Mobius changes / no export claims** in this slice. No full-size
performance claim.

### Tests (native CUDA vs `onnx-genai-paged-attention` oracle, verified
on an idle A100)
- Parity: tiny prefill/decode; GLM dims (qk=192/v=128, partial RoPE)
first-token/prefill/decode; block pow2≥16 boundary + multi-request; slot
`-1` skip; no-rotary; eager equivalence (fp16 err ~2e-4, bf16 ~2e-3 —
well under tol).
- Capture + 3 replays bit-equal to eager (`check_capture_error()==0`).
- Rejections: non-aliased `key_cache_out` and missing required input
(`block_table`).
- Measurement (tiny-shape correctness gate, **not** a perf claim):
CUDA-event timing, n=5, prefill + decode, page/VRAM accounting, op-side
alloc = 0. Prefill med ~0.050 ms, decode med ~0.073 ms on tiny GLM
shapes.
- 5 pure unit tests cover the typed rejections without a GPU.
- Existing `onnx-runtime-ep-cuda` lib suite green (559 passed) — no
regressions; KV geometry validation (#1955 + Gaff's sibling constraints)
unchanged.

### Reviewer / gates
- **Draft** — for independent review by **Gaff or Roy** (final approval
required; reviewer excludes Leon/Sapper). Do **not** merge without
explicit approval.
- Remaining gates: full-size GLM checkpoint runs (correctness +
measurement) coordinated with the ongoing GLM GGUF/safetensors work; 3B
Mobius opt-in `--paged-attention` export only starts **after** 3A
approval.

_GPU tests are gated behind the `gpu-tests` cargo feature and the CUDA
env; run on A100 with `--features gpu-tests`._

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant