Skip to content

feat(pipeline): native present-KV mirroring for paged native decode (GAP-3 Inc-C) - #566

Merged
justinchuby merged 2 commits into
mainfrom
feat/gap3-inc-c-present-kv
Jul 31, 2026
Merged

justinchuby merged 2 commits into
mainfrom
feat/gap3-inc-c-present-kv

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

GAP-3 Inc-C — native present-KV mirroring → paged native pipeline decode

Closes the S2 bail in NativePipelineDecoder::mirror_last_present_kv so the pure-native pipeline decoder (Inc-A #565) can run paged (cross-request KV prefix reuse), not only the non-paged flat-AR path. Host-KV path; device-resident present-KV (35B-A3B GPU) is explicitly deferred to Inc-D.

What's wired

  • native_decode/mod.rs: supports_host_kv_mirror gate, host_present_kv (read), seed_growable_kv (seed).
  • pipeline/decoder_component.rs: real mirror_last_present_kv — byte-identical geometry to ORT's mirror_present_kv_to_pages (shared extract_present_token/append_token_kv); supports_paged_kv + load_paged_prefix seams.
  • pipeline/flat_autoregressive.rs: DRY-factored claim_paged_prefix shared by ORT + native admit paths (only the KV sink differs).
  • pipeline/mod.rs: read-only #[cfg(feature="native-backend")] test accessor materialize_published_prefix_kv (non-mutating prefix-cache read-back).
  • tests/native_pipeline_backend_selection_parity.rs: native_paged_prefix_reuse_matches_fresh_and_ort — warm==cold==ORT tokens plus direct native-vs-ORT paged-KV byte-equality assertion.

Correctness

Deferred to Inc-D (35B-A3B GPU end-to-end)

Device-resident present-KV read-out, in-place-GQA CPU KV, f16/non-rank-4 round-trip, MoE routed-expert specifics if needed. These decoders correctly gate supports_paged_kv=false → Inc-A non-paged fallback (no silent-wrong paged run).

Reviews

Independent opus review with strict author-lockout: Mary REJECTED the original (vacuous geometry test) → revised the test as authorized non-author → Harry independent re-review APPROVE (all mutations re-run, byte-equality confirmed, production byte-identical, deferral gate verified).

Note: run the parity test binary with --test-threads=1 (two tests share a process-global decoder-device env var; parallel races flake — pre-existing on HEAD).

justinchuby and others added 2 commits July 31, 2026 18:21
Close the S2 bail at pipeline/decoder_component.rs
(NativePipelineDecoder::mirror_last_present_kv) so the pure-native pipeline
path from Inc-A (#565) can run PAGED, not just the non-paged flat-AR path.
This is the present-KV threading increment paged multi-component native
serving needs.

Host-resident growable f32 KV path only (supports_host_kv_mirror gate):
implements the full paged round-trip — present-KV mirror-write + shared-prefix
seed-read — reusing the SAME kv_bridge geometry primitives the ORT decoder
uses (extract_present_token + append_token_kv), so native and ORT mirror
byte-identical pages. Device-resident (CUDA) and in-place-GQA (CPU) present-KV
read-out is Inc-D; those decoders keep the Inc-A non-paged path (no regression).

- native_decode/mod.rs: +supports_host_kv_mirror, host_present_kv,
  seed_growable_kv on NativeDecodeSession.
- decoder_component.rs: real mirror_last_present_kv; trait supports_paged_kv +
  load_paged_prefix seams; native load_paged_prefix seeds the growable cache.
- flat_autoregressive.rs: build native decoder up front, admit it to the paged
  gate, DRY-factor claim_paged_prefix shared by ORT and native admit paths.
- No decode-loop / native_decode core / capture-core changes.

Correctness (token-exact, two-tier): new paged cross-request reuse test on
tiny-gemma4-vlm asserts warm native reuse (prefix_reused_tokens=4>0) ==
cold pure-native == ORT oracle. Non-vacuous: reverting the bail errors,
wrong geometry diverges tokens, silent no-op reuse fails reused>0.

Regressions green: #565, #384, #541, #543, #554; #544 env-gated; 350 lib
unit tests. fmt/clippy clean; no-feature build verified.

35B-A3B: not yet native end-to-end on GPU (device-resident KV → Inc-D).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…l (test-rigor)

The Inc-C parity test native_paged_prefix_reuse_matches_fresh_and_ort was
vacuous for KV geometry: it asserted warm==cold==ORT on argmax tokens, but on
tiny-gemma4-vlm the argmax is invariant to the reused-prefix KV, so a key/value
SWAP and even a fully ZEROED mirrored KV still produced identical tokens and
PASSED. Only the no-op mirror and forced-reused=0 mutations failed.

Add a direct byte/element-equality check on the mirrored paged KV, giving the
test discriminating power over the mirror geometry regardless of argmax
sensitivity:

- pipeline/mod.rs: add a minimal, read-only, native-backend-gated test-support
  accessor PipelineEngine::materialize_published_prefix_kv. It reconstructs the
  published prefix key (digest_request_identity + prefix_key), does a
  non-mutating PrefixCache::lookup, attaches the pages to a throwaway sequence
  to read them, materializes per-layer K/V, then drops the sequence WITHOUT
  freeing the pages (the prefix cache keeps sole ownership). No decode-path
  behavior changes.
- test: after the warm native run, assert the native-mirrored paged KV for the
  shared prefix equals the ORT-mirrored KV byte-for-byte (MaterializedKv:
  PartialEq). Catches key/value swap, zeroing, head-stride, seq-offset and
  page-index errors. Reuse>0 and token asserts retained.

Test-only: native_decode/mod.rs, pipeline/decoder_component.rs and
pipeline/flat_autoregressive.rs are byte-identical to HEAD.

Non-vacuity (each applied, run on GPU 0, reverted): (a) no-op mirror FAILS
(publishes unmirrored pages), (b) key/value swap FAILS on the new byte assert
(tokens still matched), (c) forced reused=0 FAILS (reused>0 assert). Clean run
green (2 passed, --test-threads=1); fmt/clippy clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent opus re-review (Harry) of 8612907f..112d6e06, author-lockout enforced (Cohaagen=Inc-C author, Mary=test-fix author, both locked out of review).

  • Non-vacuity re-verified independently: no-op mirror, key/value swap, zeroed mirror, forced reused=0 — ALL fail. The byte-equality assert catches the geometry corruption that argmax-only missed on tiny-gemma4-vlm.
  • Test accessor materialize_published_prefix_kv confirmed read-only/non-mutating (no refcount/page mutation); production code byte-identical to the Inc-C impl commit.
  • Present-KV geometry identical to ORT (shared extract/append primitives, same layer/head order).
  • Scope clean: zero edits to the parked capture core (plan_capture_region, executor/capture.rs, CudaGraphLifecycle).
  • Deferral gate honest: cuda/cpu-kv/non-f32-rank4 → supports_paged_kv=false → Inc-A non-paged fallback; seed_growable_kv hard-bails.
  • Green (GPU 6, --test-threads=1): parity, multi_session, prefix_speedup, multimodal_reuse_e2e (14), native full/decoder/step parity — no regressions.

Prior REJECT (Mary) on vacuous geometry test was correctly remediated test-only; underlying Inc-C production code was correct-by-construction throughout.

@justinchuby
justinchuby enabled auto-merge (squash) July 31, 2026 19:45
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.55%. Comparing base (8612907) to head (112d6e0).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main     #566   +/-   ##
=======================================
  Coverage   81.55%   81.55%           
=======================================
  Files         315      315           
  Lines      123574   123574           
  Branches   123574   123574           
=======================================
+ Hits       100777   100779    +2     
+ Misses      18739    18738    -1     
+ Partials     4058     4057    -1     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.67% <ø> (ø)
mlas 78.72% <ø> (ø)
offline 81.51% <ø> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@justinchuby
justinchuby merged commit 096dfbc into main Jul 31, 2026
14 checks passed
@justinchuby
justinchuby deleted the feat/gap3-inc-c-present-kv branch July 31, 2026 19:54
@github-actions

Copy link
Copy Markdown

✅ Benchmarks — No Regression

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
✅ sampling_latency/top_p_per_token 1.02 ms 1.06 ms +3.7%
✅ kv_cache/alloc_dealloc_pages 37.63 µs 38.78 µs +3.1%
✅ sampling_latency/top_k_per_token 485.61 µs 500.34 µs +3.0%
✅ sampling_latency/min_p_per_token 348.91 µs 355.83 µs +2.0%
✅ logit_processing/seven_processor_chain_per_step 1.16 ms 1.18 ms +1.5%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.26 ms 2.20 ms -2.3%
✅ tokenization/encode_tokens_per_second 415.71 µs 402.92 µs -3.1%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 81.43 µs 78.45 µs -3.7%
✅ grammar_masking/llguidance_compute_mask/32 81.13 µs 74.95 µs -7.6%
✅ add/small_f32_threads=1-internal/1024 220.8 ns 203.7 ns -7.7%
✅ add/medium_f32_threads=1-internal/262144 2.56 ms 2.35 ms -8.4%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.02 ms 1.84 ms -8.9%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 4.01 ms 3.65 ms -9.0%
✅ tokenization/decode_tokens_per_second 7.37 ms 6.69 ms -9.2%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 530.13 µs 480.82 µs -9.3%
✅ matmul/small_generic_bf16_threads=8/1x256x256 31.51 µs 28.58 µs -9.3%
✅ add/large_bf16_threads=1-internal/4194304 43.45 ms 39.24 ms -9.7%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 88.14 µs 79.01 µs -10.4%
✅ add/medium_bf16_threads=1-internal/262144 2.76 ms 2.44 ms -11.8%
✅ matmul/medium_generic_f16_threads=1/32x512x512 31.30 µs 27.59 µs -11.9%
✅ add/medium_f16_threads=1-internal/262144 2.69 ms 2.36 ms -12.1%
✅ add/small_f16_threads=1-internal/1024 14.79 µs 12.95 µs -12.4%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.76 ms 8.53 ms -12.6%
✅ add/large_f32_threads=1-internal/4194304 43.71 ms 38.06 ms -12.9%
🟢 add/small_bf16_threads=1-internal/1024 14.20 µs 11.95 µs -15.8%
🟢 sampling_latency/greedy_per_token 3.92 µs 3.27 µs -16.6%
🟢 reduce_mean/small_f32_threads=1-internal/4096 16.99 µs 14.13 µs -16.8%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.06 ms 879.72 µs -17.2%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 34.19 µs 28.17 µs -17.6%
🟢 matmul/small_generic_f16_threads=1/1x256x256 33.85 µs 27.82 µs -17.8%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 34.81 µs 28.47 µs -18.2%
🟢 gather/small_f32_threads=1-internal/4096 884.4 ns 663.3 ns -25.0%
🟢 matmul/medium_generic_bf16_threads=8/32x512x512 481.41 µs 359.58 µs -25.3%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.29 ms 951.19 µs -26.4%
🟢 add/large_f16_threads=1-internal/4194304 52.53 ms 38.19 ms -27.3%
🟢 matmul/small_generic_f32_threads=8/1x256x256 49.30 µs 35.22 µs -28.6%
🟢 gather/small_f16_threads=1-internal/4096 619.3 ns 440.9 ns -28.8%
🟢 gather/medium_bf16_threads=1-internal/32768 3.10 µs 2.20 µs -29.0%
🟢 gather/small_bf16_threads=1-internal/4096 652.2 ns 444.5 ns -31.8%
🟢 matmul/small_generic_f32_threads=1/1x256x256 52.77 µs 34.07 µs -35.4%
🟢 matmul/small_generic_f16_threads=8/1x256x256 43.35 µs 27.93 µs -35.6%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 363.01 µs 227.09 µs -37.4%
🟢 gather/medium_f16_threads=1-internal/32768 4.36 µs 2.57 µs -41.0%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.20 ms 1.25 ms -43.5%
🟢 gather/large_f16_threads=1-internal/131072 20.28 µs 11.39 µs -43.8%
🟢 gather/medium_f32_threads=1-internal/32768 7.19 µs 3.72 µs -48.2%
🟢 gather/large_f32_threads=1-internal/131072 53.61 µs 26.62 µs -50.3%
🟢 gather/large_bf16_threads=1-internal/131072 39.62 µs 12.28 µs -69.0%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.51 4.46 7.51 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Jul 31, 2026
…DA decode (GAP-3 Inc-D) (#567)

## GAP-3 Inc-D — device-resident present-KV read-out → paged native CUDA
decode

Lifts the Inc-C (`#566`) `supports_paged_kv=false` gate for
**device-resident f32 rank-4 CUDA GQA present-KV**, so a native CUDA
pipeline decoder now runs **paged** (cross-request KV reuse) instead of
the Inc-A non-paged fallback. Closes the present-KV threading gap for
Qwen3.6-35B-A3B GPU decode.

### How (pure post-step plumbing — no kernel/capture-core edits)
- `native_decode/cuda.rs`: `read_present_kv` reads the KV binding
**after** the decode step's existing stream sync
(`read_bytes`→`copy_to_host`→`dtoh`, which synchronizes) using the
**physical/capacity** shape `[1,H,max_len,Dh]` so strides address the
padded buffer; `seed_prefix` is the device counterpart of Inc-C's host
seed; `device_present_kv_view` isolates the physical-shape handling.
- `native_decode/mod.rs`:
`present_kv`/`seed_kv`/`supports_device_kv_mirror` unify host-growable
(Inc-C) and device-CUDA (Inc-D) onto the **same**
`extract_present_token`/`append_token_kv` geometry + host f32 paged
store (DRY, byte-comparable with ORT).
- `pipeline/decoder_component.rs`: `supports_paged_kv = host OR device`;
`mirror_last_present_kv &self→&mut self` (rippled to trait + ORT impl);
`load_paged_prefix → seed_kv`.

### Correctness
- `native_paged_prefix_reuse_matches_ort_on_cuda_device`:
paged-native-CUDA == non-paged-native-CUDA == ORT oracle == closed-form
tokens; mirrored pages **byte-equal** CUDA-vs-ORT; `reused=4`.
- **H=2 unit test** for the physical-vs-logical stride bug (all existing
fixtures are H=1, where the head-stride error is invisible). Mutating to
the logical stride fails it.
- Non-vacuity (independently re-run by reviewer): gate-revert,
logical-stride, and forced-reused=0 mutations ALL fail.
- Honest gating: f16 / non-rank-4 / CPU-in-place-GQA /
sink-discontinuous stay `supports_paged_kv=false` → Inc-A non-paged
fallback (no silent-wrong paged run).
- No regressions: Inc-C #566, Inc-A #565, #541/#543 hybrid, #554 reuse
(14/14), native_decode lib (54); scope clean (no standard_attention/GQA
kernel, capture core, or `plan_capture_region` edits).

### Remaining (follow-up Inc-D.1)
Real 35B-A3B export in **f16** device KV → f16 read-out + lossless paged
round-trip; and CPU-in-place-GQA f32 (needs its own H≥2 ORT-oracle
fixture — not free). Both correctly gated to non-paged today.

### Reviews
Independent opus review, author-lockout enforced: Mary (native-decode
specialist) **APPROVE** — 8/8 items verified with reproduced evidence on
GPU 0, full mutation battery.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 31, 2026
…code (GAP-3 Inc-D.1) (#568)

## GAP-3 Inc-D.1 — f16 device-resident present-KV read-out → paged
native CUDA decode

Relaxes the Inc-D (`#567`) device paged gate `cuda && f32 && rank==4` →
`cuda && (f32||f16) && rank==4`, so **f16 device-resident rank-4 CUDA
GQA present-KV** runs **paged** instead of falling to the Inc-A
non-paged path. This is the unlock for real fp16 models (e.g.
gemma4-e2b, whose decoder KV is confirmed FLOAT16) to decode natively
paged.

### How (dtype handling in native read/seed only)
- `native_decode/tensor.rs`: `kv_dtype_to_f32` widens f16→f32 via `half`
`to_f32_vec` (identical to ORT `to_vec_f32_lossy`);
`f32_slice_to_dtype_bytes` narrows f32→f16 via `half::f16::from_f32`
(identical to ORT `from_f32_slice_as`) — NOT the logits bit-twiddle.
Narrower shared with the embedding-input path (DRY).
- `native_decode/cuda.rs`: gate `kv_bindings_paged_rank4`; dtype-branch
the read-out + seed. Host paged store stays **f32 for f16 models**
(identical to ORT) → the Inc-C/D byte-equality oracle is preserved
unchanged.
- bf16, CPU-in-place-GQA, non-rank-4, sink-discontinuous stay gated →
non-paged fallback (bf16 additionally bails defensively inside the
convert).

### Correctness
- New `native_paged_prefix_reuse_matches_ort_on_cuda_device_f16`:
paged-native-f16 == non-paged-native-cold == ORT-cold tokens;
device-mirrored pages **byte-equal** CUDA-vs-ORT (f32 store both sides);
`reused=4`.
- Convert unit tests: native widen == `half` reference; **f16→f32→f16
bit-exact across all 65536 non-NaN f16 patterns**; f32 identity.
- New fixture `tiny-gemma4-vlm-cuda-f16` (Concat-KV so an ORT oracle
exists; `value=key*2` bit-exact — the `+0.5` variant was rejected for
hitting an f16 round-to-even midpoint).
- Non-vacuity (independently re-run by reviewer): gate→f32-only,
raw-u16-as-f32 wrong-convert, and mirror-disabled mutations ALL fail.
- No regressions: Inc-D #567 f32 path still green, Inc-C #566, Inc-A
#565, #541/#543 hybrid, #554 reuse (14/14), 354 lib tests. Scope clean
(no attention/GQA kernel, capture core, provider, `CudaGraphLifecycle`,
or ORT-bridge edits).

### Remaining (follow-up Inc-D.2)
qwen3-30b-a3b is `torch_dtype=bfloat16`; if its ONNX export keeps bf16
KV it stays gated → Inc-D.2 flips the bf16 arm (helpers already have it)
after confirming ORT widens bf16→f32 in its paged store. MoE FFN
produces no KV (orthogonal); present-KV dtype is the only decode-path
gate.

### Reviews
Independent opus review with strict author-lockout: Harry **APPROVE** —
convert matches ORT exactly, round-trip bit-exact (all 65536 patterns),
bf16 excluded, full mutation battery fires, scope clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 31, 2026
#568) (#569)

Consolidates the accumulated GAP-3 native-paged-decode decision inbox (7
notes) into the canonical `.squad/decisions.md`, clearing the drop-box.

Folds in the decisions from the 4 merged GAP-3 PRs plus the Scan-capture
deferral:
- **#565** Inc-A — native multi-component pipeline construction
- **#566** Inc-C — host present-KV mirroring → paged native decode
- **#567** Inc-D — device-resident f32 CUDA present-KV read-out → paged
- **#568** Inc-D.1 — f16 device present-KV read-out → paged
- Inc-C test-rigor fix (byte-equality oracle pattern)
- Scan-capture slice-1b **deferral** (needs per-EP device-graph
handle-keyed registry; resumption trigger recorded)
- Scan slice-1a (#564)

decisions.md 20465 → 28227 bytes (under the 30KB healthy threshold; no
archiving needed). Inbox now README-only. Docs-only change under
`.squad/`.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 3, 2026
…+ GAP-3 decomposition (increment 1/N) (#613)

## What & why

Scope pass on **GAP-3 (native pipeline decode)**. The finding is that
GAP-3's core is
**already implemented and conformance-locked on `origin/main`** — the
task was authored
against a local checkout (`1ba215ee`) that is ~35+ merged PRs behind
`origin/main`
(`be6d4e34`, #612). So this PR does **not** add new decode
functionality. It is a
bounded, **zero-behavior-change** truth-up + a design/decomposition
drop.

This is **GAP-3 housekeeping increment 1 of N** (see the design note in
this diff:

`.squad/decisions/inbox/cohaagen-gap3-native-pipeline-decode-design.md`).
The
substantive increments already landed:

- Backend-neutral component ownership seam — #546
- `NativePipelineDecoder` via `PipelineDecoderComponent` (Inc2b) — #479
- Pure-native multi-component decode wiring (Inc-A) — #565
- Native present-KV mirroring, paged (Inc-C) — #566
- Device-resident present-KV read-out (Inc-D / D.1) — #567/#568
- rank-3 mrope native positions — #543 · text-only decode pipeline —
#535 · fp16 TopK MoE router — #612

The pipeline decode loop (`PipelineDecodeLoopBackend`) already owns
`Box<dyn PipelineDecoderComponent>` + `Box<dyn ComponentSession>`, not
an ORT `Session`;
`DecodeState`/ORT `Value` are confined to `OrtPipelineDecoder`. The
Qwen3.5-0.8B hybrid
(same class as Qwen3.6-35B-A3B) decodes natively with **token-for-token
parity vs ORT**
under `tests/qwen35_0_8b_hybrid_native_cuda_e2e.rs`.

## The change

`native_component.rs`'s module doc still claimed wiring native sessions
into *"the
ORT-owned pipeline decode loop is the remaining GAP 3 work"* — false
since #546/#565.
Corrected to describe the now-backend-neutral loop and the merged
Inc-A/C/D, and to name
the genuinely-remaining feature-sized gaps (non-flat plans, native
cross-attn/vision KV).

**Behavior-preserving:** comment-only in `native_component.rs` + a
tracked decision drop.
No code path, signature, or data change.

## Remaining decomposition (in the design note)

Each is feature-sized (needs op/attention support and/or fixtures —
**not** zero-behavior),
none blocks the text-only 35B-A3B native number: R2 native
sliding-window paged mirror ·
R3 Inc-D.2 discontinuous prefix reuse · R4 native cross-attn/vision KV
(Inc3) · R5 non-flat
plans native · R6 (optional) neutral host tensor in the shared pool.

## Verification for the first native pipeline model

Byte/token-exact differential vs an ORT-backend decode of the **same
artifact** (ORT
front-end for both arms, decoder EP isolated), greedy, token-for-token —
already in place
for the 0.8B hybrid; same harness pattern applied to the real 35B-A3B is
the next step.

## Checks

- `cargo fmt --all` clean
- `cargo clippy -p onnx-genai-engine --features "native-backend cuda" --
-D warnings` clean

Left **open for Harry review**; do not merge without it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant