Repository navigation
feat(pipeline): native present-KV mirroring for paged native decode (GAP-3 Inc-C) - #566
Conversation
Close the S2 bail at pipeline/decoder_component.rs (NativePipelineDecoder::mirror_last_present_kv) so the pure-native pipeline path from Inc-A (#565) can run PAGED, not just the non-paged flat-AR path. This is the present-KV threading increment paged multi-component native serving needs. Host-resident growable f32 KV path only (supports_host_kv_mirror gate): implements the full paged round-trip — present-KV mirror-write + shared-prefix seed-read — reusing the SAME kv_bridge geometry primitives the ORT decoder uses (extract_present_token + append_token_kv), so native and ORT mirror byte-identical pages. Device-resident (CUDA) and in-place-GQA (CPU) present-KV read-out is Inc-D; those decoders keep the Inc-A non-paged path (no regression). - native_decode/mod.rs: +supports_host_kv_mirror, host_present_kv, seed_growable_kv on NativeDecodeSession. - decoder_component.rs: real mirror_last_present_kv; trait supports_paged_kv + load_paged_prefix seams; native load_paged_prefix seeds the growable cache. - flat_autoregressive.rs: build native decoder up front, admit it to the paged gate, DRY-factor claim_paged_prefix shared by ORT and native admit paths. - No decode-loop / native_decode core / capture-core changes. Correctness (token-exact, two-tier): new paged cross-request reuse test on tiny-gemma4-vlm asserts warm native reuse (prefix_reused_tokens=4>0) == cold pure-native == ORT oracle. Non-vacuous: reverting the bail errors, wrong geometry diverges tokens, silent no-op reuse fails reused>0. Regressions green: #565, #384, #541, #543, #554; #544 env-gated; 350 lib unit tests. fmt/clippy clean; no-feature build verified. 35B-A3B: not yet native end-to-end on GPU (device-resident KV → Inc-D). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…l (test-rigor) The Inc-C parity test native_paged_prefix_reuse_matches_fresh_and_ort was vacuous for KV geometry: it asserted warm==cold==ORT on argmax tokens, but on tiny-gemma4-vlm the argmax is invariant to the reused-prefix KV, so a key/value SWAP and even a fully ZEROED mirrored KV still produced identical tokens and PASSED. Only the no-op mirror and forced-reused=0 mutations failed. Add a direct byte/element-equality check on the mirrored paged KV, giving the test discriminating power over the mirror geometry regardless of argmax sensitivity: - pipeline/mod.rs: add a minimal, read-only, native-backend-gated test-support accessor PipelineEngine::materialize_published_prefix_kv. It reconstructs the published prefix key (digest_request_identity + prefix_key), does a non-mutating PrefixCache::lookup, attaches the pages to a throwaway sequence to read them, materializes per-layer K/V, then drops the sequence WITHOUT freeing the pages (the prefix cache keeps sole ownership). No decode-path behavior changes. - test: after the warm native run, assert the native-mirrored paged KV for the shared prefix equals the ORT-mirrored KV byte-for-byte (MaterializedKv: PartialEq). Catches key/value swap, zeroing, head-stride, seq-offset and page-index errors. Reuse>0 and token asserts retained. Test-only: native_decode/mod.rs, pipeline/decoder_component.rs and pipeline/flat_autoregressive.rs are byte-identical to HEAD. Non-vacuity (each applied, run on GPU 0, reverted): (a) no-op mirror FAILS (publishes unmirrored pages), (b) key/value swap FAILS on the new byte assert (tokens still matched), (c) forced reused=0 FAILS (reused>0 assert). Clean run green (2 passed, --test-threads=1); fmt/clippy clean. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
VERDICT: APPROVE Independent opus re-review (Harry) of
Prior REJECT (Mary) on vacuous geometry test was correctly remediated test-only; underlying Inc-C production code was correct-by-construction throughout. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #566 +/- ##
=======================================
Coverage 81.55% 81.55%
=======================================
Files 315 315
Lines 123574 123574
Branches 123574 123574
=======================================
+ Hits 100777 100779 +2
+ Misses 18739 18738 -1
+ Partials 4058 4057 -1
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
✅ Benchmarks — No RegressionComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
…DA decode (GAP-3 Inc-D) (#567) ## GAP-3 Inc-D — device-resident present-KV read-out → paged native CUDA decode Lifts the Inc-C (`#566`) `supports_paged_kv=false` gate for **device-resident f32 rank-4 CUDA GQA present-KV**, so a native CUDA pipeline decoder now runs **paged** (cross-request KV reuse) instead of the Inc-A non-paged fallback. Closes the present-KV threading gap for Qwen3.6-35B-A3B GPU decode. ### How (pure post-step plumbing — no kernel/capture-core edits) - `native_decode/cuda.rs`: `read_present_kv` reads the KV binding **after** the decode step's existing stream sync (`read_bytes`→`copy_to_host`→`dtoh`, which synchronizes) using the **physical/capacity** shape `[1,H,max_len,Dh]` so strides address the padded buffer; `seed_prefix` is the device counterpart of Inc-C's host seed; `device_present_kv_view` isolates the physical-shape handling. - `native_decode/mod.rs`: `present_kv`/`seed_kv`/`supports_device_kv_mirror` unify host-growable (Inc-C) and device-CUDA (Inc-D) onto the **same** `extract_present_token`/`append_token_kv` geometry + host f32 paged store (DRY, byte-comparable with ORT). - `pipeline/decoder_component.rs`: `supports_paged_kv = host OR device`; `mirror_last_present_kv &self→&mut self` (rippled to trait + ORT impl); `load_paged_prefix → seed_kv`. ### Correctness - `native_paged_prefix_reuse_matches_ort_on_cuda_device`: paged-native-CUDA == non-paged-native-CUDA == ORT oracle == closed-form tokens; mirrored pages **byte-equal** CUDA-vs-ORT; `reused=4`. - **H=2 unit test** for the physical-vs-logical stride bug (all existing fixtures are H=1, where the head-stride error is invisible). Mutating to the logical stride fails it. - Non-vacuity (independently re-run by reviewer): gate-revert, logical-stride, and forced-reused=0 mutations ALL fail. - Honest gating: f16 / non-rank-4 / CPU-in-place-GQA / sink-discontinuous stay `supports_paged_kv=false` → Inc-A non-paged fallback (no silent-wrong paged run). - No regressions: Inc-C #566, Inc-A #565, #541/#543 hybrid, #554 reuse (14/14), native_decode lib (54); scope clean (no standard_attention/GQA kernel, capture core, or `plan_capture_region` edits). ### Remaining (follow-up Inc-D.1) Real 35B-A3B export in **f16** device KV → f16 read-out + lossless paged round-trip; and CPU-in-place-GQA f32 (needs its own H≥2 ORT-oracle fixture — not free). Both correctly gated to non-paged today. ### Reviews Independent opus review, author-lockout enforced: Mary (native-decode specialist) **APPROVE** — 8/8 items verified with reproduced evidence on GPU 0, full mutation battery. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…code (GAP-3 Inc-D.1) (#568) ## GAP-3 Inc-D.1 — f16 device-resident present-KV read-out → paged native CUDA decode Relaxes the Inc-D (`#567`) device paged gate `cuda && f32 && rank==4` → `cuda && (f32||f16) && rank==4`, so **f16 device-resident rank-4 CUDA GQA present-KV** runs **paged** instead of falling to the Inc-A non-paged path. This is the unlock for real fp16 models (e.g. gemma4-e2b, whose decoder KV is confirmed FLOAT16) to decode natively paged. ### How (dtype handling in native read/seed only) - `native_decode/tensor.rs`: `kv_dtype_to_f32` widens f16→f32 via `half` `to_f32_vec` (identical to ORT `to_vec_f32_lossy`); `f32_slice_to_dtype_bytes` narrows f32→f16 via `half::f16::from_f32` (identical to ORT `from_f32_slice_as`) — NOT the logits bit-twiddle. Narrower shared with the embedding-input path (DRY). - `native_decode/cuda.rs`: gate `kv_bindings_paged_rank4`; dtype-branch the read-out + seed. Host paged store stays **f32 for f16 models** (identical to ORT) → the Inc-C/D byte-equality oracle is preserved unchanged. - bf16, CPU-in-place-GQA, non-rank-4, sink-discontinuous stay gated → non-paged fallback (bf16 additionally bails defensively inside the convert). ### Correctness - New `native_paged_prefix_reuse_matches_ort_on_cuda_device_f16`: paged-native-f16 == non-paged-native-cold == ORT-cold tokens; device-mirrored pages **byte-equal** CUDA-vs-ORT (f32 store both sides); `reused=4`. - Convert unit tests: native widen == `half` reference; **f16→f32→f16 bit-exact across all 65536 non-NaN f16 patterns**; f32 identity. - New fixture `tiny-gemma4-vlm-cuda-f16` (Concat-KV so an ORT oracle exists; `value=key*2` bit-exact — the `+0.5` variant was rejected for hitting an f16 round-to-even midpoint). - Non-vacuity (independently re-run by reviewer): gate→f32-only, raw-u16-as-f32 wrong-convert, and mirror-disabled mutations ALL fail. - No regressions: Inc-D #567 f32 path still green, Inc-C #566, Inc-A #565, #541/#543 hybrid, #554 reuse (14/14), 354 lib tests. Scope clean (no attention/GQA kernel, capture core, provider, `CudaGraphLifecycle`, or ORT-bridge edits). ### Remaining (follow-up Inc-D.2) qwen3-30b-a3b is `torch_dtype=bfloat16`; if its ONNX export keeps bf16 KV it stays gated → Inc-D.2 flips the bf16 arm (helpers already have it) after confirming ORT widens bf16→f32 in its paged store. MoE FFN produces no KV (orthogonal); present-KV dtype is the only decode-path gate. ### Reviews Independent opus review with strict author-lockout: Harry **APPROVE** — convert matches ORT exactly, round-trip bit-exact (all 65536 patterns), bf16 excluded, full mutation battery fires, scope clean. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#568) (#569) Consolidates the accumulated GAP-3 native-paged-decode decision inbox (7 notes) into the canonical `.squad/decisions.md`, clearing the drop-box. Folds in the decisions from the 4 merged GAP-3 PRs plus the Scan-capture deferral: - **#565** Inc-A — native multi-component pipeline construction - **#566** Inc-C — host present-KV mirroring → paged native decode - **#567** Inc-D — device-resident f32 CUDA present-KV read-out → paged - **#568** Inc-D.1 — f16 device present-KV read-out → paged - Inc-C test-rigor fix (byte-equality oracle pattern) - Scan-capture slice-1b **deferral** (needs per-EP device-graph handle-keyed registry; resumption trigger recorded) - Scan slice-1a (#564) decisions.md 20465 → 28227 bytes (under the 30KB healthy threshold; no archiving needed). Inbox now README-only. Docs-only change under `.squad/`. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…+ GAP-3 decomposition (increment 1/N) (#613) ## What & why Scope pass on **GAP-3 (native pipeline decode)**. The finding is that GAP-3's core is **already implemented and conformance-locked on `origin/main`** — the task was authored against a local checkout (`1ba215ee`) that is ~35+ merged PRs behind `origin/main` (`be6d4e34`, #612). So this PR does **not** add new decode functionality. It is a bounded, **zero-behavior-change** truth-up + a design/decomposition drop. This is **GAP-3 housekeeping increment 1 of N** (see the design note in this diff: `.squad/decisions/inbox/cohaagen-gap3-native-pipeline-decode-design.md`). The substantive increments already landed: - Backend-neutral component ownership seam — #546 - `NativePipelineDecoder` via `PipelineDecoderComponent` (Inc2b) — #479 - Pure-native multi-component decode wiring (Inc-A) — #565 - Native present-KV mirroring, paged (Inc-C) — #566 - Device-resident present-KV read-out (Inc-D / D.1) — #567/#568 - rank-3 mrope native positions — #543 · text-only decode pipeline — #535 · fp16 TopK MoE router — #612 The pipeline decode loop (`PipelineDecodeLoopBackend`) already owns `Box<dyn PipelineDecoderComponent>` + `Box<dyn ComponentSession>`, not an ORT `Session`; `DecodeState`/ORT `Value` are confined to `OrtPipelineDecoder`. The Qwen3.5-0.8B hybrid (same class as Qwen3.6-35B-A3B) decodes natively with **token-for-token parity vs ORT** under `tests/qwen35_0_8b_hybrid_native_cuda_e2e.rs`. ## The change `native_component.rs`'s module doc still claimed wiring native sessions into *"the ORT-owned pipeline decode loop is the remaining GAP 3 work"* — false since #546/#565. Corrected to describe the now-backend-neutral loop and the merged Inc-A/C/D, and to name the genuinely-remaining feature-sized gaps (non-flat plans, native cross-attn/vision KV). **Behavior-preserving:** comment-only in `native_component.rs` + a tracked decision drop. No code path, signature, or data change. ## Remaining decomposition (in the design note) Each is feature-sized (needs op/attention support and/or fixtures — **not** zero-behavior), none blocks the text-only 35B-A3B native number: R2 native sliding-window paged mirror · R3 Inc-D.2 discontinuous prefix reuse · R4 native cross-attn/vision KV (Inc3) · R5 non-flat plans native · R6 (optional) neutral host tensor in the shared pool. ## Verification for the first native pipeline model Byte/token-exact differential vs an ORT-backend decode of the **same artifact** (ORT front-end for both arms, decoder EP isolated), greedy, token-for-token — already in place for the 0.8B hybrid; same harness pattern applied to the real 35B-A3B is the next step. ## Checks - `cargo fmt --all` clean - `cargo clippy -p onnx-genai-engine --features "native-backend cuda" -- -D warnings` clean Left **open for Harry review**; do not merge without it. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
GAP-3 Inc-C — native present-KV mirroring → paged native pipeline decode
Closes the S2 bail in
NativePipelineDecoder::mirror_last_present_kvso the pure-native pipeline decoder (Inc-A #565) can run paged (cross-request KV prefix reuse), not only the non-paged flat-AR path. Host-KV path; device-resident present-KV (35B-A3B GPU) is explicitly deferred to Inc-D.What's wired
native_decode/mod.rs:supports_host_kv_mirrorgate,host_present_kv(read),seed_growable_kv(seed).pipeline/decoder_component.rs: realmirror_last_present_kv— byte-identical geometry to ORT'smirror_present_kv_to_pages(sharedextract_present_token/append_token_kv);supports_paged_kv+load_paged_prefixseams.pipeline/flat_autoregressive.rs: DRY-factoredclaim_paged_prefixshared by ORT + native admit paths (only the KV sink differs).pipeline/mod.rs: read-only#[cfg(feature="native-backend")]test accessormaterialize_published_prefix_kv(non-mutating prefix-cache read-back).tests/native_pipeline_backend_selection_parity.rs:native_paged_prefix_reuse_matches_fresh_and_ort— warm==cold==ORT tokens plus direct native-vs-ORT paged-KV byte-equality assertion.Correctness
reused=4.plan_capture_region,executor/capture.rs,CudaGraphLifecycle).Deferred to Inc-D (35B-A3B GPU end-to-end)
Device-resident present-KV read-out, in-place-GQA CPU KV, f16/non-rank-4 round-trip, MoE routed-expert specifics if needed. These decoders correctly gate
supports_paged_kv=false→ Inc-A non-paged fallback (no silent-wrong paged run).Reviews
Independent opus review with strict author-lockout: Mary REJECTED the original (vacuous geometry test) → revised the test as authorized non-author → Harry independent re-review APPROVE (all mutations re-run, byte-equality confirmed, production byte-identical, deferral gate verified).
Note: run the parity test binary with
--test-threads=1(two tests share a process-global decoder-device env var; parallel races flake — pre-existing on HEAD).