Repository navigation
feat(pipeline): native CUDA device-KV decoder via inputs_embeds (native multi-component inc3a) - #485
Conversation
…l root-cause + Inc3a/Inc3b split) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ve multi-component inc3a) Lift the CUDA-target refusal of metadata-declared `inputs_embeds` step inputs in the native decoder so a fused VLM decoder can run on the CUDA EP while keeping its KV cache device-resident — the 35B-A3B GPU native-decode unblock flagged in Inc2b. The prior refusal (native_decode/load.rs + mod.rs) guarded an unimplemented on-device binding path, not a correctness barrier: `DecodeCudaState` hardwired an Int64 `[1,1]` token binding with no float `inputs_embeds [1,1,hidden]` sequence binding. Inc3a implements that path by routing the inputs_embeds decode through the existing eager device forward (`run_cuda_eager_rows`), which already binds the sequence tensor as an owned per-step host input against the persistent device mask + KV bindings. So only one token's embedding (`[1,1,hidden]`) crosses host->device each step; the KV cache never round-trips — the same guarantee as the CPU Inc2b path, now on the CUDA EP. - native_decode/cuda.rs: `DecodeCudaIo.inputs_embeds` + `CudaEmbedsBinding`; a shared `run_cuda_eager_rows_owned` body (token path byte-identical) reused by a new `decode_cuda_inputs_embeds` step; `DecodeCudaState::new` allocates a float `[1,1,hidden]` sequence binding when embeds are present, else the Int64 token binding (default path unchanged). - native_decode/load.rs: relax the CUDA refusal — accept a declared `inputs_embeds` sequence source (resolve dtype/hidden), still refuse arbitrary generic `Routed` ports. - native_decode/mod.rs, backend.rs: thread `step_inputs` into `decode_cuda`. - pipeline/mod.rs: `ONNX_GENAI_PIPELINE_NATIVE_DECODER_DEVICE=cuda[:index]` selects the native decoder device (default cpu), gated by the existing `ONNX_GENAI_PIPELINE_NATIVE_DECODER` flag. - scripts/build_tiny_gemma4_vlm_cuda.py + tests/fixtures/tiny-gemma4-vlm-cuda: CUDA-capable fixture (declares attention_mask + position_ids, closed-form tokens `[0,5,6,7]` unchanged). - tests/native_cuda_pipeline_decoder_parity.rs: native-CUDA decoder in the pipeline produces token ids identical to the native-CPU baseline (both `[0,5,6,7]`), proving on-GPU device-KV parity. Refs #384. Stacks on #479 (native CPU device-KV decoder). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #485 +/- ##
=======================================
Coverage 81.53% 81.53%
=======================================
Files 315 315
Lines 122793 122793
Branches 122793 122793
=======================================
Hits 100117 100117
Misses 18648 18648
Partials 4028 4028
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
VERDICT: APPROVE Independent review of Inc3a (native CUDA device-KV decoder via inputs_embeds), branch squad/native-multi-component-pipeline-inc3 @ fa9a626, base origin/main. Reviewer: Lori. Correctness is the gate; all gates pass. Two advisory items to track on #384 (neither blocks this plumbing slice). 1. On-GPU token parity is genuine (not a CPU fallback)
2. KV stays device-resident; only the one-token embedding crosses per step
3. Eager-forward reuse is sound (with a perf caveat)
4. The unconsumed-mask / ReduceSum fixture caveat — legitimate, but a coverage gap to flag
5. Default ORT + CPU path unchanged; env-gated; cfg-correct
6. Verification evidence
VerdictAPPROVE. On-GPU execution is proven, KV stays device-resident (only the one-token embedding uploads per step), the eager-forward reuse is semantically identical to the captured path, the default/ORT/CPU paths are untouched, goldens are unchanged, clippy/fmt are clean, and no new test failures are introduced. The two advisory items for #384 — (a) eager-per-step vs graph-capture perf for the inputs_embeds decode, and (b) the coverage gap that no CONSUMED attention mask is exercised on the native CUDA path — should be tracked but do not block this Inc3a plumbing slice. Inc3b (generic Routed ports on CUDA) remains correctly refused. |
…ve multi-component inc3b) Lift the remaining CUDA refusal deferred in Inc3a: the native CUDA decoder now binds arbitrary declared non-KV `Routed` step-input ports on-device per step, so cross-component handoffs (a routed hidden/state edge, and eventually static_cross_kv) work on the CUDA EP — generalizing the Inc3a inputs_embeds path rather than forking it. The CPU path already builds its owned per-step input set generically (`prepare_cpu_step_inputs`). Inc3b gives the CUDA eager path the same treatment: a new `prepare_cuda_owned_step_inputs` iterates every declared step input, generating token/position ids and pulling inputs_embeds/routed tensors from the supplied set by exact graph-port name, with the one CUDA-specific exclusion that `attention_mask` is a persistent device binding (filled by extend_mask), never an owned upload. Routed ports are owned per-step uploads — no new persistent device binding and no DecodeCudaState binding-table change. Only the small per-step tensors (embedding + routed state) cross host->device; the mask and KV cache stay device-resident on the GPU. - native_decode/cuda.rs: replace the embeds-only `decode_cuda_inputs_embeds` with a generic `decode_cuda_eager_step_inputs` + `prepare_cuda_owned_step_inputs` (mirrors the CPU contract); `decode_cuda` takes the eager path whenever any inputs_embeds/routed port is declared (`has_eager_step_inputs`), keeping the pure token-id captured fast path byte-identical. - native_decode/load.rs: remove the CUDA `Routed` refusal (routed ports are now bound generically); inputs_embeds metadata resolution from Inc3a unchanged. - scripts/build_tiny_gemma4_vlm_cuda_routed.py + tests/fixtures/ tiny-gemma4-vlm-cuda-routed: fixture where the every_step embedding emits a second `router_state` output routed to a decoder `router_state` port, consumed via a real MatMul-by-zero (closed-form tokens `[0,5,6,7]` unchanged). - tests/native_cuda_routed_pipeline_decoder_parity.rs: native decoder CPU vs CUDA EP (device 4) through the pipeline both produce `[0,5,6,7]`, proving the routed port binds on-device with the KV kept resident. Scope out (deferred): vision cross-KV (needs the vision Attention float-mask fixes) and the static_cross_kv upload-once optimization. Refs #384. Stacks on #485 (native CUDA inputs_embeds decoder). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ve multi-component inc3b) (#487) ## Inc3b — generic routed ports on the native CUDA decoder Lifts the remaining CUDA refusal deferred in Inc3a. The native CUDA decoder now binds arbitrary declared non-KV **`Routed`** step-input ports on-device per step (beyond just `inputs_embeds`), enabling cross-component handoffs (a routed hidden/state edge, and eventually `static_cross_kv`) on the CUDA EP while keeping the KV cache device-resident. ### What changed - **`native_decode/cuda.rs`**: replaced the embeds-only `decode_cuda_inputs_embeds` with a generic `decode_cuda_eager_step_inputs` + `prepare_cuda_owned_step_inputs`, mirroring the CPU `prepare_cpu_step_inputs` contract. `decode_cuda` takes the eager (uncaptured) path whenever any inputs_embeds/routed port is declared (`has_eager_step_inputs`), keeping the pure token-id **captured** fast path byte-identical. The one CUDA-specific exclusion: `attention_mask` is a persistent device binding (filled by `extend_mask`), never an owned upload — so routed ports need **no new persistent binding** and **no `DecodeCudaState` change**; only the small per-step tensors cross host→device. - **`native_decode/load.rs`**: removed the CUDA `Routed` refusal; Inc3a inputs_embeds metadata resolution unchanged. - **Fixture + test**: `tiny-gemma4-vlm-cuda-routed` — every_step embedding emits a second `router_state` output routed to a decoder `router_state` port consumed via a real MatMul-by-zero (closed-form tokens `[0,5,6,7]` unchanged). `native_cuda_routed_pipeline_decoder_parity.rs` proves native decoder **CPU vs CUDA EP (device 4)** through the pipeline both produce `[0,5,6,7]`. ### Proof `native_cuda_routed_pipeline_decoder_matches_cpu_token_ids` — GREEN on device 4 (native-CPU == native-CUDA == `[0,5,6,7]`). ORT-default goldens unchanged; full native-backend failing set identical to base (12 pre-existing, zero new). fmt clean; clippy clean ×4 (default / native-backend / cuda / cuda,native-backend). ### Scope out (deferred) Vision cross-KV (needs the vision Attention float-mask fixes, separate #384 blocker) and the `static_cross_kv` upload-once optimization. Stacks on #485 (native CUDA inputs_embeds decoder, now merged). Refs #384. Do not merge yet. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…e/hybrid wave logs (#532) Scribe round 5. Records the native multi-component CUDA pipeline decode wave (#484/#485/#486/#487/#525) and distills decisions.md 28520→19858 bytes (under the 20480 gate). State-only; no production code. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…s native multi-component decode, #82/#384/35B-A3B) (#546) ## Summary The task was to introduce a backend-neutral ownership seam so the pipeline decode loop can drive **either** ORT **or** native component sessions per step, then route the existing ORT path through it byte-identically (increment-1). **On inspection, that seam already fully exists on `origin/main`** — it landed across the inc1→inc3c chain (#450, #478, #479, #485, #487, #533) and was hardened by #543. `PipelineDecodeLoopBackend` owns **no** ORT `Session`/decode-state; it holds only `Box<dyn PipelineDecoderComponent>` (stateful decoder seam) and `Vec<(_, Box<dyn ComponentSession>)>` (stateless every_step seam), and both ORT and native backends are driven through one decode loop via runtime env selection. So increment-1 here is the piece the chain had **not** locked: a parity test proving the seam's **keystone** end-state — *every declared component running natively at once* (native every_step embedding **+** native device-KV decoder in the same loop), the exact shape a large multi-component package (up to the 35B-A3B 3-component package) decodes through. ## What changed (test-only, zero production change) - **`crates/onnx-genai-engine/tests/native_full_pipeline_parity.rs`** — drives the `tiny-gemma4-vlm` composite with both `ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding` **and** `ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder`, asserting the fully-native run is token-identical to the ORT baseline `[0, 5, 6, 7]`. - **`crates/onnx-genai-engine/Cargo.toml`** — registers the test (`required-features = ["native-backend"]`, CPU-only). - **`.squad/decisions/inbox/mary-pipeline-native-ownership.md`** — full assessment (ownership map), design affirmation, and the deferred next increment. Prior increments proved each slice in isolation: inc1 (native embedding + ORT decoder), inc2b (ORT embedding + native decoder). Nothing exercised **both** natively at once until now. ## ORT byte-identical proof No production source is touched, so the ORT decode path is byte-identical to `origin/main` by construction. Empirically the ORT baseline `[0,5,6,7]` and the fully-native run `[0,5,6,7]` match exactly. ## Tests - `native_full_pipeline_parity` — **pass** (new) - `native_step_component_parity`, `native_pipeline_decoder_parity` — **pass** - 343 engine lib unit tests — **pass**, 1 ignored - `cargo fmt --all --check` — clean CUDA-gated native tests were not run in this CPU environment (unchanged by this PR). ## Deferred to the next increment (native wiring completion) The one genuine remaining hard limitation the code itself flags (`decoder_component.rs:244-260`): `NativePipelineDecoder::mirror_last_present_kv` bails — the native decoder keeps KV session-resident and does not expose host present tensors, so native selection runs the non-paged, fresh-decode path with no cross-request KV reuse. Wiring native present-KV exposure + paged mirroring is higher blast radius and is intentionally **not** bundled here. Refs #82, #384. Working as Mary (native-decode / pipeline engineer). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Increment 3a — native CUDA device-KV decoder with
inputs_embedsLifts the CUDA-target refusal of metadata-declared
inputs_embedsstep inputs in the native decoder, so a fused VLM decoder runs on the CUDA EP while keeping its KV cache device-resident — the 35B-A3B GPU native-decode unblock flagged in Inc2b.Builds on #479 (native CPU device-KV decoder), now merged to
main; this PR is rebased ontomain. Refs #384.Refusal root cause
The prior refusal (
native_decode/load.rs+mod.rs) guarded an unimplemented on-device binding path, not a correctness barrier:DecodeCudaStatehardwired anInt64 [1,1]token binding with no floatinputs_embeds [1,1,hidden]sequence binding and no routed-tensor upload.Approach (minimal, low-risk)
Route the
inputs_embedsdecode through the existing eager device forward (run_cuda_eager_rows), which already binds the sequence tensor as an owned per-step host input against the persistent device mask + KV bindings. Result: only one token's embedding ([1,1,hidden]) crosses host→device per step; the KV cache never round-trips — identical guarantee to the CPU Inc2b path, now on the CUDA EP. No captured-graph / persistent-embeds-binding surgery.Inc3a / Inc3b split
inputs_embeds, on-GPU device-KV, native-CUDA-in-pipeline token parity proven.Routedports on CUDA (still refused here); paged cross-request reuse and vision cross-KV remain scoped out.On-GPU token-parity proof
native_cuda_pipeline_decoder_parityruns the composite pipeline overtiny-gemma4-vlm-cudawithONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder, comparing the native decoder on CPU (..._DEVICE=cpu) vs the CUDA EP (..._DEVICE=cuda:0, GPU device 4 viaCUDA_VISIBLE_DEVICES=4). Both produce[0, 5, 6, 7](≥2 real decode steps). GREEN.Verification
cargo test -p onnx-genai-engine --features cuda,native-backend --test native_cuda_pipeline_decoder_parity→ 1 passed (GPU device 4).cargo test -p onnx-genai-engine --features native-backend→ ORT/pipeline goldens unchanged (multimodal_reuse 14, optional_modality 8, vlm_multibinding 2, pipeline_executor 1, tts_*, inc1native_step_component_parity, inc2bnative_pipeline_decoder_parityall pass). Pre-existingnative_engine(7),native_speculative_driver(2),gemma4_assistant_full(1) failures reproduce identically on the base (same "3 ports matchmodel.io.token_input" fixture heuristic) — not introduced here.cargo fmt --all --checkclean.Do not merge.