Skip to content

feat(pipeline): native CUDA device-KV decoder via inputs_embeds (native multi-component inc3a) - #485

Merged
justinchuby merged 2 commits into
mainfrom
squad/native-multi-component-pipeline-inc3
Jul 30, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/native-multi-component-pipeline-inc3

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Increment 3a — native CUDA device-KV decoder with inputs_embeds

Lifts the CUDA-target refusal of metadata-declared inputs_embeds step inputs in the native decoder, so a fused VLM decoder runs on the CUDA EP while keeping its KV cache device-resident — the 35B-A3B GPU native-decode unblock flagged in Inc2b.

Builds on #479 (native CPU device-KV decoder), now merged to main; this PR is rebased onto main. Refs #384.

Refusal root cause

The prior refusal (native_decode/load.rs + mod.rs) guarded an unimplemented on-device binding path, not a correctness barrier: DecodeCudaState hardwired an Int64 [1,1] token binding with no float inputs_embeds [1,1,hidden] sequence binding and no routed-tensor upload.

Approach (minimal, low-risk)

Route the inputs_embeds decode through the existing eager device forward (run_cuda_eager_rows), which already binds the sequence tensor as an owned per-step host input against the persistent device mask + KV bindings. Result: only one token's embedding ([1,1,hidden]) crosses host→device per step; the KV cache never round-trips — identical guarantee to the CPU Inc2b path, now on the CUDA EP. No captured-graph / persistent-embeds-binding surgery.

Inc3a / Inc3b split

  • Inc3a (this PR): CUDA native decoder accepts inputs_embeds, on-GPU device-KV, native-CUDA-in-pipeline token parity proven.
  • Inc3b (deferred): generic arbitrary Routed ports on CUDA (still refused here); paged cross-request reuse and vision cross-KV remain scoped out.

On-GPU token-parity proof

native_cuda_pipeline_decoder_parity runs the composite pipeline over tiny-gemma4-vlm-cuda with ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder, comparing the native decoder on CPU (..._DEVICE=cpu) vs the CUDA EP (..._DEVICE=cuda:0, GPU device 4 via CUDA_VISIBLE_DEVICES=4). Both produce [0, 5, 6, 7] (≥2 real decode steps). GREEN.

Verification

  • cargo test -p onnx-genai-engine --features cuda,native-backend --test native_cuda_pipeline_decoder_parity → 1 passed (GPU device 4).
  • cargo test -p onnx-genai-engine --features native-backend → ORT/pipeline goldens unchanged (multimodal_reuse 14, optional_modality 8, vlm_multibinding 2, pipeline_executor 1, tts_*, inc1 native_step_component_parity, inc2b native_pipeline_decoder_parity all pass). Pre-existing native_engine (7), native_speculative_driver (2), gemma4_assistant_full (1) failures reproduce identically on the base (same "3 ports match model.io.token_input" fixture heuristic) — not introduced here.
  • cargo fmt --all --check clean.
  • clippy clean across default / native-backend / cuda / cuda,native-backend (cfg-correct; CUDA-decoder code guarded behind the right cfgs).

Do not merge.

justinchuby and others added 2 commits July 30, 2026 18:22
…l root-cause + Inc3a/Inc3b split)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ve multi-component inc3a)

Lift the CUDA-target refusal of metadata-declared `inputs_embeds` step inputs
in the native decoder so a fused VLM decoder can run on the CUDA EP while
keeping its KV cache device-resident — the 35B-A3B GPU native-decode unblock
flagged in Inc2b.

The prior refusal (native_decode/load.rs + mod.rs) guarded an unimplemented
on-device binding path, not a correctness barrier: `DecodeCudaState` hardwired
an Int64 `[1,1]` token binding with no float `inputs_embeds [1,1,hidden]`
sequence binding. Inc3a implements that path by routing the inputs_embeds decode
through the existing eager device forward (`run_cuda_eager_rows`), which already
binds the sequence tensor as an owned per-step host input against the persistent
device mask + KV bindings. So only one token's embedding (`[1,1,hidden]`) crosses
host->device each step; the KV cache never round-trips — the same guarantee as
the CPU Inc2b path, now on the CUDA EP.

- native_decode/cuda.rs: `DecodeCudaIo.inputs_embeds` + `CudaEmbedsBinding`; a
  shared `run_cuda_eager_rows_owned` body (token path byte-identical) reused by a
  new `decode_cuda_inputs_embeds` step; `DecodeCudaState::new` allocates a float
  `[1,1,hidden]` sequence binding when embeds are present, else the Int64 token
  binding (default path unchanged).
- native_decode/load.rs: relax the CUDA refusal — accept a declared
  `inputs_embeds` sequence source (resolve dtype/hidden), still refuse arbitrary
  generic `Routed` ports.
- native_decode/mod.rs, backend.rs: thread `step_inputs` into `decode_cuda`.
- pipeline/mod.rs: `ONNX_GENAI_PIPELINE_NATIVE_DECODER_DEVICE=cuda[:index]`
  selects the native decoder device (default cpu), gated by the existing
  `ONNX_GENAI_PIPELINE_NATIVE_DECODER` flag.
- scripts/build_tiny_gemma4_vlm_cuda.py + tests/fixtures/tiny-gemma4-vlm-cuda:
  CUDA-capable fixture (declares attention_mask + position_ids, closed-form
  tokens `[0,5,6,7]` unchanged).
- tests/native_cuda_pipeline_decoder_parity.rs: native-CUDA decoder in the
  pipeline produces token ids identical to the native-CPU baseline (both
  `[0,5,6,7]`), proving on-GPU device-KV parity.

Refs #384. Stacks on #479 (native CPU device-KV decoder).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Jul 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.53%. Comparing base (d79e258) to head (fa9a626).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main     #485   +/-   ##
=======================================
  Coverage   81.53%   81.53%           
=======================================
  Files         315      315           
  Lines      122793   122793           
  Branches   122793   122793           
=======================================
  Hits       100117   100117           
  Misses      18648    18648           
  Partials     4028     4028           
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.78% <ø> (ø)
mlas 78.72% <ø> (ø)
offline 81.49% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 3.72 ms 6.63 ms +78.5%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 368.18 µs 640.35 µs +73.9%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 942.17 µs 1.56 ms +65.3%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 31.09 µs 42.11 µs +35.4%
🔴 matmul/small_generic_f32_threads=8/1x256x256 33.96 µs 45.82 µs +34.9%
🔴 add/large_bf16_threads=1-internal/4194304 42.28 ms 55.74 ms +31.8%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 82.81 µs 105.75 µs +27.7%
⚠️ matmul/small_generic_bf16_threads=1/1x256x256 31.22 µs 38.45 µs +23.1%
⚠️ matmul/small_generic_f32_threads=1/1x256x256 34.55 µs 42.41 µs +22.8%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 74.08 µs 90.50 µs +22.2%
⚠️ matmul/large_generic_bf16_threads=8/32x1024x1024 1.87 ms 2.26 ms +20.6%
⚠️ matmul/medium_generic_bf16_threads=1/32x512x512 494.89 µs 594.78 µs +20.2%
⚠️ gather/large_bf16_threads=1-internal/131072 14.64 µs 17.08 µs +16.7%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 32.39 µs 37.77 µs +16.6%
⚠️ add/medium_f32_threads=1-internal/262144 2.47 ms 2.88 ms +16.5%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 8.58 ms 9.85 ms +14.8%
✅ matmul/small_generic_f16_threads=8/1x256x256 30.36 µs 34.78 µs +14.6%
✅ add/large_f16_threads=1-internal/4194304 42.54 ms 48.34 ms +13.6%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.91 ms 2.16 ms +13.2%
✅ add/small_f32_threads=1-internal/1024 200.2 ns 225.9 ns +12.9%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.24 ms 2.52 ms +12.7%
✅ gather/large_f16_threads=1-internal/131072 15.18 µs 16.85 µs +11.0%
✅ gather/small_bf16_threads=1-internal/4096 529.3 ns 584.2 ns +10.4%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.26 µs 16.77 µs +9.9%
✅ reduce_mean/medium_f32_threads=1-internal/65536 247.85 µs 272.16 µs +9.8%
✅ add/medium_bf16_threads=1-internal/262144 2.80 ms 3.06 ms +9.5%
✅ add/large_f32_threads=1-internal/4194304 40.24 ms 43.99 ms +9.3%
✅ reduce_mean/large_f32_threads=1-internal/262144 949.11 µs 1.03 ms +8.1%
✅ add/small_bf16_threads=1-internal/1024 12.92 µs 13.81 µs +6.9%
✅ tokenization/encode_tokens_per_second 380.06 µs 404.34 µs +6.4%
✅ matmul/small_generic_bf16_threads=8/1x256x256 32.09 µs 34.11 µs +6.3%
✅ add/small_f16_threads=1-internal/1024 12.97 µs 13.73 µs +5.9%
✅ add/medium_f16_threads=1-internal/262144 2.60 ms 2.72 ms +4.4%
✅ gather/small_f32_threads=1-internal/4096 664.8 ns 687.4 ns +3.4%
✅ sampling_latency/top_p_per_token 1.10 ms 1.14 ms +3.2%
✅ gather/medium_f32_threads=1-internal/32768 4.12 µs 4.23 µs +2.7%
✅ gather/small_f16_threads=1-internal/4096 484.1 ns 494.2 ns +2.1%
✅ matmul/small_generic_f16_threads=1/1x256x256 34.44 µs 35.10 µs +1.9%
✅ sampling_latency/greedy_per_token 3.42 µs 3.38 µs -1.2%
✅ gather/medium_f16_threads=1-internal/32768 2.51 µs 2.46 µs -2.0%
✅ sampling_latency/top_k_per_token 507.25 µs 492.29 µs -2.9%
✅ tokenization/decode_tokens_per_second 7.27 ms 7.01 ms -3.6%
✅ gather/medium_bf16_threads=1-internal/32768 2.56 µs 2.47 µs -3.7%
✅ sampling_latency/min_p_per_token 377.38 µs 355.62 µs -5.8%
✅ logit_processing/seven_processor_chain_per_step 1.45 ms 1.32 ms -9.1%
✅ grammar_masking/llguidance_compute_mask/32 89.41 µs 80.11 µs -10.4%
🟢 kv_cache/alloc_dealloc_pages 48.23 µs 39.86 µs -17.4%
🟢 gather/large_f32_threads=1-internal/131072 47.02 µs 34.92 µs -25.7%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.71 5.70 8.07 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review of Inc3a (native CUDA device-KV decoder via inputs_embeds), branch squad/native-multi-component-pipeline-inc3 @ fa9a626, base origin/main. Reviewer: Lori. Correctness is the gate; all gates pass. Two advisory items to track on #384 (neither blocks this plumbing slice).

1. On-GPU token parity is genuine (not a CPU fallback)

  • Test: tests/native_cuda_pipeline_decoder_parity.rs runs the native decoder on CPU (asserts ids [0,5,6,7], line 91-95) then on cuda:index (asserts equal to the CPU baseline, line 100-103). GREEN on device 5:
    test native_cuda_pipeline_decoder_matches_cpu_token_ids ... ok (1 passed).
  • Genuinely on GPU: while polling nvidia-smi -i 5 during the run, device-5 memory rose 1 -> 434 -> 533 MiB, i.e. real device allocation, not a CPU run.
  • No silent fallback in the native path: load.rs:44-49 requires CudaExecutionProvider::initialized(index) to succeed or the load errors; load.rs:434 builds DecodeCudaState only when session.device_id().device_type == Cuda; mod.rs:275 dispatches to decode_cuda only when self.cuda.is_some(). The inputs_embeds decode branch is the one exercised: the fixture declares sequence_source: inputs_embeds (inference_metadata.yaml), so decode_cuda routes to decode_cuda_inputs_embeds (cuda.rs:329-341).
  • Advisory: the test proves GPU via device string + no-fallback semantics + memory, but does not programmatically assert session.device_id() == Cuda. A one-line EP assertion inside the cuda arm would make the on-GPU guarantee self-evident in CI. Not a blocker.

2. KV stays device-resident; only the one-token embedding crosses per step

  • run_cuda_eager_rows_owned (cuda.rs:262-314) binds only owned (the [1,1,hidden] embedding + optional [1,1] position_ids) via run_with_device_bindings against state.bindings[..base_binding_count] — the persistent device KV/mask bindings. KV present outputs stay device-bound: the guard at cuda.rs:301-306 bails if any bound output materialized to host, structurally enforcing that KV never round-trips. Per step: embedding + position_ids upload host->device, logits [1,vocab] read back; KV device-resident. Confirmed — the perf guarantee holds.

3. Eager-forward reuse is sound (with a perf caveat)

  • decode_cuda_inputs_embeds (cuda.rs:456-473) does ensure_capacity + extend_mask(if grew {0} else {past_len}, total_len, total_len) + set_logical_len(total_len) — identical growing-KV / mask semantics to the multi-token eager token path run_cuda_eager_rows (cuda.rs:362-376, 405-412). No correctness gap between eager-forward decode and the captured decode.
  • PERF FLAG (advisory): the single-token token-id path uses the CUDA-graph-captured run_one_token (cuda.rs:378-402) for fast replay, but the inputs_embeds path runs an eager (uncaptured) forward EVERY step because the embedding is a fresh host upload and cannot be a captured device-write target (acknowledged at cuda.rs:457-459). For the 35B-A3B target this is a real per-step cost (no graph capture / re-bind each step). Correct, but worth tracking as a follow-up perf item on Large-model (27B) native E2E offload blocked: Unsqueeze rank bug + missing recurrent_state + mobius#432 #384.

4. The unconsumed-mask / ReduceSum fixture caveat — legitimate, but a coverage gap to flag

  • Judgment: LEGITIMATE tiny-fixture detail, NOT masking a real-model bug. The tiny decoder has no real softmax attention (only MatMul + Concat-KV); attention_mask and position_ids are declared as genuine graph inputs but referenced by no node (build_tiny_gemma4_vlm_cuda.py:63-66, 103-113). cuDNN rejecting a degenerate all-axes ReduceSum-to-scalar is an artifact of forcing mask consumption in a trivial graph; a real model consumes the mask additively inside attention, not via a scalar reduce, so this is not a real-model cuDNN blocker.
  • BUT this means the parity test does NOT exercise a CONSUMED attention mask on the native CUDA decode path. Real-model attention-mask numerics on the CUDA decoder (35B-A3B / 27B) remain UNPROVEN by this fixture. Please flag on Large-model (27B) native E2E offload blocked: Unsqueeze rank bug + missing recurrent_state + mobius#432 #384 as a known coverage gap: a real-model or larger-fixture CUDA validation with a consumed mask is still owed before the GPU native-decode milestone is numerically signed off. Not a blocker for this plumbing slice.
  • Minor nit: the build-script docstring (lines 9-11) says the inputs are "wired into a zero contribution" but the code leaves them fully unconsumed — stale wording; align the comment with the actual construction.

5. Default ORT + CPU path unchanged; env-gated; cfg-correct

  • build_native_pipeline_decoder now resolves the device via native_decoder_device() (pipeline/mod.rs), which returns Cpu when ONNX_GENAI_PIPELINE_NATIVE_DECODER_DEVICE is unset. The CUDA path is strictly opt-in. Goldens unchanged:
    multimodal_reuse_e2e 14; optional_modality_pipeline_e2e 8; vlm_multibinding_pipeline_e2e 2; pipeline_executor 1; native_pipeline_decoder_parity (inc1/inc2b) 1 — all pass.
  • Clippy clean across default / native-backend / cuda / cuda,native-backend, including cargo clippy -p onnx-genai-engine --features cuda,native-backend --all-targets -- -D warnings (EXIT=0). cargo fmt --all --check clean.
  • Minor robustness nit (advisory): native_decoder_device() maps values like "cudabad" or "cuda:abc" to Cuda{index:None} (device 0), whereas the doc-comment says unrecognized values fall back to CPU. Only strings starting exactly with "cuda" reach CUDA, so risk is low; tightening the parse to reject non-empty non-numeric suffixes would match the documented intent.

6. Verification evidence

  • cargo test -p onnx-genai-engine (default): the only failures are gemma4_assistant_full 1, gemma4_assistant_metadata_smoke 1, gemma4_assistant_mixed 2 (token_input port ambiguity — untouched by this PR).
  • cargo test -p onnx-genai-engine --features native-backend: failing binaries native_engine (1 passed / 7 failed), native_speculative_driver (1/2), gemma4_assistant_full (0/1), gemma4_assistant_metadata_smoke (0/1), gemma4_assistant_mixed (1/2).
  • Baseline on origin/main (fresh worktree, same features): IDENTICAL failing set and counts — native_engine 7, native_speculative_driver 2, gemma4_assistant_full 1, gemma4_assistant_metadata_smoke 1, gemma4_assistant_mixed 2. This PR introduces ZERO new failures. (Note: the pre-existing set is slightly broader than the three named in the task, but every one reproduces identically on main.)
  • CUDA parity test GREEN on device 5 (see First milestone #1). cargo fmt --all --check: clean.

Verdict

APPROVE. On-GPU execution is proven, KV stays device-resident (only the one-token embedding uploads per step), the eager-forward reuse is semantically identical to the captured path, the default/ORT/CPU paths are untouched, goldens are unchanged, clippy/fmt are clean, and no new test failures are introduced. The two advisory items for #384 — (a) eager-per-step vs graph-capture perf for the inputs_embeds decode, and (b) the coverage gap that no CONSUMED attention mask is exercised on the native CUDA path — should be tracked but do not block this Inc3a plumbing slice. Inc3b (generic Routed ports on CUDA) remains correctly refused.

@justinchuby
justinchuby marked this pull request as ready for review July 30, 2026 19:17
@justinchuby
justinchuby merged commit 7ba70cb into main Jul 30, 2026
14 checks passed
@justinchuby
justinchuby deleted the squad/native-multi-component-pipeline-inc3 branch July 30, 2026 19:17
justinchuby added a commit that referenced this pull request Jul 30, 2026
…ve multi-component inc3b)

Lift the remaining CUDA refusal deferred in Inc3a: the native CUDA decoder now
binds arbitrary declared non-KV `Routed` step-input ports on-device per step, so
cross-component handoffs (a routed hidden/state edge, and eventually
static_cross_kv) work on the CUDA EP — generalizing the Inc3a inputs_embeds path
rather than forking it.

The CPU path already builds its owned per-step input set generically
(`prepare_cpu_step_inputs`). Inc3b gives the CUDA eager path the same treatment:
a new `prepare_cuda_owned_step_inputs` iterates every declared step input,
generating token/position ids and pulling inputs_embeds/routed tensors from the
supplied set by exact graph-port name, with the one CUDA-specific exclusion that
`attention_mask` is a persistent device binding (filled by extend_mask), never an
owned upload. Routed ports are owned per-step uploads — no new persistent device
binding and no DecodeCudaState binding-table change. Only the small per-step
tensors (embedding + routed state) cross host->device; the mask and KV cache stay
device-resident on the GPU.

- native_decode/cuda.rs: replace the embeds-only `decode_cuda_inputs_embeds` with
  a generic `decode_cuda_eager_step_inputs` + `prepare_cuda_owned_step_inputs`
  (mirrors the CPU contract); `decode_cuda` takes the eager path whenever any
  inputs_embeds/routed port is declared (`has_eager_step_inputs`), keeping the
  pure token-id captured fast path byte-identical.
- native_decode/load.rs: remove the CUDA `Routed` refusal (routed ports are now
  bound generically); inputs_embeds metadata resolution from Inc3a unchanged.
- scripts/build_tiny_gemma4_vlm_cuda_routed.py + tests/fixtures/
  tiny-gemma4-vlm-cuda-routed: fixture where the every_step embedding emits a
  second `router_state` output routed to a decoder `router_state` port, consumed
  via a real MatMul-by-zero (closed-form tokens `[0,5,6,7]` unchanged).
- tests/native_cuda_routed_pipeline_decoder_parity.rs: native decoder CPU vs CUDA
  EP (device 4) through the pipeline both produce `[0,5,6,7]`, proving the routed
  port binds on-device with the KV kept resident.

Scope out (deferred): vision cross-KV (needs the vision Attention float-mask
fixes) and the static_cross_kv upload-once optimization.

Refs #384. Stacks on #485 (native CUDA inputs_embeds decoder).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
…ve multi-component inc3b) (#487)

## Inc3b — generic routed ports on the native CUDA decoder

Lifts the remaining CUDA refusal deferred in Inc3a. The native CUDA
decoder now binds arbitrary declared non-KV **`Routed`** step-input
ports on-device per step (beyond just `inputs_embeds`), enabling
cross-component handoffs (a routed hidden/state edge, and eventually
`static_cross_kv`) on the CUDA EP while keeping the KV cache
device-resident.

### What changed
- **`native_decode/cuda.rs`**: replaced the embeds-only
`decode_cuda_inputs_embeds` with a generic
`decode_cuda_eager_step_inputs` + `prepare_cuda_owned_step_inputs`,
mirroring the CPU `prepare_cpu_step_inputs` contract. `decode_cuda`
takes the eager (uncaptured) path whenever any inputs_embeds/routed port
is declared (`has_eager_step_inputs`), keeping the pure token-id
**captured** fast path byte-identical. The one CUDA-specific exclusion:
`attention_mask` is a persistent device binding (filled by
`extend_mask`), never an owned upload — so routed ports need **no new
persistent binding** and **no `DecodeCudaState` change**; only the small
per-step tensors cross host→device.
- **`native_decode/load.rs`**: removed the CUDA `Routed` refusal; Inc3a
inputs_embeds metadata resolution unchanged.
- **Fixture + test**: `tiny-gemma4-vlm-cuda-routed` — every_step
embedding emits a second `router_state` output routed to a decoder
`router_state` port consumed via a real MatMul-by-zero (closed-form
tokens `[0,5,6,7]` unchanged).
`native_cuda_routed_pipeline_decoder_parity.rs` proves native decoder
**CPU vs CUDA EP (device 4)** through the pipeline both produce
`[0,5,6,7]`.

### Proof
`native_cuda_routed_pipeline_decoder_matches_cpu_token_ids` — GREEN on
device 4 (native-CPU == native-CUDA == `[0,5,6,7]`). ORT-default goldens
unchanged; full native-backend failing set identical to base (12
pre-existing, zero new). fmt clean; clippy clean ×4 (default /
native-backend / cuda / cuda,native-backend).

### Scope out (deferred)
Vision cross-KV (needs the vision Attention float-mask fixes, separate
#384 blocker) and the `static_cross_kv` upload-once optimization.

Stacks on #485 (native CUDA inputs_embeds decoder, now merged). Refs
#384. Do not merge yet.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
…e/hybrid wave logs (#532)

Scribe round 5. Records the native multi-component CUDA pipeline decode
wave (#484/#485/#486/#487/#525) and distills decisions.md 28520→19858
bytes (under the 20480 gate). State-only; no production code.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 31, 2026
…s native multi-component decode, #82/#384/35B-A3B) (#546)

## Summary

The task was to introduce a backend-neutral ownership seam so the
pipeline decode
loop can drive **either** ORT **or** native component sessions per step,
then route
the existing ORT path through it byte-identically (increment-1).

**On inspection, that seam already fully exists on `origin/main`** — it
landed
across the inc1→inc3c chain (#450, #478, #479, #485, #487, #533) and was
hardened by
#543. `PipelineDecodeLoopBackend` owns **no** ORT
`Session`/decode-state; it holds
only `Box<dyn PipelineDecoderComponent>` (stateful decoder seam) and
`Vec<(_, Box<dyn ComponentSession>)>` (stateless every_step seam), and
both ORT and
native backends are driven through one decode loop via runtime env
selection.

So increment-1 here is the piece the chain had **not** locked: a parity
test proving
the seam's **keystone** end-state — *every declared component running
natively at
once* (native every_step embedding **+** native device-KV decoder in the
same loop),
the exact shape a large multi-component package (up to the 35B-A3B
3-component
package) decodes through.

## What changed (test-only, zero production change)

- **`crates/onnx-genai-engine/tests/native_full_pipeline_parity.rs`** —
drives the
  `tiny-gemma4-vlm` composite with both
  `ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding` **and**
`ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder`, asserting the fully-native
run is
  token-identical to the ORT baseline `[0, 5, 6, 7]`.
- **`crates/onnx-genai-engine/Cargo.toml`** — registers the test
  (`required-features = ["native-backend"]`, CPU-only).
- **`.squad/decisions/inbox/mary-pipeline-native-ownership.md`** — full
assessment
  (ownership map), design affirmation, and the deferred next increment.

Prior increments proved each slice in isolation: inc1 (native embedding
+ ORT
decoder), inc2b (ORT embedding + native decoder). Nothing exercised
**both** natively
at once until now.

## ORT byte-identical proof

No production source is touched, so the ORT decode path is
byte-identical to
`origin/main` by construction. Empirically the ORT baseline `[0,5,6,7]`
and the
fully-native run `[0,5,6,7]` match exactly.

## Tests

- `native_full_pipeline_parity` — **pass** (new)
- `native_step_component_parity`, `native_pipeline_decoder_parity` —
**pass**
- 343 engine lib unit tests — **pass**, 1 ignored
- `cargo fmt --all --check` — clean

CUDA-gated native tests were not run in this CPU environment (unchanged
by this PR).

## Deferred to the next increment (native wiring completion)

The one genuine remaining hard limitation the code itself flags
(`decoder_component.rs:244-260`):
`NativePipelineDecoder::mirror_last_present_kv`
bails — the native decoder keeps KV session-resident and does not expose
host present
tensors, so native selection runs the non-paged, fresh-decode path with
no
cross-request KV reuse. Wiring native present-KV exposure + paged
mirroring is higher
blast radius and is intentionally **not** bundled here.

Refs #82, #384. Working as Mary (native-decode / pipeline engineer).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant