Skip to content

feat(pipeline): drive the decoder via a stateful PipelineDecoderComponent trait (native multi-component inc2a) - #478

Merged
justinchuby merged 2 commits into
mainfrom
squad/native-multi-component-pipeline-inc2
Jul 30, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/native-multi-component-pipeline-inc2

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What & why

Increment 2a of the native multi-component pipeline work (refs #384). Inc1 (#450, merged) routed the stateless every_step components through the backend-neutral ComponentSession trait. The decoder cannot reuse that seam: it is stateful — its KV cache grows across steps and, for the native backend, lives device-resident — so a stateless host round-trip would drop KV continuity and re-stage the whole cache every step, destroying decode throughput.

This PR introduces the stateful counterpart and lands it as a pure, provably-behavior-identical refactor (no native decoder yet), de-risking the native wiring (Inc2b).

The stateful decoder seam

  • New trait PipelineDecoderComponent (pipeline/decoder_component.rs): the flat autoregressive decode loop calls step() once per token; the impl retains its own per-step outputs, so the loop reads next_token_logits() / mirror_last_present_kv() without ever handling a concrete tensor type. Same DRY principle as Inc1, but stateful (KV stays inside the backend).
  • OrtPipelineDecoder is the ONNX Runtime impl, behaviour-identical to the previous inline run_decode_step_with_extra → mirror → extract path.
  • PipelineDecodeLoopBackend now holds decoder: Box<dyn PipelineDecoderComponent> instead of a concrete &Session + &mut DecodeState; next_logits() drives it entirely through trait methods.
  • Removed the now-redundant one-line extract_next_token_logits_with_io wrapper in favour of the slice-based extract_next_token_logits_from_outputs (no per-step clone of retained KV outputs); test callers updated.

Inc2a / Inc2b split

  • Inc2a (this PR): stateful trait + ORT adapter, loop calls .step(), ORT token output UNCHANGED. Pure refactor.
  • Inc2b (follow-up): native decoder impl wrapping NativeDecodeSession keeping device-resident KV across steps + native-decoder-in-pipeline token parity. Requires extending NativeDecodeSession::step to accept routed/inputs_embeds per-step inputs and expose present-KV — substantial, its own increment.

Design note: .squad/decisions/inbox/mary-pipeline-inc2-design.md (committed 6209729).

Proof the ORT path is unchanged

  • New equivalence unit test pipeline::decoder_component::tests::ort_decoder_component_matches_inline_step_path: asserts the trait wrapper's logits equal the inline helper path bit-for-bit across a 3-token prefill + two decode steps on tiny-multiaxis-state-decoder.
  • Flat-AR e2e token-id goldens unchanged: gemma4_vlm (embedding every_step + decoder → [0,5,6,7]), vlm_multibinding (2), multimodal_reuse (14), optional_modality (8), pipeline_executor (1).
  • Inc1 native every_step parity still green.

Verification

  • cargo test -p onnx-genai-engine --lib: 283 passed, 0 failed, 1 ignored.
  • Flat-AR e2e: gemma4_vlm 1, vlm_multibinding 2, multimodal_reuse 14, optional_modality 8, pipeline_executor 1 — all pass.
  • native_step_component_parity (native-backend): 1 passed.
  • cargo fmt --all --check: clean.
  • clippy clean (no warnings) for default / native-backend / cuda / cuda,native-backend (cfg-correct imports).

⚠️ Draft — do not merge. Stops at the largest PROVEN slice (Inc2a); Inc2b native decoder is deferred.

justinchuby and others added 2 commits July 30, 2026 14:58
…#384)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…nent trait (native multi-component inc2a)

Inc1 (#450) routed the stateless every_step components through the
backend-neutral ComponentSession trait. The decoder cannot reuse that
seam: it is stateful — its KV cache grows across steps and, for the
native backend, lives device-resident — so a stateless host round-trip
would drop KV continuity and re-stage the whole cache every step.

Inc2a introduces the stateful counterpart, PipelineDecoderComponent: the
flat autoregressive decode loop calls step() once per token and the
implementation retains its own per-step outputs, so the loop never
touches a concrete tensor type. OrtPipelineDecoder is the ONNX Runtime
implementation, behaviour-identical to the previous inline
run_decode_step_with_extra / mirror / extract path. A native decoder
keeping KV device-resident is the follow-up (Inc2b); see
.squad/decisions/inbox/mary-pipeline-inc2-design.md.

This is a pure refactor: no native decoder yet, ORT token output
unchanged. PipelineDecodeLoopBackend now holds
`decoder: Box<dyn PipelineDecoderComponent>` instead of a concrete
`&Session` + `&mut DecodeState`, and next_logits() drives it entirely
through trait methods. The now-redundant one-line
extract_next_token_logits_with_io wrapper is removed in favour of the
slice-based extract_next_token_logits_from_outputs (no per-step clone).

Proof the ORT path is unchanged:
- new unit test pipeline::decoder_component::tests::
  ort_decoder_component_matches_inline_step_path asserts the trait
  wrapper's logits equal the inline helper path bit-for-bit across a
  3-token prefill + two decode steps on tiny-multiaxis-state-decoder.
- flat-AR e2e token-id goldens unchanged: gemma4_vlm (embedding
  every_step + decoder -> [0,5,6,7]), vlm_multibinding (2),
  multimodal_reuse (14), optional_modality (8), pipeline_executor (1).
- Inc1 native every_step parity still green.

Refs #384.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Jul 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.52%. Comparing base (6e4257e) to head (2b01b12).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #478      +/-   ##
==========================================
+ Coverage   80.56%   81.52%   +0.95%     
==========================================
  Files         314      314              
  Lines      122637   122637              
  Branches   122637   122637              
==========================================
+ Hits        98804    99975    +1171     
+ Misses      19814    18637    -1177     
- Partials     4019     4025       +6     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.67% <ø> (ø)
mlas 78.72% <ø> (+0.81%) ⬆️
offline 81.48% <ø> (+0.99%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 7 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review of PR #478 — Inc2a stateful PipelineDecoderComponent seam (commit 2b01b12, refs #384). Reviewer: Melina (author Mary locked out). CPU-only worktree at origin/squad/native-multi-component-pipeline-inc2.

This is a genuine pure refactor. Verified with evidence against origin/main.

1. Behavior-identical guarantee — REAL, not tautological

The equivalence test ort_decoder_component_matches_inline_step_path (decoder_component.rs:150-231) builds a golden via the inline helpers over an independent inline_state (run_decode_step_with_extra + extract_next_token_logits_from_outputs, lines ~197-215), then drives a separate wrapped_state through OrtPipelineDecoder::step()/next_token_logits() and asserts assert_eq!(&got, expected) — exact Vec<f32> equality (bit-for-bit, not approx) over a real 3-token prefill + two decode steps (steps = [(vec![1,2,3],0),(vec![6],3),(vec![15],4)]). Because the two paths thread distinct DecodeState objects, it exercises the seam's state/extras forwarding and output retention rather than comparing the trait path to itself. Passes.

  • Goldens UNCHANGED: git diff origin/main...HEAD --stat touches only 9 files, none under tests/. gemma4_vlm [0,5,6,7] confirmed at gemma4_vlm_pipeline_e2e.rs:65; vlm_multibinding / native_step_component_parity golden files untouched by the diff.

2. Stateful seam correctness — CONFIRMED

  • The loop no longer touches DecodeState: every decoder_state.* in paged_decode.rs is now a trait call (use_kv(), retained_kv_len(), sliding_window(), sink_tokens(), mirror_last_present_kv()); DecodeState is owned inside OrtPipelineDecoder (state: &'a mut DecodeState, decoder_component.rs:70). flat_autoregressive.rs:229 retains_kv = backend.decoder.use_kv().
  • step(input_tokens, past_len, extras) carries exactly the old per-step inputs. decoder_extras() (paged_decode.rs:183) is unchanged — still threads decoder_in_edges (routed every_step/inputs_embeds/positions) and static_cross_kv; no dropped routed/cross inputs.
  • Call ordering preserved: step → set kv_len → mirror → sliding-window → logits (paged_decode.rs:235-302), identical to the old inline sequence; retained_kv_len(past_len) still applied before mirroring.
  • No aliasing change: the Box<dyn ...> holds the single &mut decoder_state borrow for the loop and is explicitly drop(backend)-ed (flat_autoregressive.rs) before downstream reuse. Reset/rewind paths (flat_autoregressive.rs:438-491) are outside the loop and untouched.

3. Removed wrapper / slice extraction — NO off-by-one

extract_next_token_logits_with_io (which merely took Vec<Value> by value and forwarded to extract_next_token_logits_from_outputs(&outputs, ...)) was deleted; callers now call the slice-based fn directly. The extractor's body is unchanged (logits.rs diff shows only pub(super)→pub(crate) visibility) — the last-token/prefill-vs-decode seq-dim indexing is untouched, so the "no per-step clone" change is ownership-only (Vec→slice), not an indexing change. All callers updated: state.rs tests, kv_bridge.rs tests, mod.rs re-export — and those tests still assert the same golden argmax/next_positions and pass.

4. Verify

  • cargo test -p onnx-genai-engine --lib: 283 passed / 0 failed / 1 ignored. The ignored test is pre-existing and env-gated: fim_generation_runs_with_fim_capable_model (requires ONNX_GENAI_FIM_MODEL_DIR), not touched by this PR.
  • cargo fmt --all --check: clean.
  • cargo clippy -p onnx-genai-engine --lib across [], native-backend, cuda, cuda,native-backend: all Finished, zero warnings (cfg-correct imports).
  • The pre-existing onnx-genai-ort loader / native gemma-assistant failures are outside the engine lib and unrelated to this refactor; the engine lib suite is fully green.

Clean stateful seam, no functional change, all evidence green. Approving.

@justinchuby
justinchuby marked this pull request as ready for review July 30, 2026 16:05
@github-actions

Copy link
Copy Markdown

✅ Benchmarks — No Regression

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
✅ kv_cache/alloc_dealloc_pages 34.09 µs 35.17 µs +3.2%
✅ matmul/small_generic_f16_threads=8/1x256x256 28.12 µs 28.65 µs +1.9%
✅ gather/small_bf16_threads=1-internal/4096 468.6 ns 473.9 ns +1.1%
✅ tokenization/encode_tokens_per_second 345.36 µs 349.10 µs +1.1%
✅ matmul/medium_generic_f32_threads=8/32x512x512 878.13 µs 886.60 µs +1.0%
✅ matmul/small_generic_bf16_threads=8/1x256x256 28.83 µs 29.08 µs +0.9%
✅ logit_processing/seven_processor_chain_per_step 1.08 ms 1.09 ms +0.8%
✅ matmul/small_generic_f32_threads=1/1x256x256 33.59 µs 33.83 µs +0.7%
✅ sampling_latency/greedy_per_token 2.98 µs 3.00 µs +0.7%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.13 ms 2.15 ms +0.6%
✅ grammar_masking/llguidance_compute_mask/32 68.05 µs 68.46 µs +0.6%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.81 ms 1.82 ms +0.6%
✅ matmul/small_generic_bf16_threads=1/1x256x256 27.79 µs 27.88 µs +0.3%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 77.61 µs 77.65 µs +0.1%
✅ matmul/medium_generic_f16_threads=1/32x512x512 27.32 µs 27.33 µs +0.0%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 72.84 µs 72.84 µs +0.0%
✅ matmul/small_generic_f16_threads=1/1x256x256 28.02 µs 28.00 µs -0.1%
✅ tokenization/decode_tokens_per_second 5.64 ms 5.62 ms -0.3%
✅ sampling_latency/top_p_per_token 913.58 µs 906.71 µs -0.8%
✅ sampling_latency/min_p_per_token 316.05 µs 313.52 µs -0.8%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 8.62 ms 8.55 ms -0.8%
✅ matmul/medium_generic_f16_threads=8/32x512x512 27.75 µs 27.50 µs -0.9%
✅ gather/medium_f32_threads=1-internal/32768 3.42 µs 3.37 µs -1.5%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 489.78 µs 481.68 µs -1.7%
✅ sampling_latency/top_k_per_token 447.25 µs 437.67 µs -2.1%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.61 ms 3.53 ms -2.2%
✅ add/medium_f16_threads=1-internal/262144 2.44 ms 2.38 ms -2.3%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.33 µs 14.85 µs -3.1%
✅ gather/medium_bf16_threads=1-internal/32768 2.23 µs 2.16 µs -3.3%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.27 ms 1.23 ms -3.3%
✅ add/medium_bf16_threads=1-internal/262144 2.52 ms 2.43 ms -3.7%
✅ matmul/small_generic_f32_threads=8/1x256x256 33.12 µs 31.69 µs -4.3%
✅ add/small_f16_threads=1-internal/1024 12.38 µs 11.84 µs -4.4%
✅ gather/small_f32_threads=1-internal/4096 660.9 ns 627.3 ns -5.1%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 377.74 µs 354.59 µs -6.1%
✅ gather/large_f32_threads=1-internal/131072 24.96 µs 23.27 µs -6.8%
✅ add/large_bf16_threads=1-internal/4194304 41.74 ms 38.89 ms -6.8%
✅ gather/large_bf16_threads=1-internal/131072 11.52 µs 10.69 µs -7.2%
✅ add/large_f16_threads=1-internal/4194304 40.89 ms 37.92 ms -7.3%
✅ add/large_f32_threads=1-internal/4194304 40.74 ms 37.53 ms -7.9%
✅ add/small_bf16_threads=1-internal/1024 13.01 µs 11.90 µs -8.5%
✅ add/medium_f32_threads=1-internal/262144 2.58 ms 2.35 ms -9.0%
✅ gather/small_f16_threads=1-internal/4096 487.5 ns 442.1 ns -9.3%
✅ add/small_f32_threads=1-internal/1024 204.6 ns 181.1 ns -11.5%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.04 ms 906.43 µs -12.9%
✅ gather/medium_f16_threads=1-internal/32768 2.55 µs 2.22 µs -13.0%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 269.86 µs 226.46 µs -16.1%
🟢 gather/large_f16_threads=1-internal/131072 14.31 µs 11.15 µs -22.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 2.90 3.44 4.89 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby
justinchuby merged commit f02a258 into main Jul 30, 2026
14 checks passed
@justinchuby
justinchuby deleted the squad/native-multi-component-pipeline-inc2 branch July 30, 2026 16:06
justinchuby added a commit that referenced this pull request Jul 30, 2026
… (native multi-component inc2b)

Inc2a (#478) introduced the stateful PipelineDecoderComponent trait and
drove the ORT decoder through it. Inc2b adds the NATIVE counterpart:
NativePipelineDecoder wraps NativeDecodeSession and keeps its KV cache
session-resident across every step() call, so the expensive KV state
never round-trips through the host pipeline pool. The same flat
autoregressive decode loop drives either backend through the trait with
no forked code path (the DRY principle of inc1/inc2a, now stateful).

Per-step seam: each step the every_step embedding component publishes
inputs_embeds into the host pool as an ort::Value; the native decoder
receives it as one small routed input (one token's embedding, converted
value -> ComponentTensor -> native Tensor, reusing the inc1 value seam),
while the KV cache stays inside the native session.

Native selection is gated behind ONNX_GENAI_PIPELINE_NATIVE_DECODER
(names the decoder component or a truthy token), mirroring inc1's
ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS. Unset keeps the ORT decoder
(default, unchanged). Requesting it without --features native-backend is
a clear error. The native decoder runs the non-paged, fresh-decode path
(paged_enabled=false, reused=0): its KV is session-resident and not
exposed as host present tensors, so paged present-KV mirroring and
cross-request reuse are deferred to inc3 — this changes cross-request KV
reuse only, never the tokens produced within a generation.

What changed:
- pipeline/decoder_component.rs: NativePipelineDecoder impl of the trait
  (step converts extras + calls decode_with_step_inputs; next_token_logits
  returns the retained final row; mirror_last_present_kv is unsupported
  and never reached on the non-paged native path). cfg(native-backend).
- native_decode/load.rs: thread the pipeline ModelIoSpec through a new
  pub(crate) load_with_io so an inputs_embeds decoder (no token input)
  binds sequence source / KV pairs from metadata instead of guessing.
- native_component.rs: pub(crate) component_tensor_to_native_tensor seam.
- pipeline/mod.rs: native_decoder_selected() flag + build_native_pipeline_decoder().
- flat_autoregressive.rs: select native vs ORT decoder; force non-paged
  fresh path when native.

Split: Inc2b-i (NativeDecodeSession::decode_with_step_inputs + resident
KV) already existed in tree and is proven by the native_decode tests;
this PR is Inc2b-ii (the adapter + wiring + parity proof). See
.squad/decisions/inbox/mary-pipeline-inc2b-design.md.

Token-parity proof (native vs ORT, exact ids):
- new test native_pipeline_decoder_parity::native_pipeline_decoder_matches_ort_token_ids:
  ORT decoder -> [0,5,6,7]; native device-KV decoder -> [0,5,6,7] (identical),
  on tiny-gemma4-vlm (embedding every_step + inputs_embeds decoder).
- ORT-path goldens unchanged: gemma4_vlm, vlm_multibinding (2),
  multimodal_reuse (14), optional_modality (8), pipeline_executor (1).
- inc1 native every_step parity still green.
- lib: 283 passed (default), 333 passed (native-backend).
- clippy clean: default / native-backend / cuda / cuda,native-backend.

Refs #384.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
… (native multi-component inc2b) (#479)

## What & why

Increment 2b of the native multi-component pipeline work (refs #384).
Inc2a (#478, now merged) introduced the stateful
`PipelineDecoderComponent` trait and drove the **ORT** decoder through
it. **Inc2b adds the native counterpart**: a `NativePipelineDecoder`
that wraps `NativeDecodeSession` and keeps its KV cache
**session-resident across every `step()` call**, so the expensive KV
state never round-trips through the host pipeline pool. The same flat
autoregressive decode loop drives either backend through the trait with
**no forked code path** — the DRY principle of inc1/inc2a, now stateful.

> 📌 Originally stacked on #478; **#478 has merged**, so this branch was
rebased onto `origin/main` and now targets `main` directly (only the two
Inc2b commits remain).

## The per-step seam (why this is the right design)

Each step the every_step embedding component publishes `inputs_embeds`
into the host pool as an `ort::Value` (routed edge
`embedding.inputs_embeds → decoder.inputs_embeds`). The native decoder
receives it as **one small routed input** — one token's embedding,
`[1,1,hidden]`, converted `value → ComponentTensor → native Tensor`
(reusing the inc1 value-type seam). **The KV cache stays inside the
native session** and is never uploaded/downloaded from the pipeline
pool. That is exactly why the decoder needs a *stateful* seam and cannot
reuse the inc1 stateless `ComponentSession` round-trip.

## What already existed vs. what this PR adds

- **Inc2b-i (already in tree):**
`NativeDecodeSession::decode_with_step_inputs` already accepts routed /
`inputs_embeds` per-step inputs and owns KV across steps
(`NativeStepInputSource::InputsEmbeds`), proven by the `native_decode`
tests. No new kernel work.
- **Inc2b-ii (THIS PR):** the `NativePipelineDecoder` adapter + flat-AR
wiring + env selection + **token-parity proof**.

## Changes

- `pipeline/decoder_component.rs`: `NativePipelineDecoder` impl of the
trait (`step` converts extras + calls `decode_with_step_inputs`;
`next_token_logits` returns the retained final logits row;
`mirror_last_present_kv` is unsupported and never reached on the
non-paged native path). `cfg(native-backend)`.
- `native_decode/load.rs`: thread the pipeline `ModelIoSpec` through a
new `pub(crate) load_with_io`, so an `inputs_embeds` decoder (no token
input) binds its sequence source / KV pairs from **metadata** instead of
guessing from tensor shapes.
- `native_component.rs`: `pub(crate) component_tensor_to_native_tensor`
conversion seam.
- `pipeline/mod.rs`: `native_decoder_selected()` env flag +
`build_native_pipeline_decoder()`.
- `flat_autoregressive.rs`: select native vs ORT decoder; force the
non-paged, fresh-decode path when native.

## Selection & paging

- Flag: `ONNX_GENAI_PIPELINE_NATIVE_DECODER` (names the decoder
component or a truthy token), mirroring inc1's
`ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS`. Unset ⇒ ORT decoder
(default, unchanged). Requesting it without `--features native-backend`
is a clear error.
- **Paging deferred to inc3:** native selection runs the non-paged path
(`paged_enabled=false`, `reused=0`) because the native KV is
session-resident and not exposed as host present tensors. Paging is a
*cross-request KV-reuse cache* and does **not** change the tokens
produced within a generation, so ORT (paged) vs native (non-paged) token
IDs still match. Native present-KV exposure + paged mirroring + vision
cross-KV are inc3.

## Token-parity proof (native vs ORT, exact IDs)

- **New test**
`native_pipeline_decoder_parity::native_pipeline_decoder_matches_ort_token_ids`:
ORT decoder → `[0,5,6,7]`; native device-KV decoder → `[0,5,6,7]`
(**identical**), on `tiny-gemma4-vlm` (embedding every_step +
`inputs_embeds` decoder).
- ORT-path goldens **unchanged**: gemma4_vlm (1), vlm_multibinding (2),
multimodal_reuse (14), optional_modality (8), pipeline_executor (1).
- inc1 native every_step parity still green.

## Verification

- `cargo test -p onnx-genai-engine --lib`: **283 passed** (default),
**333 passed** (native-backend), 0 failed.
- `cargo fmt --all --check`: clean.
- clippy clean (no warnings): **default / native-backend / cuda /
cuda,native-backend** (cfg-correct imports; all native decoder code
behind `cfg(feature = "native-backend")`).

Design note: `.squad/decisions/inbox/mary-pipeline-inc2b-design.md`.

⚠️ Draft — do not merge. Stops at the largest PROVEN slice (Inc2b-ii,
CPU device-KV text decoder with green token parity). Inc3 (native
present-KV exposure + paged reuse + vision cross-KV + CUDA target)
deferred.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
…CUDA-hybrid merge logs (#483)

Scribe round 4 state consolidation (squad-internal, no code). Merges 4
design-note inbox drops into decisions/archive, distils 2 standing
directives, logs #477/#478/#479/#480 merges, appends histories.
decisions.md 25.4KB→28.5KB (older 07-29 entries flagged for next-round
distillation).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 31, 2026
…s native multi-component decode, #82/#384/35B-A3B) (#546)

## Summary

The task was to introduce a backend-neutral ownership seam so the
pipeline decode
loop can drive **either** ORT **or** native component sessions per step,
then route
the existing ORT path through it byte-identically (increment-1).

**On inspection, that seam already fully exists on `origin/main`** — it
landed
across the inc1→inc3c chain (#450, #478, #479, #485, #487, #533) and was
hardened by
#543. `PipelineDecodeLoopBackend` owns **no** ORT
`Session`/decode-state; it holds
only `Box<dyn PipelineDecoderComponent>` (stateful decoder seam) and
`Vec<(_, Box<dyn ComponentSession>)>` (stateless every_step seam), and
both ORT and
native backends are driven through one decode loop via runtime env
selection.

So increment-1 here is the piece the chain had **not** locked: a parity
test proving
the seam's **keystone** end-state — *every declared component running
natively at
once* (native every_step embedding **+** native device-KV decoder in the
same loop),
the exact shape a large multi-component package (up to the 35B-A3B
3-component
package) decodes through.

## What changed (test-only, zero production change)

- **`crates/onnx-genai-engine/tests/native_full_pipeline_parity.rs`** —
drives the
  `tiny-gemma4-vlm` composite with both
  `ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding` **and**
`ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder`, asserting the fully-native
run is
  token-identical to the ORT baseline `[0, 5, 6, 7]`.
- **`crates/onnx-genai-engine/Cargo.toml`** — registers the test
  (`required-features = ["native-backend"]`, CPU-only).
- **`.squad/decisions/inbox/mary-pipeline-native-ownership.md`** — full
assessment
  (ownership map), design affirmation, and the deferred next increment.

Prior increments proved each slice in isolation: inc1 (native embedding
+ ORT
decoder), inc2b (ORT embedding + native decoder). Nothing exercised
**both** natively
at once until now.

## ORT byte-identical proof

No production source is touched, so the ORT decode path is
byte-identical to
`origin/main` by construction. Empirically the ORT baseline `[0,5,6,7]`
and the
fully-native run `[0,5,6,7]` match exactly.

## Tests

- `native_full_pipeline_parity` — **pass** (new)
- `native_step_component_parity`, `native_pipeline_decoder_parity` —
**pass**
- 343 engine lib unit tests — **pass**, 1 ignored
- `cargo fmt --all --check` — clean

CUDA-gated native tests were not run in this CPU environment (unchanged
by this PR).

## Deferred to the next increment (native wiring completion)

The one genuine remaining hard limitation the code itself flags
(`decoder_component.rs:244-260`):
`NativePipelineDecoder::mirror_last_present_kv`
bails — the native decoder keeps KV session-resident and does not expose
host present
tensors, so native selection runs the non-paged, fresh-decode path with
no
cross-request KV reuse. Wiring native present-KV exposure + paged
mirroring is higher
blast radius and is intentionally **not** bundled here.

Refs #82, #384. Working as Mary (native-decode / pipeline engineer).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant