Skip to content

feat(pipeline): drive every_step components via ComponentSession trait (native multi-component inc1) - #450

Merged
justinchuby merged 2 commits into
mainfrom
squad/native-multi-component-pipeline-inc1
Jul 30, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/native-multi-component-pipeline-inc1

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What & why (refs #384)

Unblocks native decode of pipelined multi-component models (Qwen3.6-35B-A3B = embedding + decoder + vision, and any multimodal pipeline). The blocker: PipelineDecodeLoopBackend hardcoded the concrete ORT Session/Value, so it could only drive ORT components even though native (nxrt custom-EP) component sessions load fine.

This is increment 1 of the refactor scoped in .squad/decisions/inbox/mary-native-pipeline-plan.md (committed here).

The value-type seam (verdict)

The ComponentSession trait boundary is a backend-neutral, host-resident ComponentTensor (raw little-endian bytes), not ORT Value and not an nxrt tensor. The decode-loop pool holds ORT Value, so routing step components through the trait requires a pool Value ⇄ ComponentTensor conversion at the loop boundary — that host round-trip is the crux of the work.

Changes

  • PipelineDecodeLoopBackend.step_components → Vec<(StepComponentBinding, Box<dyn ComponentSession>)>.
  • run_step_components crosses the seam generically: pool Value → ComponentTensor → ComponentSession::run → ComponentTensor → pool Value. One code path, no forked native copy.
  • New OrtComponentSessionRef<'a>: a borrowing ORT ComponentSession adapter so the default path drives already-loaded sessions unchanged (behaviour-identical); shares its run body with the owning OrtComponentSession.
  • Native every_step components selected at runtime via ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS (empty/unset ⇒ all ORT, so the ORT decode path is untouched by default).

Validation

  • Native-vs-ORT token parity: native_step_component_parity::native_every_step_embedding_matches_ort_token_ids runs the tiny-gemma4-vlm composite pipeline with its embedding every_step component on ORT (baseline [0,5,6,7]) and on native nxrt (decoder stays ORT in both), asserting identical token ids. CPU-only, deterministic.
  • New unit test borrowing_ref_adapter_matches_owning_adapter (onnx-genai-ort).
  • All pipeline e2e/unit tests pass (gemma4-vlm, vlm-multibinding, optional-modality, multimodal-reuse, pipeline-executor, 67 pipeline lib tests).
  • cargo fmt --all --check clean; clippy clean for default, native-backend, cuda, and cuda,native-backend.

Scope / follow-ups (honest increment boundary)

The decoder itself (decoder: &Session, run_decode_step_with_extra, logits) and cross-component value handoff / static_cross_kv / device placement remain ORT-owned — those are inc2/inc3 in the plan doc. This PR proves the seam on the every_step slice only.

Note: a pre-existing, unrelated failure in onnx-genai-ort (loader::model_package_tests::flat_directory_with_unrelated_manifest_remains_backward_compatible) reproduces on origin/main with this branch's changes stashed — not introduced here.

Do not merge (draft).

justinchuby and others added 2 commits July 30, 2026 13:46
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…t (native multi-component inc1)

Route the pipeline decode loop's every_step (step) components through the
backend-neutral ComponentSession trait instead of the concrete ORT Session,
so the same run_step_components code path drives an ORT session or a native
nxrt component with no forked native copy.

- PipelineDecodeLoopBackend.step_components becomes
  Vec<(StepComponentBinding, Box<dyn ComponentSession>)>.
- run_step_components crosses the value-type seam: pool ORT Value ->
  neutral host ComponentTensor -> trait run -> ComponentTensor -> pool Value.
- Add OrtComponentSessionRef, a borrowing ORT ComponentSession adapter, so the
  default path drives already-loaded sessions unchanged (behaviour-identical).
- Select native every_step components at runtime via
  ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS (empty/unset => all ORT).
- Parity test: the tiny-gemma4-vlm embedding every_step component produces
  identical token ids [0,5,6,7] on ORT and native while the decoder stays ORT.

The decoder itself and cross-component value handoff remain ORT-owned; that is
inc2/inc3 (see .squad/decisions/inbox/mary-native-pipeline-plan.md). Refs #384.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Jul 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.56%. Comparing base (d233366) to head (72d7374).
⚠️ Report is 3 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #450      +/-   ##
==========================================
- Coverage   80.56%   80.56%   -0.01%     
==========================================
  Files         314      314              
  Lines      122637   122637              
  Branches   122637   122637              
==========================================
- Hits        98801    98799       -2     
- Misses      19817    19818       +1     
- Partials     4019     4020       +1     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.67% <ø> (-0.11%) ⬇️
mlas 77.91% <ø> (ø)
offline 80.48% <ø> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.42 ms 2.11 ms +49.1%
⚠️ tokenization/decode_tokens_per_second 6.56 ms 8.15 ms +24.1%
⚠️ tokenization/encode_tokens_per_second 401.47 µs 469.77 µs +17.0%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 95.51 µs 106.96 µs +12.0%
✅ sampling_latency/greedy_per_token 3.04 µs 3.33 µs +9.7%
✅ sampling_latency/top_k_per_token 466.60 µs 506.41 µs +8.5%
✅ sampling_latency/top_p_per_token 980.94 µs 1.00 ms +2.0%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.14 ms 2.16 ms +1.0%
✅ kv_cache/alloc_dealloc_pages 40.77 µs 40.35 µs -1.0%
✅ sampling_latency/min_p_per_token 371.88 µs 351.08 µs -5.6%
✅ matmul/small_generic_bf16_threads=8/1x256x256 40.07 µs 37.63 µs -6.1%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 553.12 µs 503.29 µs -9.0%
✅ grammar_masking/llguidance_compute_mask/32 82.31 µs 74.51 µs -9.5%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 100.13 µs 90.59 µs -9.5%
✅ add/small_bf16_threads=1-internal/1024 15.26 µs 13.55 µs -11.2%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 5.52 ms 4.85 ms -12.2%
✅ add/medium_f32_threads=1-internal/262144 2.89 ms 2.52 ms -12.7%
✅ gather/medium_bf16_threads=1-internal/32768 2.58 µs 2.23 µs -13.8%
🟢 add/medium_f16_threads=1-internal/262144 3.30 ms 2.73 ms -17.1%
🟢 matmul/medium_generic_bf16_threads=1/32x512x512 601.85 µs 491.37 µs -18.4%
🟢 matmul/small_generic_f16_threads=8/1x256x256 37.15 µs 30.06 µs -19.1%
🟢 matmul/large_generic_f32_threads=1/32x1024x1024 11.29 ms 9.04 ms -19.9%
🟢 gather/large_f32_threads=1-internal/131072 32.28 µs 25.75 µs -20.2%
🟢 matmul/medium_generic_f32_threads=1/32x512x512 2.71 ms 2.15 ms -20.9%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 36.61 µs 28.80 µs -21.3%
🟢 matmul/small_generic_f16_threads=1/1x256x256 36.26 µs 28.11 µs -22.5%
🟢 matmul/medium_generic_f16_threads=1/32x512x512 36.05 µs 27.73 µs -23.1%
🟢 logit_processing/seven_processor_chain_per_step 1.61 ms 1.21 ms -24.9%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 38.80 µs 29.02 µs -25.2%
🟢 add/small_f16_threads=1-internal/1024 17.67 µs 13.04 µs -26.2%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.27 ms 933.14 µs -26.3%
🟢 add/large_bf16_threads=1-internal/4194304 58.40 ms 42.26 ms -27.6%
🟢 add/small_f32_threads=1-internal/1024 290.7 ns 208.6 ns -28.2%
🟢 add/large_f32_threads=1-internal/4194304 56.91 ms 40.70 ms -28.5%
🟢 gather/small_f32_threads=1-internal/4096 839.2 ns 598.2 ns -28.7%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 316.73 µs 225.31 µs -28.9%
🟢 gather/small_f16_threads=1-internal/4096 610.5 ns 434.1 ns -28.9%
🟢 reduce_mean/small_f32_threads=1-internal/4096 19.58 µs 13.83 µs -29.4%
🟢 add/large_f16_threads=1-internal/4194304 61.72 ms 41.81 ms -32.3%
🟢 gather/small_bf16_threads=1-internal/4096 650.9 ns 435.3 ns -33.1%
🟢 gather/medium_f16_threads=1-internal/32768 3.21 µs 2.13 µs -33.6%
🟢 add/medium_bf16_threads=1-internal/262144 4.17 ms 2.73 ms -34.6%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.42 ms 925.57 µs -34.6%
🟢 gather/large_bf16_threads=1-internal/131072 17.48 µs 10.16 µs -41.9%
🟢 gather/medium_f32_threads=1-internal/32768 5.96 µs 3.41 µs -42.7%
🟢 gather/large_f16_threads=1-internal/131072 20.88 µs 10.86 µs -48.0%
🟢 matmul/small_generic_f32_threads=1/1x256x256 67.65 µs 34.15 µs -49.5%
🟢 matmul/small_generic_f32_threads=8/1x256x256 69.64 µs 32.31 µs -53.6%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.83 4.39 6.64 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review by Melina (senior reviewer). Author locked out; reviewed at commit 72d7374 in a clean worktree. Evidence below with file:line.

1. Seam correctness (the critical point) — LOSSLESS + unsupported-dtype-safe

  • Dtype mapping is complete and exhaustive, both directions. DataType to ComponentDataType and back are exhaustive 1:1 matches over all 14 variants (crates/onnx-genai-ort/src/component.rs:21-61). ORT's DataType enum has exactly those 14 variants, and both From impls are exhaustive match arms with no wildcard, so on the ORT-to-Component path there is NO reachable unsupported dtype and NO silent wrong-dtype coercion. The native side does surface it loudly: ir_dtype_to_component returns ComponentError::UnsupportedDataType (native_component.rs:44).
  • Byte layout is raw little-endian, shape preserved exactly. to_value/from_value carry raw bytes only (component.rs:107-131). value_to_component_tensor uses value.to_raw_bytes() + value.shape().to_vec(); component_tensor_to_value uses from_raw_bytes with tensor.shape() (pipeline/mod.rs:1381-1400). No re-interpretation.
  • No truncation/padding bug for non-8-bit dtypes. Value::from_raw_bytes rejects any length != numel * dtype.size_of() (value.rs:275-287); ComponentTensor::from_raw does the same via ByteLengthMismatch (metadata/src/component.rs:146-160). The two size_of tables are byte-for-byte identical (value.rs:29-39 vs component.rs:62-77), so i64 token ids (8B), fp16/bf16 (2B), and fp32 (4B) all round-trip without drift.
  • Unsupported dtype errors, never guesses. ComponentError::UnsupportedDataType exists and is the documented failure surface (metadata/src/component.rs:233-238).

Round-trip verdict: lossless and unsupported-dtype-safe.

2. Default ORT path — round-trips, proven behavior-identical

The env-unset path is NOT short-circuited: every_step ORT components are driven through OrtComponentSessionRef (pipeline/mod.rs:426-430), so each step does pool Value -> ComponentTensor -> Value -> Session::run -> Value -> ComponentTensor -> pool Value. Because the conversion is raw-byte + 1:1 dtype + exact shape (section 1), the output Value is bit-identical to the pre-PR direct Session::run. Confirmed empirically: gemma4_vlm_pipeline_e2e (exercises the ORT every_step embedding) passes 1/1, and the parity test baselines the ORT token ids to the fixture's closed-form [0,5,6,7]. Cost note: this adds one host copy per every_step component per step; the author documents it as negligible for the small embedding outputs (mod.rs:1374-1380) — acceptable.

3. DRY — one path, no native fork

step_components is a single Vec<(StepComponentBinding, Box)> driven by one run_step_components loop (paged_decode.rs:87-92, 131-171); there is no forked native copy. OrtComponentSessionRef genuinely shares its run body with the owning OrtComponentSession via the single run_ort_component fn (component.rs:132-215). The new unit test borrowing_ref_adapter_matches_owning_adapter asserts the two adapters produce identical bytes — passes.

4. Token-parity — reproduced, real, deterministic

native_step_component_parity::native_every_step_embedding_matches_ort_token_ids passes. It sets ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding inside the test, constructs the native session, greedy-decodes (temperature 0.0), asserts native == ort AND ort == [0,5,6,7] on tiny-gemma4-vlm. Real assertion on token ids, deterministic, and genuinely exercises the native path.

5. Scope honesty

Genuinely the every_step slice only: the decoder still runs as a concrete ORT Session (flat_autoregressive.rs:153-156); only step_components are boxed behind the trait. Native selection is opt-in per named component via the env var; without native-backend the native request bails with an actionable message (mod.rs:412-419). Inc2/inc3 (decoder KV-cache ownership, cross-component handoff) are correctly deferred and documented, not half-wired.

6. Tests + feature sets

  • cargo fmt --all --check: clean.
  • clippy default (engine/ort/metadata, --all-targets): clean. clippy --features native-backend (engine, --all-targets): clean.
  • onnx-genai-ort lib: 77 passed / 1 failed — the single failure is loader::model_package_tests::flat_directory_with_unrelated_manifest_remains_backward_compatible (tiny-llm-scatter missing model.onnx), which reproduces on origin/main and is NOT from this PR. The 3 new/relevant component tests (exposes_graph_io_metadata, borrowing_ref_adapter_matches_owning_adapter, named_tensor_run_round_trip_matches_session) all pass.
  • onnx-genai-engine --features native-backend --no-fail-fast: native_step_component_parity 1/1 pass; ORT pipeline e2e pass (gemma4_vlm_pipeline_e2e 1, iterative_pipeline_e2e 32, pipeline_executor 1; vlm/whisper e2e ignored — need real artifacts). The 12 named failures (gemma4_assistant x3, native_engine x7, native_speculative_driver x2) are an identical set on origin/main (diff of failing test names: identical) — pre-existing, caused by native-backend operator support / fixture limits in this environment, not this PR.

Conclusion

The value-type seam is lossless and fails loudly on any out-of-vocabulary dtype; the default ORT path is proven behavior-identical (extra host copy only); there is a single trait-parameterized decode path with no native fork; the token-parity test is a real, deterministic assertion that exercises the native path. No new test failures vs main. APPROVE.

@justinchuby
justinchuby marked this pull request as ready for review July 30, 2026 14:47
@justinchuby
justinchuby merged commit 6e4257e into main Jul 30, 2026
14 checks passed
@justinchuby
justinchuby deleted the squad/native-multi-component-pipeline-inc1 branch July 30, 2026 14:47
justinchuby added a commit that referenced this pull request Jul 30, 2026
…nent trait (native multi-component inc2a) (#478)

## What & why

Increment 2a of the native multi-component pipeline work (refs #384).
Inc1 (#450, merged) routed the **stateless** `every_step` components
through the backend-neutral `ComponentSession` trait. The **decoder**
cannot reuse that seam: it is **stateful** — its KV cache grows across
steps and, for the native backend, lives device-resident — so a
stateless host round-trip would drop KV continuity and re-stage the
whole cache every step, destroying decode throughput.

This PR introduces the **stateful** counterpart and lands it as a
**pure, provably-behavior-identical refactor** (no native decoder yet),
de-risking the native wiring (Inc2b).

## The stateful decoder seam

- New `trait PipelineDecoderComponent`
(`pipeline/decoder_component.rs`): the flat autoregressive decode loop
calls `step()` once per token; the impl **retains its own per-step
outputs**, so the loop reads `next_token_logits()` /
`mirror_last_present_kv()` without ever handling a concrete tensor type.
Same DRY principle as Inc1, but stateful (KV stays inside the backend).
- `OrtPipelineDecoder` is the ONNX Runtime impl, behaviour-identical to
the previous inline `run_decode_step_with_extra` → mirror → extract
path.
- `PipelineDecodeLoopBackend` now holds `decoder: Box<dyn
PipelineDecoderComponent>` instead of a concrete `&Session` + `&mut
DecodeState`; `next_logits()` drives it entirely through trait methods.
- Removed the now-redundant one-line `extract_next_token_logits_with_io`
wrapper in favour of the slice-based
`extract_next_token_logits_from_outputs` (no per-step clone of retained
KV outputs); test callers updated.

## Inc2a / Inc2b split

- **Inc2a (this PR):** stateful trait + ORT adapter, loop calls
`.step()`, ORT token output UNCHANGED. Pure refactor.
- **Inc2b (follow-up):** native decoder impl wrapping
`NativeDecodeSession` keeping device-resident KV across steps +
native-decoder-in-pipeline token parity. Requires extending
`NativeDecodeSession::step` to accept routed/`inputs_embeds` per-step
inputs and expose present-KV — substantial, its own increment.

Design note: `.squad/decisions/inbox/mary-pipeline-inc2-design.md`
(committed 6209729).

## Proof the ORT path is unchanged

- **New equivalence unit test**
`pipeline::decoder_component::tests::ort_decoder_component_matches_inline_step_path`:
asserts the trait wrapper's logits equal the inline helper path
**bit-for-bit** across a 3-token prefill + two decode steps on
`tiny-multiaxis-state-decoder`.
- Flat-AR e2e token-id goldens unchanged: `gemma4_vlm` (embedding
every_step + decoder → `[0,5,6,7]`), `vlm_multibinding` (2),
`multimodal_reuse` (14), `optional_modality` (8), `pipeline_executor`
(1).
- Inc1 native `every_step` parity still green.

## Verification

- `cargo test -p onnx-genai-engine --lib`: **283 passed, 0 failed, 1
ignored**.
- Flat-AR e2e: gemma4_vlm 1, vlm_multibinding 2, multimodal_reuse 14,
optional_modality 8, pipeline_executor 1 — all pass.
- `native_step_component_parity` (native-backend): 1 passed.
- `cargo fmt --all --check`: clean.
- clippy clean (no warnings) for **default / native-backend / cuda /
cuda,native-backend** (cfg-correct imports).

⚠️ Draft — do not merge. Stops at the largest PROVEN slice (Inc2a);
Inc2b native decoder is deferred.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 31, 2026
…s native multi-component decode, #82/#384/35B-A3B) (#546)

## Summary

The task was to introduce a backend-neutral ownership seam so the
pipeline decode
loop can drive **either** ORT **or** native component sessions per step,
then route
the existing ORT path through it byte-identically (increment-1).

**On inspection, that seam already fully exists on `origin/main`** — it
landed
across the inc1→inc3c chain (#450, #478, #479, #485, #487, #533) and was
hardened by
#543. `PipelineDecodeLoopBackend` owns **no** ORT
`Session`/decode-state; it holds
only `Box<dyn PipelineDecoderComponent>` (stateful decoder seam) and
`Vec<(_, Box<dyn ComponentSession>)>` (stateless every_step seam), and
both ORT and
native backends are driven through one decode loop via runtime env
selection.

So increment-1 here is the piece the chain had **not** locked: a parity
test proving
the seam's **keystone** end-state — *every declared component running
natively at
once* (native every_step embedding **+** native device-KV decoder in the
same loop),
the exact shape a large multi-component package (up to the 35B-A3B
3-component
package) decodes through.

## What changed (test-only, zero production change)

- **`crates/onnx-genai-engine/tests/native_full_pipeline_parity.rs`** —
drives the
  `tiny-gemma4-vlm` composite with both
  `ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding` **and**
`ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder`, asserting the fully-native
run is
  token-identical to the ORT baseline `[0, 5, 6, 7]`.
- **`crates/onnx-genai-engine/Cargo.toml`** — registers the test
  (`required-features = ["native-backend"]`, CPU-only).
- **`.squad/decisions/inbox/mary-pipeline-native-ownership.md`** — full
assessment
  (ownership map), design affirmation, and the deferred next increment.

Prior increments proved each slice in isolation: inc1 (native embedding
+ ORT
decoder), inc2b (ORT embedding + native decoder). Nothing exercised
**both** natively
at once until now.

## ORT byte-identical proof

No production source is touched, so the ORT decode path is
byte-identical to
`origin/main` by construction. Empirically the ORT baseline `[0,5,6,7]`
and the
fully-native run `[0,5,6,7]` match exactly.

## Tests

- `native_full_pipeline_parity` — **pass** (new)
- `native_step_component_parity`, `native_pipeline_decoder_parity` —
**pass**
- 343 engine lib unit tests — **pass**, 1 ignored
- `cargo fmt --all --check` — clean

CUDA-gated native tests were not run in this CPU environment (unchanged
by this PR).

## Deferred to the next increment (native wiring completion)

The one genuine remaining hard limitation the code itself flags
(`decoder_component.rs:244-260`):
`NativePipelineDecoder::mirror_last_present_kv`
bails — the native decoder keeps KV session-resident and does not expose
host present
tensors, so native selection runs the non-paged, fresh-decode path with
no
cross-request KV reuse. Wiring native present-KV exposure + paged
mirroring is higher
blast radius and is intentionally **not** bundled here.

Refs #82, #384. Working as Mary (native-decode / pipeline engineer).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant