Skip to content

refactor(pipeline): backend-neutral component ownership seam (unblocks native multi-component decode, #82/#384/35B-A3B) - #546

Merged
justinchuby merged 1 commit into
mainfrom
squad/pipeline-native-ownership
Jul 31, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/pipeline-native-ownership

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

The task was to introduce a backend-neutral ownership seam so the pipeline decode
loop can drive either ORT or native component sessions per step, then route
the existing ORT path through it byte-identically (increment-1).

On inspection, that seam already fully exists on origin/main — it landed
across the inc1→inc3c chain (#450, #478, #479, #485, #487, #533) and was hardened by
#543. PipelineDecodeLoopBackend owns no ORT Session/decode-state; it holds
only Box<dyn PipelineDecoderComponent> (stateful decoder seam) and
Vec<(_, Box<dyn ComponentSession>)> (stateless every_step seam), and both ORT and
native backends are driven through one decode loop via runtime env selection.

So increment-1 here is the piece the chain had not locked: a parity test proving
the seam's keystone end-state — every declared component running natively at
once
(native every_step embedding + native device-KV decoder in the same loop),
the exact shape a large multi-component package (up to the 35B-A3B 3-component
package) decodes through.

What changed (test-only, zero production change)

  • crates/onnx-genai-engine/tests/native_full_pipeline_parity.rs — drives the
    tiny-gemma4-vlm composite with both
    ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding and
    ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder, asserting the fully-native run is
    token-identical to the ORT baseline [0, 5, 6, 7].
  • crates/onnx-genai-engine/Cargo.toml — registers the test
    (required-features = ["native-backend"], CPU-only).
  • .squad/decisions/inbox/mary-pipeline-native-ownership.md — full assessment
    (ownership map), design affirmation, and the deferred next increment.

Prior increments proved each slice in isolation: inc1 (native embedding + ORT
decoder), inc2b (ORT embedding + native decoder). Nothing exercised both natively
at once until now.

ORT byte-identical proof

No production source is touched, so the ORT decode path is byte-identical to
origin/main by construction. Empirically the ORT baseline [0,5,6,7] and the
fully-native run [0,5,6,7] match exactly.

Tests

  • native_full_pipeline_parity — pass (new)
  • native_step_component_parity, native_pipeline_decoder_parity — pass
  • 343 engine lib unit tests — pass, 1 ignored
  • cargo fmt --all --check — clean

CUDA-gated native tests were not run in this CPU environment (unchanged by this PR).

Deferred to the next increment (native wiring completion)

The one genuine remaining hard limitation the code itself flags
(decoder_component.rs:244-260): NativePipelineDecoder::mirror_last_present_kv
bails — the native decoder keeps KV session-resident and does not expose host present
tensors, so native selection runs the non-paged, fresh-decode path with no
cross-request KV reuse. Wiring native present-KV exposure + paged mirroring is higher
blast radius and is intentionally not bundled here.

Refs #82, #384. Working as Mary (native-decode / pipeline engineer).

…-neutral seam

The backend-neutral component ownership seam (PipelineDecoderComponent +
ComponentSession) already landed across inc1..inc3c (#450, #478, #479, #485,
#487, #533, #543): PipelineDecodeLoopBackend owns no ORT Session/decode-state,
only Box<dyn PipelineDecoderComponent> and Box<dyn ComponentSession>, and both
ORT and native backends are driven through the same decode loop.

Prior increments proved each slice in isolation — native every_step embedding
with an ORT decoder (inc1) and a native device-KV decoder with an ORT embedding
(inc2b) — but nothing locked BOTH natively at once, which is exactly the shape a
large multi-component package (up to the 35B-A3B 3-component package) decodes
through.

Add native_full_pipeline_parity: it drives the tiny-gemma4-vlm composite with
ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding AND
ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder and asserts the fully-native run is
token-identical to the ORT baseline [0,5,6,7]. Test-only, zero production
change, so the ORT decode path is byte-identical to origin/main.

Assessment/design and the deferred next increment (native present-KV exposure +
paged cross-request reuse, which NativePipelineDecoder::mirror_last_present_kv
still bails on) are captured in
.squad/decisions/inbox/mary-pipeline-native-ownership.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.31%. Comparing base (cdc5af9) to head (5d908e8).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #546      +/-   ##
==========================================
+ Coverage   80.59%   81.31%   +0.72%     
==========================================
  Files         315      315              
  Lines      123446   123446              
  Branches   123446   123446              
==========================================
+ Hits        99487   100384     +897     
+ Misses      19908    19006     -902     
- Partials     4051     4056       +5     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.67% <ø> (-0.11%) ⬇️
mlas 78.72% <ø> (+0.81%) ⬆️
offline 81.27% <ø> (+0.75%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 7 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 gather/large_bf16_threads=1-internal/131072 9.77 µs 13.68 µs +40.0%
⚠️ gather/large_f32_threads=1-internal/131072 22.61 µs 28.48 µs +26.0%
⚠️ matmul/small_generic_f16_threads=8/1x256x256 30.07 µs 37.77 µs +25.6%
⚠️ gather/medium_bf16_threads=1-internal/32768 2.16 µs 2.48 µs +15.2%
✅ gather/small_bf16_threads=1-internal/4096 441.0 ns 496.1 ns +12.5%
✅ add/large_bf16_threads=1-internal/4194304 39.38 ms 44.23 ms +12.3%
✅ gather/small_f16_threads=1-internal/4096 445.1 ns 495.7 ns +11.4%
✅ gather/medium_f32_threads=1-internal/32768 3.72 µs 4.06 µs +9.0%
✅ gather/small_f32_threads=1-internal/4096 615.0 ns 663.9 ns +8.0%
✅ gather/medium_f16_threads=1-internal/32768 2.19 µs 2.30 µs +5.5%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.48 µs 15.23 µs +5.2%
✅ matmul/small_generic_f32_threads=1/1x256x256 33.91 µs 35.17 µs +3.7%
✅ gather/large_f16_threads=1-internal/131072 10.84 µs 10.93 µs +0.8%
✅ reduce_mean/large_f32_threads=1-internal/262144 983.84 µs 991.52 µs +0.8%
✅ add/small_f16_threads=1-internal/1024 13.12 µs 13.21 µs +0.7%
✅ matmul/small_generic_f32_threads=8/1x256x256 34.42 µs 34.20 µs -0.6%
✅ add/large_f16_threads=1-internal/4194304 44.09 ms 43.80 ms -0.6%
✅ add/large_f32_threads=1-internal/4194304 42.62 ms 41.82 ms -1.9%
✅ sampling_latency/top_k_per_token 482.74 µs 470.27 µs -2.6%
✅ sampling_latency/greedy_per_token 3.29 µs 3.19 µs -3.1%
✅ add/small_f32_threads=1-internal/1024 199.6 ns 193.3 ns -3.2%
✅ add/medium_f32_threads=1-internal/262144 2.73 ms 2.59 ms -5.1%
✅ reduce_mean/medium_f32_threads=1-internal/65536 267.36 µs 243.40 µs -9.0%
✅ add/medium_bf16_threads=1-internal/262144 2.91 ms 2.63 ms -9.4%
✅ add/medium_f16_threads=1-internal/262144 3.02 ms 2.70 ms -10.3%
✅ tokenization/decode_tokens_per_second 6.73 ms 5.97 ms -11.3%
✅ tokenization/encode_tokens_per_second 414.21 µs 363.54 µs -12.2%
✅ logit_processing/seven_processor_chain_per_step 1.29 ms 1.11 ms -14.4%
🟢 kv_cache/alloc_dealloc_pages 44.75 µs 36.38 µs -18.7%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 41.17 µs 33.25 µs -19.2%
🟢 add/small_bf16_threads=1-internal/1024 15.18 µs 12.10 µs -20.2%
🟢 matmul/large_generic_bf16_threads=1/32x1024x1024 2.47 ms 1.97 ms -20.5%
🟢 grammar_masking/llguidance_compute_mask/32 86.83 µs 68.92 µs -20.6%
🟢 matmul/medium_generic_f32_threads=1/32x512x512 2.79 ms 2.18 ms -22.0%
🟢 sampling_latency/top_p_per_token 1.23 ms 952.90 µs -22.4%
🟢 sampling_latency/min_p_per_token 431.46 µs 321.28 µs -25.5%
🟢 matmul/large_generic_f16_threads=1/32x1024x1024 104.36 µs 77.49 µs -25.7%
🟢 matmul/small_generic_f16_threads=1/1x256x256 37.77 µs 27.86 µs -26.3%
🟢 matmul/large_generic_f32_threads=1/32x1024x1024 12.18 ms 8.90 ms -27.0%
🟢 matmul/small_generic_bf16_threads=8/1x256x256 42.38 µs 29.25 µs -31.0%
🟢 matmul/medium_generic_bf16_threads=1/32x512x512 750.36 µs 481.11 µs -35.9%
🟢 matmul/medium_generic_f16_threads=1/32x512x512 44.91 µs 27.88 µs -37.9%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.57 ms 1.43 ms -44.6%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 50.57 µs 27.83 µs -45.0%
🟢 matmul/large_generic_f32_threads=8/32x1024x1024 7.05 ms 3.85 ms -45.4%
🟢 matmul/medium_generic_bf16_threads=8/32x512x512 692.94 µs 372.82 µs -46.2%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.77 ms 887.31 µs -49.8%
🟢 matmul/large_generic_f16_threads=8/32x1024x1024 173.41 µs 82.52 µs -52.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 2.88 3.87 6.13 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Reviewer: Lori (independent) — test-only PR, so the entire review is: is the new regression test NON-VACUOUS and CORRECT? Verified empirically in an independent worktree off origin/squad/pipeline-native-ownership. Answer: yes.

1. The test drives the FULL composite natively (both slices)

native_full_pipeline_parity.rs sets BOTH knobs before the native run and clears them for the ORT run:

  • ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding (every_step embedding component)
  • ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder (autoregressive decoder)

Same prompt [3, 7], greedy (temperature: 0.0, max_new_tokens: 4). It asserts native token IDs == ORT token IDs, AND independently pins the ORT baseline to the fixture's closed-form [0, 5, 6, 7] (so a regression that shifts BOTH backends identically still fails). Good defensive design.

2. It actually RUNS on the CPU CI env (not ignored/skipped)

required-features = ["native-backend"] only — NOT cuda-gated. Ran on CPU:

running 1 test
test full_native_pipeline_matches_ort_token_ids ... ok
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out

0 ignored, real pass. Within CI reach.

3. NON-VACUITY proven by perturbation (the key evidence)

I independently broke each native slice and confirmed the parity assert FIRES.

(a) Perturb native DECODER — spiked next_token_logits() in decoder_component.rs to force argmax:

assertion `left == right` failed: fully-native composite pipeline ... diverged from the ORT baseline
  left: [2, 2, 2, 2]
 right: [0, 5, 6, 7]

(b) Perturb native EMBEDDING — zeroed the NativeComponentSession::run output tensor in native_component.rs:

assertion `left == right` failed: fully-native composite pipeline ... diverged from the ORT baseline
  left: [7, 7, 7, 7]
 right: [0, 5, 6, 7]

Both native components are genuinely engaged and both feed the asserted tokens; the test is NOT vacuous. In both runs the ORT side (right) stayed [0,5,6,7], confirming the perturbation only touched the native path and the ORT/native separation is real. All perturbations reverted; tree is clean; test green again afterward.

(Note: a tiny low-mantissa 1-byte flip on the embedding did NOT flip tokens — that reflects the closed-form fixture's numerical robustness, not vacuity; the drastic zeroing fired cleanly.)

4. Fixture is a genuine multi-component composite

tests/fixtures/tiny-gemma4-vlm/inference_metadata.yaml declares strategy.kind: composite with THREE models: vision_encoder (run_on prompt_only), embedding (run_on every_step, io.token_input: input_ids), decoder (run_on every_step autoregressive), plus embedding.inputs_embeds -> decoder.inputs_embeds dataflow. This exercises the every-step embedding component + decoder simultaneously — exactly the multi-slice shape, not a single-component decoder.

5. Hygiene

  • cargo fmt --all --check — clean (exit 0).
  • Cargo.toml change is ONLY a new [[test]] entry (required-features = ["native-backend"]) — test wiring, no runtime dependency creep.
  • No production source touched, so the ORT decode path is byte-identical to main by construction.

Conclusion: The regression test locks the previously-unguarded keystone (native embedding + native decoder in one loop) and is proven non-vacuous by two independent perturbations. Approving.

@justinchuby
justinchuby marked this pull request as ready for review July 31, 2026 05:17
@justinchuby
justinchuby merged commit 080735a into main Jul 31, 2026
14 checks passed
@justinchuby
justinchuby deleted the squad/pipeline-native-ownership branch July 31, 2026 05:17
justinchuby added a commit that referenced this pull request Aug 3, 2026
…+ GAP-3 decomposition (increment 1/N) (#613)

## What & why

Scope pass on **GAP-3 (native pipeline decode)**. The finding is that
GAP-3's core is
**already implemented and conformance-locked on `origin/main`** — the
task was authored
against a local checkout (`1ba215ee`) that is ~35+ merged PRs behind
`origin/main`
(`be6d4e34`, #612). So this PR does **not** add new decode
functionality. It is a
bounded, **zero-behavior-change** truth-up + a design/decomposition
drop.

This is **GAP-3 housekeeping increment 1 of N** (see the design note in
this diff:

`.squad/decisions/inbox/cohaagen-gap3-native-pipeline-decode-design.md`).
The
substantive increments already landed:

- Backend-neutral component ownership seam — #546
- `NativePipelineDecoder` via `PipelineDecoderComponent` (Inc2b) — #479
- Pure-native multi-component decode wiring (Inc-A) — #565
- Native present-KV mirroring, paged (Inc-C) — #566
- Device-resident present-KV read-out (Inc-D / D.1) — #567/#568
- rank-3 mrope native positions — #543 · text-only decode pipeline —
#535 · fp16 TopK MoE router — #612

The pipeline decode loop (`PipelineDecodeLoopBackend`) already owns
`Box<dyn PipelineDecoderComponent>` + `Box<dyn ComponentSession>`, not
an ORT `Session`;
`DecodeState`/ORT `Value` are confined to `OrtPipelineDecoder`. The
Qwen3.5-0.8B hybrid
(same class as Qwen3.6-35B-A3B) decodes natively with **token-for-token
parity vs ORT**
under `tests/qwen35_0_8b_hybrid_native_cuda_e2e.rs`.

## The change

`native_component.rs`'s module doc still claimed wiring native sessions
into *"the
ORT-owned pipeline decode loop is the remaining GAP 3 work"* — false
since #546/#565.
Corrected to describe the now-backend-neutral loop and the merged
Inc-A/C/D, and to name
the genuinely-remaining feature-sized gaps (non-flat plans, native
cross-attn/vision KV).

**Behavior-preserving:** comment-only in `native_component.rs` + a
tracked decision drop.
No code path, signature, or data change.

## Remaining decomposition (in the design note)

Each is feature-sized (needs op/attention support and/or fixtures —
**not** zero-behavior),
none blocks the text-only 35B-A3B native number: R2 native
sliding-window paged mirror ·
R3 Inc-D.2 discontinuous prefix reuse · R4 native cross-attn/vision KV
(Inc3) · R5 non-flat
plans native · R6 (optional) neutral host tensor in the shared pool.

## Verification for the first native pipeline model

Byte/token-exact differential vs an ORT-backend decode of the **same
artifact** (ORT
front-end for both arms, decoder EP isolated), greedy, token-for-token —
already in place
for the 0.8B hybrid; same harness pattern applied to the real 35B-A3B is
the next step.

## Checks

- `cargo fmt --all` clean
- `cargo clippy -p onnx-genai-engine --features "native-backend cuda" --
-D warnings` clean

Left **open for Harry review**; do not merge without it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 22, 2026
…ined contract

gemma4-real-packages published the faithful E2B fix (parent kept the
schema strict: folded_carry_seed names a real target output, no
request-input escape hatch). mobius gemma4 now emits the post-final-norm
hidden as hidden_states.{idx} (PR #546 @ 710d4927, backward-compatible),
and the real E2B packages carry the SAME folded_carry_seed/token_embedding
contract at scale with real ports (target hidden_states.34,
model.embed_tokens.weight [262144,1536] fp16).

This empirically confirms the Phase-1 field-reading chained driver
generalizes from the tiny fixture to real models with no model-name gate.
Record the real-package coordinates as an optional post-parity scale case;
the tiny gemma4_chained @ 8a66e2c stays the required hermetic parity
fixture (unchanged, revalidated). No seam change; still gated on #1716.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant