Repository navigation
refactor(pipeline): backend-neutral component ownership seam (unblocks native multi-component decode, #82/#384/35B-A3B) - #546
Conversation
…-neutral seam The backend-neutral component ownership seam (PipelineDecoderComponent + ComponentSession) already landed across inc1..inc3c (#450, #478, #479, #485, #487, #533, #543): PipelineDecodeLoopBackend owns no ORT Session/decode-state, only Box<dyn PipelineDecoderComponent> and Box<dyn ComponentSession>, and both ORT and native backends are driven through the same decode loop. Prior increments proved each slice in isolation — native every_step embedding with an ORT decoder (inc1) and a native device-KV decoder with an ORT embedding (inc2b) — but nothing locked BOTH natively at once, which is exactly the shape a large multi-component package (up to the 35B-A3B 3-component package) decodes through. Add native_full_pipeline_parity: it drives the tiny-gemma4-vlm composite with ONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embedding AND ONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder and asserts the fully-native run is token-identical to the ORT baseline [0,5,6,7]. Test-only, zero production change, so the ORT decode path is byte-identical to origin/main. Assessment/design and the deferred next increment (native present-KV exposure + paged cross-request reuse, which NativePipelineDecoder::mirror_last_present_kv still bails on) are captured in .squad/decisions/inbox/mary-pipeline-native-ownership.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #546 +/- ##
==========================================
+ Coverage 80.59% 81.31% +0.72%
==========================================
Files 315 315
Lines 123446 123446
Branches 123446 123446
==========================================
+ Hits 99487 100384 +897
+ Misses 19908 19006 -902
- Partials 4051 4056 +5
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
VERDICT: APPROVE Reviewer: Lori (independent) — test-only PR, so the entire review is: is the new regression test NON-VACUOUS and CORRECT? Verified empirically in an independent worktree off origin/squad/pipeline-native-ownership. Answer: yes. 1. The test drives the FULL composite natively (both slices)
Same prompt 2. It actually RUNS on the CPU CI env (not ignored/skipped)
0 ignored, real pass. Within CI reach. 3. NON-VACUITY proven by perturbation (the key evidence)I independently broke each native slice and confirmed the parity assert FIRES. (a) Perturb native DECODER — spiked (b) Perturb native EMBEDDING — zeroed the Both native components are genuinely engaged and both feed the asserted tokens; the test is NOT vacuous. In both runs the ORT side ( (Note: a tiny low-mantissa 1-byte flip on the embedding did NOT flip tokens — that reflects the closed-form fixture's numerical robustness, not vacuity; the drastic zeroing fired cleanly.) 4. Fixture is a genuine multi-component composite
5. Hygiene
Conclusion: The regression test locks the previously-unguarded keystone (native embedding + native decoder in one loop) and is proven non-vacuous by two independent perturbations. Approving. |
…+ GAP-3 decomposition (increment 1/N) (#613) ## What & why Scope pass on **GAP-3 (native pipeline decode)**. The finding is that GAP-3's core is **already implemented and conformance-locked on `origin/main`** — the task was authored against a local checkout (`1ba215ee`) that is ~35+ merged PRs behind `origin/main` (`be6d4e34`, #612). So this PR does **not** add new decode functionality. It is a bounded, **zero-behavior-change** truth-up + a design/decomposition drop. This is **GAP-3 housekeeping increment 1 of N** (see the design note in this diff: `.squad/decisions/inbox/cohaagen-gap3-native-pipeline-decode-design.md`). The substantive increments already landed: - Backend-neutral component ownership seam — #546 - `NativePipelineDecoder` via `PipelineDecoderComponent` (Inc2b) — #479 - Pure-native multi-component decode wiring (Inc-A) — #565 - Native present-KV mirroring, paged (Inc-C) — #566 - Device-resident present-KV read-out (Inc-D / D.1) — #567/#568 - rank-3 mrope native positions — #543 · text-only decode pipeline — #535 · fp16 TopK MoE router — #612 The pipeline decode loop (`PipelineDecodeLoopBackend`) already owns `Box<dyn PipelineDecoderComponent>` + `Box<dyn ComponentSession>`, not an ORT `Session`; `DecodeState`/ORT `Value` are confined to `OrtPipelineDecoder`. The Qwen3.5-0.8B hybrid (same class as Qwen3.6-35B-A3B) decodes natively with **token-for-token parity vs ORT** under `tests/qwen35_0_8b_hybrid_native_cuda_e2e.rs`. ## The change `native_component.rs`'s module doc still claimed wiring native sessions into *"the ORT-owned pipeline decode loop is the remaining GAP 3 work"* — false since #546/#565. Corrected to describe the now-backend-neutral loop and the merged Inc-A/C/D, and to name the genuinely-remaining feature-sized gaps (non-flat plans, native cross-attn/vision KV). **Behavior-preserving:** comment-only in `native_component.rs` + a tracked decision drop. No code path, signature, or data change. ## Remaining decomposition (in the design note) Each is feature-sized (needs op/attention support and/or fixtures — **not** zero-behavior), none blocks the text-only 35B-A3B native number: R2 native sliding-window paged mirror · R3 Inc-D.2 discontinuous prefix reuse · R4 native cross-attn/vision KV (Inc3) · R5 non-flat plans native · R6 (optional) neutral host tensor in the shared pool. ## Verification for the first native pipeline model Byte/token-exact differential vs an ORT-backend decode of the **same artifact** (ORT front-end for both arms, decoder EP isolated), greedy, token-for-token — already in place for the 0.8B hybrid; same harness pattern applied to the real 35B-A3B is the next step. ## Checks - `cargo fmt --all` clean - `cargo clippy -p onnx-genai-engine --features "native-backend cuda" -- -D warnings` clean Left **open for Harry review**; do not merge without it. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ined contract
gemma4-real-packages published the faithful E2B fix (parent kept the
schema strict: folded_carry_seed names a real target output, no
request-input escape hatch). mobius gemma4 now emits the post-final-norm
hidden as hidden_states.{idx} (PR #546 @ 710d4927, backward-compatible),
and the real E2B packages carry the SAME folded_carry_seed/token_embedding
contract at scale with real ports (target hidden_states.34,
model.embed_tokens.weight [262144,1536] fp16).
This empirically confirms the Phase-1 field-reading chained driver
generalizes from the tiny fixture to real models with no model-name gate.
Record the real-package coordinates as an optional post-parity scale case;
the tiny gemma4_chained @ 8a66e2c stays the required hermetic parity
fixture (unchanged, revalidated). No seam change; still gated on #1716.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Summary
The task was to introduce a backend-neutral ownership seam so the pipeline decode
loop can drive either ORT or native component sessions per step, then route
the existing ORT path through it byte-identically (increment-1).
On inspection, that seam already fully exists on
origin/main— it landedacross the inc1→inc3c chain (#450, #478, #479, #485, #487, #533) and was hardened by
#543.
PipelineDecodeLoopBackendowns no ORTSession/decode-state; it holdsonly
Box<dyn PipelineDecoderComponent>(stateful decoder seam) andVec<(_, Box<dyn ComponentSession>)>(stateless every_step seam), and both ORT andnative backends are driven through one decode loop via runtime env selection.
So increment-1 here is the piece the chain had not locked: a parity test proving
the seam's keystone end-state — every declared component running natively at
once (native every_step embedding + native device-KV decoder in the same loop),
the exact shape a large multi-component package (up to the 35B-A3B 3-component
package) decodes through.
What changed (test-only, zero production change)
crates/onnx-genai-engine/tests/native_full_pipeline_parity.rs— drives thetiny-gemma4-vlmcomposite with bothONNX_GENAI_PIPELINE_NATIVE_STEP_COMPONENTS=embeddingandONNX_GENAI_PIPELINE_NATIVE_DECODER=decoder, asserting the fully-native run istoken-identical to the ORT baseline
[0, 5, 6, 7].crates/onnx-genai-engine/Cargo.toml— registers the test(
required-features = ["native-backend"], CPU-only)..squad/decisions/inbox/mary-pipeline-native-ownership.md— full assessment(ownership map), design affirmation, and the deferred next increment.
Prior increments proved each slice in isolation: inc1 (native embedding + ORT
decoder), inc2b (ORT embedding + native decoder). Nothing exercised both natively
at once until now.
ORT byte-identical proof
No production source is touched, so the ORT decode path is byte-identical to
origin/mainby construction. Empirically the ORT baseline[0,5,6,7]and thefully-native run
[0,5,6,7]match exactly.Tests
native_full_pipeline_parity— pass (new)native_step_component_parity,native_pipeline_decoder_parity— passcargo fmt --all --check— cleanCUDA-gated native tests were not run in this CPU environment (unchanged by this PR).
Deferred to the next increment (native wiring completion)
The one genuine remaining hard limitation the code itself flags
(
decoder_component.rs:244-260):NativePipelineDecoder::mirror_last_present_kvbails — the native decoder keeps KV session-resident and does not expose host present
tensors, so native selection runs the non-paged, fresh-decode path with no
cross-request KV reuse. Wiring native present-KV exposure + paged mirroring is higher
blast radius and is intentionally not bundled here.
Refs #82, #384. Working as Mary (native-decode / pipeline engineer).