Repository navigation
Inc-1b PR-2: decode-specialized inlined-body Executor (dual-plan EAGER, flag default-OFF) - #588
Conversation
… EAGER, flag default-OFF) Build a second, decode-specialized `Executor` (`decode_inline_exec`) from `inline_single_trip_scan_bodies(&graph)` + Permissive `infer_graph`, sharing the main exec's `Arc<WeightStore>` and `Arc<dyn ExecutionProvider>`. Route single-token decode steps (greedy device-argmax fast path, CUDA logits path, and CPU in-place path) to it; prefill and the main/prefill exec stay byte-identical. Capture stays OFF (eager) — device-graph capture of the inlined body is PR-3. Gated behind `ONNX_GENAI_DECODE_INLINE_SCAN` (default OFF): zero behavior change unless explicitly enabled. Lazy build at the first single-token decode step; a model with no single-trip-eligible recurrent Scan latches Disabled and stays on today's Scan child-session path. Harry's 4 guards -> tests: 1. Byte-identical parity + final state: decode_inline_sibling_is_byte_exact_with_scan_and_preserves_state (session) 2. Runtime scan-axis extent==1 fallback (single-token only routes; multi-token falls back to main Scan exec): route_decode_inline_decision + decode_inline_routes_only_single_token_when_enabled / decode_inline_never_routes_when_disabled_or_unbuilt (engine) 3. Persistent state-buffer continuity across prefill->decode: decode_inline_sibling_preserves_persistent_state_across_prefill_handoff (session) 4. state_pairs ordering + inlined-interior shape resolution: decode_inline_sibling_preserves_state_output_order_and_resolves_shapes (session) Measured on Qwen3.6-27B int4 hybrid (H200, profile_native --steady --decode-skip 8 --warmups 2 --runs 3 --tokens 64): decode ~153 -> ~122 ms/tok (~6.5 -> ~8.2 tok/s, ~1.26x), generated ids byte-identical OFF vs ON. Design: .squad/decisions/inbox/cohaagen-27b-inc1b-design.md §1-3; transform PR #580. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #588 +/- ##
==========================================
- Coverage 81.36% 80.75% -0.62%
==========================================
Files 317 318 +1
Lines 124181 126895 +2714
Branches 124181 126895 +2714
==========================================
+ Hits 101042 102472 +1430
- Misses 19043 20289 +1246
- Partials 4096 4134 +38
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Byte-exact GPU e2e gate — PASSED with the flag ON
GPU (flag ON, inlined-body plan) is byte-identical to the CPU fp32 oracle and to the expected sequence. Measured decode perf — Qwen3.6-27B int4 hybrid (H200)
~1.26× decode speedup, generated ids byte-identical OFF vs ON on every run. |
|
VERDICT: APPROVE Independent review of PR #588 (Inc-1b PR-2 — decode-specialized inlined-body Executor, dual-plan EAGER, flag default-OFF). Head 7ef5b6d. I re-ran every check first-hand in a detached worktree and did not rely on author claims. Blast radius — CONFINED. Local main was stale; the true PR base is 70bac71 (branch is 3 commits). git diff and gh pr view both report 10 files, +678/-1: native_decode/{backend,cpu,cuda,load,mod,tests}.rs, executor/{build,tests}.rs, session lib.rs, and the decision note. ZERO changes to plan_capture_segments / run_plan_segmented / graph.rs capture / #443/#543 / onnx-runtime-ep-cuda. Flag default-OFF = zero behavior change — PROVEN. enable_decode_inline (the only sibling builder) is reachable only via maybe_enable_decode_inline, which checks decode_inline_scan_enabled() first and latches Disabled when unset/falsy, so the sibling is never constructed. run_decode_inline_with_device_bindings runs only when route_decode_inline_decision returns true (Enabled + sibling_ready + token_count==1). OFF path is the pre-existing Scan child-session path, byte-identical to today. Tests — all 8 named guards pass. onnx-runtime-session executor: 65 passed/0 failed. onnx-genai-engine native_decode: 58 passed/0 failed. Mutation testing — every guard NON-VACUOUS (broke code, mapped test failed, reverted, re-verified green + git diff --quiet):
Semantic equivalence — verified by reading the transform and the CUDA/CPU branches: same persistent state/KV buffers bound by name across the prefill->decode boundary (graph input/output names+order preserved); greedy device-argmax uses the identical read_greedy_result kernel (no host round-trip); only single-token decode is diverted; the sibling shares only read-only weights + EP and does not mutate the main exec; inline branches replicate mask/input/logical-len bookkeeping and correctly omit the capture-error poll because they run eager. Byte-exact GPU e2e with flag ON (H200, CUDA_VISIBLE_DEVICES=2, ONNX_GENAI_DECODE_INLINE_SCAN=1, --test-threads=1) — PASS: Independent perf (27B int4 hybrid, release, CUDA_VISIBLE_DEVICES=3, --steady --decode-skip 8 --warmups 2 --runs 3 --tokens 64): fmt/clippy — clean: cargo fmt --all --check exit 0; clippy on onnx-runtime-session and onnx-genai-engine (features cuda,native-backend) both -D warnings exit 0. Non-blocking recs for PR-3 (capture; greenlight + capture-team sign-off required): re-introduce the check_device_capture_error poll once the inlined body is captured; add the design section 4 capture-engagement (segment-count growth) test; confirm inlined interior shapes join capture_warm snapshots and body sync-ops stay quarantined; add an assertion that inputs_embeds/routed-port paths never route to the sibling to lock the stated scope. Full evidence: .squad/decisions/inbox/harry-review-588.md. |
…et-A) (#589) # Inc-1b PR-3 — capture-fold the decode-inline sibling (flag-gated, bucket-A) Part of the Inc-1b 27B decode-perf lane. PR-1 (#580, inline_single_trip_scan_bodies) and PR-2 (#588, the eager decode-inline sibling behind ONNX_GENAI_DECODE_INLINE_SCAN, default OFF) are merged. This is **PR-3, the capture step**: let the decode-inline sibling's inlined body ops fold into the CUDA-graph capture. Bucket-(A)-only per my accepted scope note (.squad/decisions/inbox/cohaagen-inc1b-pr3-scope.md): **no change to the shared #443/#543 capture surface**. The sibling is an ordinary Executor whose plan has no Scan after inlining, so it reuses the segmenter / warm-seed / quarantine machinery verbatim; PR-3 only drives it through the existing capture state machine. ## Flag-gated, default-OFF (structural no-op) Capture engages only when ONNX_GENAI_DECODE_INLINE_SCAN gates route_inline. Flag off: the sibling is never built and the inline branch is never taken — byte-identical to current main. ## Harry's 4 PR-3 invariants (each with a non-vacuous test) 1. **Re-introduce check_device_capture_error()** on the sibling capture path, piggybacked on the single logits/greedy device-to-host sync (detection-before-consumption). The latch lives on the shared EP, so the poll observes the sibling's captured-replay result; a latched violation rejects the token and invalidates the graph. 2. **Capture-engagement test** decode_inline_sibling_folds_body_into_captured_graph_byte_exact: the inlined body folds into >= 1 captured segment while staying byte-exact with the eager sibling run. 3. **Inlined-interior shapes join the warm-seeded snapshots** — proven by the engagement test capturing after an eager warmup (warm-seed is the precondition for capture engaging). 4. **Scope-lock** route_decode_inline_decision refuses inputs_embeds / Routed step-input decoders (new has_eager_step_inputs arg); covered by decode_inline_never_routes_when_decoder_has_eager_step_inputs. ## Single-slot / single-latch EP safety Prefill runs the main exec (multi-token, non-capturable); ALL single-token decode routes to the sibling; the main capture machine stays dormant. So the shared EP's one graph slot + one capture-error latch are owned solely by the sibling — no double-capture, no cross-latch bleed. invalidate_graph now also resets the sibling's host-side capture schedule so KV-growth / shape-change re-warms instead of replaying a dropped graph. ## Correctness gate (real 27B, H200, qwen3.6-27b-int4-cuda) Byte-exact vs the CPU fp32 oracle with capture engaged (flag ON), greedy ids: [11751, 13, 271, 248068, 271, 248069, 271, 4639, 369, 4252, 13, 11751, 369, 279, 6511, 321] native_autoderive_io_cuda_e2e passed with ONNX_GENAI_DECODE_INLINE_SCAN=1 (CUDA tokens == CPU oracle tokens). ## Measured decode perf (27B, H200, ms/tok, prefill+load cancelled by token-count delta) - flag OFF (main eager path): 143.8 ms/tok - flag ON + capture: 70.1 ms/tok => **2.05x** ORT-CUDA crashes on this hybrid linear-attention export (documented stl_vector assertion), so there is no live ORT baseline for this model; the trusted reference is our native CPU fp32 oracle. ## Mutation map - crates/onnx-runtime-session/src/lib.rs — 5 additive sibling capture wrappers on InferenceSession + the CUDA capture-engagement test. - crates/onnx-genai-engine/src/native_decode/cuda.rs — inline_graph_phase field; new run_one_token_inline capture state machine; invalidate_graph resets the sibling too; both inline branches drive capture + re-add the capture-error poll; has_eager_step_inputs widened to pub(super). - crates/onnx-genai-engine/src/native_decode/mod.rs — route_decode_inline_decision gains has_eager_step_inputs. - crates/onnx-genai-engine/src/native_decode/tests.rs — updated calls + new scope-lock test. ## Not in scope Does NOT flip the default to ON — that stays Justin's decision. Build evidence: .squad/decisions/inbox/cohaagen-inc1b-pr3-build.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Inc-1b PR-2 — wire a decode-specialized inlined-body Executor (dual-plan EAGER, flag default-OFF)
Wires PR-1's
inline_single_trip_scan_bodiestransform (#580) into a second, decode-specializedExecutorand routes single-token decode to it. Eager (capture stays OFF). Flag default-OFF ⇒ zero behavior change unless explicitly enabled.What this does
onnx-runtime-session):Executor::build_decode_inline_siblingruns the transform, re-resolves interior shapes with Permissive inference (mirrorsChildExecutor::compile), and builds a second executor that shares the main exec'sArc<WeightStore>andArc<dyn ExecutionProvider>. ReturnsNonefor a dense (non-hybrid) decoder. The prefill/main exec is left byte-identical.InferenceSession::{enable_decode_inline, decode_inline_ready, run_decode_inline_with_device_bindings}. The sibling binds the identical persistent device state buffers the main exec used at the prefill→decode hand-off (bindings resolve by name; the transform leaves graph input/output names+order unchanged ⇒ recurrent-state continuity is automatic — design §3, the integration invariant).onnx-genai-engine):ONNX_GENAI_DECODE_INLINE_SCAN(default OFF; truthy1/true/yes/on). Lazy build at the first single-token decode step. Single-token decode routes to the sibling on all three native single-token paths — CUDA greedy device-argmax fast path, CUDA logits path, and CPU in-place path. Greedy reuses the existing device-argmax kernel, so tie-breaking is byte-identical and full logits never round-trip to host.Flag default-OFF — zero behavior change when off
When
ONNX_GENAI_DECODE_INLINE_SCANis unset/falsy, the sibling is never built,decode_inlinelatchesDisabled, and every decode step uses today's Scan child-session path unchanged. An ordinary session is byte-identical tomainand pays nothing.Harry's 4 mandatory guards → tests
decode_inline_sibling_is_byte_exact_with_scan_and_preserves_state(session): N decode steps, Scan plan vs inline plan, per-token outputs byte-identical AND final recurrent state identical.route_decode_inline_decision(pure) +decode_inline_routes_only_single_token_when_enabled/decode_inline_never_routes_when_disabled_or_unbuilt(engine): only single-token (extent-1) steps route to the sibling; multi-token steps fall back to the main Scan exec so a wrongly-collapsed graph is never run.decode_inline_sibling_preserves_persistent_state_across_prefill_handoff(session).decode_inline_sibling_preserves_state_output_order_and_resolves_shapes(session): firstnum_statepresent outputs map to present-state inio.state_pairsorder; inlined-interior shapes resolve (Permissive) before use.Plus
decode_inline_sibling_none_for_dense_graphanddecode_inline_flag_defaults_off_and_parses_truthy.Measured OFF vs ON — Qwen3.6-27B int4 hybrid (H200)
profile_native --ep cuda --backend native --steady --decode-skip 8 --warmups 2 --runs 3 --tokens 64~1.26× decode speedup (design §5 predicted ~1.28×; the 167→130 ms/tok absolutes were on a slower baseline — same ratio here). Generated token ids byte-identical OFF vs ON on every run. The eager inline plan beats even the CUDA-graph-captured baseline because the captured Scan operator still pays real per-step child-dispatch + loop-state-collect work inside each replay; inlining removes that boundary entirely (design §5).
Byte-exact GPU e2e
native_autoderive_io_cuda_e2e.rs(#[ignore], stock 27b == CPU fp32 oracle, expected ids[11751,13,271,248068,271,248069,271,4639,369,4252,13,11751,369,279,6511,321]) run with the flag ON — see PR comment for the run result.Scope
inputs_embedsand multi-token steps keep the main exec.clippy -D warningsexit 0 ononnx-runtime-sessionandonnx-genai-engine(incl.--features cuda,native-backend).Design:
cohaagen-27b-inc1b-design.md§1–3; transform PR #580.Do not auto-merge — Harry reviews first (author-lockout).
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com