Skip to content

Inc-1b PR-2: decode-specialized inlined-body Executor (dual-plan EAGER, flag default-OFF) - #588

Merged
justinchuby merged 3 commits into
mainfrom
squad/inc1b-pr2-decode-inline
Aug 2, 2026
Merged

justinchuby merged 3 commits into
mainfrom
squad/inc1b-pr2-decode-inline

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Inc-1b PR-2 — wire a decode-specialized inlined-body Executor (dual-plan EAGER, flag default-OFF)

Wires PR-1's inline_single_trip_scan_bodies transform (#580) into a second, decode-specialized Executor and routes single-token decode to it. Eager (capture stays OFF). Flag default-OFF ⇒ zero behavior change unless explicitly enabled.

What this does

  • Sibling Executor (onnx-runtime-session): Executor::build_decode_inline_sibling runs the transform, re-resolves interior shapes with Permissive inference (mirrors ChildExecutor::compile), and builds a second executor that shares the main exec's Arc<WeightStore> and Arc<dyn ExecutionProvider>. Returns None for a dense (non-hybrid) decoder. The prefill/main exec is left byte-identical.
  • Session API: InferenceSession::{enable_decode_inline, decode_inline_ready, run_decode_inline_with_device_bindings}. The sibling binds the identical persistent device state buffers the main exec used at the prefill→decode hand-off (bindings resolve by name; the transform leaves graph input/output names+order unchanged ⇒ recurrent-state continuity is automatic — design §3, the integration invariant).
  • Engine flag + routing (onnx-genai-engine): ONNX_GENAI_DECODE_INLINE_SCAN (default OFF; truthy 1/true/yes/on). Lazy build at the first single-token decode step. Single-token decode routes to the sibling on all three native single-token paths — CUDA greedy device-argmax fast path, CUDA logits path, and CPU in-place path. Greedy reuses the existing device-argmax kernel, so tie-breaking is byte-identical and full logits never round-trip to host.

Flag default-OFF — zero behavior change when off

When ONNX_GENAI_DECODE_INLINE_SCAN is unset/falsy, the sibling is never built, decode_inline latches Disabled, and every decode step uses today's Scan child-session path unchanged. An ordinary session is byte-identical to main and pays nothing.

Harry's 4 mandatory guards → tests

  1. Byte-identical parity + final recurrent state — decode_inline_sibling_is_byte_exact_with_scan_and_preserves_state (session): N decode steps, Scan plan vs inline plan, per-token outputs byte-identical AND final recurrent state identical.
  2. Runtime scan-axis extent==1 assertion + fallback — route_decode_inline_decision (pure) + decode_inline_routes_only_single_token_when_enabled / decode_inline_never_routes_when_disabled_or_unbuilt (engine): only single-token (extent-1) steps route to the sibling; multi-token steps fall back to the main Scan exec so a wrongly-collapsed graph is never run.
  3. Persistent state-buffer continuity — decode_inline_sibling_preserves_persistent_state_across_prefill_handoff (session).
  4. state_pairs ordering + shape check — decode_inline_sibling_preserves_state_output_order_and_resolves_shapes (session): first num_state present outputs map to present-state in io.state_pairs order; inlined-interior shapes resolve (Permissive) before use.

Plus decode_inline_sibling_none_for_dense_graph and decode_inline_flag_defaults_off_and_parses_truthy.

Measured OFF vs ON — Qwen3.6-27B int4 hybrid (H200)

profile_native --ep cuda --backend native --steady --decode-skip 8 --warmups 2 --runs 3 --tokens 64

flag decode ms/tok (medians, 4 runs) tok/s
OFF ~153 (149.2 / 155.9 / 157.7 / 150.6) ~6.5
ON ~122 (117.7 / 124.0 / 121.9 / 126.7) ~8.2

~1.26× decode speedup (design §5 predicted ~1.28×; the 167→130 ms/tok absolutes were on a slower baseline — same ratio here). Generated token ids byte-identical OFF vs ON on every run. The eager inline plan beats even the CUDA-graph-captured baseline because the captured Scan operator still pays real per-step child-dispatch + loop-state-collect work inside each replay; inlining removes that boundary entirely (design §5).

Byte-exact GPU e2e

native_autoderive_io_cuda_e2e.rs (#[ignore], stock 27b == CPU fp32 oracle, expected ids [11751,13,271,248068,271,248069,271,4639,369,4252,13,11751,369,279,6511,321]) run with the flag ON — see PR comment for the run result.

Scope

Design: cohaagen-27b-inc1b-design.md §1–3; transform PR #580.

Do not auto-merge — Harry reviews first (author-lockout).

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

justinchuby and others added 2 commits August 1, 2026 21:55
… EAGER, flag default-OFF)

Build a second, decode-specialized `Executor` (`decode_inline_exec`) from
`inline_single_trip_scan_bodies(&graph)` + Permissive `infer_graph`, sharing the
main exec's `Arc<WeightStore>` and `Arc<dyn ExecutionProvider>`. Route
single-token decode steps (greedy device-argmax fast path, CUDA logits path, and
CPU in-place path) to it; prefill and the main/prefill exec stay byte-identical.
Capture stays OFF (eager) — device-graph capture of the inlined body is PR-3.

Gated behind `ONNX_GENAI_DECODE_INLINE_SCAN` (default OFF): zero behavior change
unless explicitly enabled. Lazy build at the first single-token decode step; a
model with no single-trip-eligible recurrent Scan latches Disabled and stays on
today's Scan child-session path.

Harry's 4 guards -> tests:
1. Byte-identical parity + final state:
   decode_inline_sibling_is_byte_exact_with_scan_and_preserves_state (session)
2. Runtime scan-axis extent==1 fallback (single-token only routes; multi-token
   falls back to main Scan exec): route_decode_inline_decision +
   decode_inline_routes_only_single_token_when_enabled /
   decode_inline_never_routes_when_disabled_or_unbuilt (engine)
3. Persistent state-buffer continuity across prefill->decode:
   decode_inline_sibling_preserves_persistent_state_across_prefill_handoff (session)
4. state_pairs ordering + inlined-interior shape resolution:
   decode_inline_sibling_preserves_state_output_order_and_resolves_shapes (session)

Measured on Qwen3.6-27B int4 hybrid (H200, profile_native --steady --decode-skip 8
--warmups 2 --runs 3 --tokens 64): decode ~153 -> ~122 ms/tok (~6.5 -> ~8.2 tok/s,
~1.26x), generated ids byte-identical OFF vs ON.

Design: .squad/decisions/inbox/cohaagen-27b-inc1b-design.md §1-3; transform PR #580.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 4.34783% with 22 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.75%. Comparing base (70bac71) to head (7ef5b6d).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
crates/onnx-runtime-session/src/lib.rs 4.34% 22 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #588      +/-   ##
==========================================
- Coverage   81.36%   80.75%   -0.62%     
==========================================
  Files         317      318       +1     
  Lines      124181   126895    +2714     
  Branches   124181   126895    +2714     
==========================================
+ Hits       101042   102472    +1430     
- Misses      19043    20289    +1246     
- Partials     4096     4134      +38     
Flag Coverage Δ
cli-ort-linux 86.69% <ø> (ø)
cli-ort-windows 83.73% <ø> (ø)
mlas 80.96% <ø> (+3.05%) ⬆️
offline 80.51% <4.34%> (-0.69%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-session/src/lib.rs 71.12% <4.34%> (-2.08%) ⬇️

... and 11 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Aug 1, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 gather/medium_f32_threads=1-internal/32768 4.08 µs 6.56 µs +60.9%
🔴 gather/small_bf16_threads=1-internal/4096 518.0 ns 773.8 ns +49.4%
🔴 matmul/small_generic_f32_threads=8/1x256x256 34.23 µs 50.37 µs +47.2%
⚠️ gather/large_f16_threads=1-internal/131072 13.99 µs 17.70 µs +26.5%
⚠️ gather/medium_f16_threads=1-internal/32768 2.52 µs 3.15 µs +24.9%
⚠️ qwen3_sampling_processors/top_k_partial_selection 141.56 µs 174.44 µs +23.2%
⚠️ matmul/small_generic_bf16_threads=8/1x256x256 32.65 µs 39.74 µs +21.7%
⚠️ gather/large_bf16_threads=1-internal/131072 13.64 µs 16.42 µs +20.4%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.55 ms 4.25 ms +19.8%
⚠️ matmul/small_generic_bf16_threads=1/1x256x256 32.76 µs 39.21 µs +19.7%
⚠️ matmul/small_generic_f32_threads=1/1x256x256 38.11 µs 44.31 µs +16.3%
✅ matmul/small_generic_f16_threads=8/1x256x256 32.81 µs 37.05 µs +12.9%
✅ matmul/small_generic_f16_threads=1/1x256x256 31.52 µs 35.16 µs +11.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.13 ms 2.36 ms +10.7%
✅ sampling_latency/greedy_per_token 3.24 µs 3.59 µs +10.7%
✅ grammar_masking/llguidance_compute_mask/32 75.08 µs 82.50 µs +9.9%
✅ gather/small_f16_threads=1-internal/4096 545.3 ns 594.0 ns +8.9%
✅ tokenization/decode_tokens_per_second 6.14 ms 6.67 ms +8.7%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.70 ms 3.86 ms +4.4%
✅ matmul/medium_generic_f32_threads=8/32x512x512 916.65 µs 954.10 µs +4.1%
✅ kv_cache/alloc_dealloc_pages 36.95 µs 38.28 µs +3.6%
✅ sampling_latency/min_p_per_token 208.31 µs 213.98 µs +2.7%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.33 ms 2.39 ms +2.7%
✅ gather/small_f32_threads=1-internal/4096 695.3 ns 713.4 ns +2.6%
✅ gather/medium_bf16_threads=1-internal/32768 2.89 µs 2.96 µs +2.3%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.74 ms 5.84 ms +1.9%
✅ reduce_mean/medium_f32_threads=1-internal/65536 256.92 µs 259.95 µs +1.2%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.32 ms 1.33 ms +0.9%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 80.22 µs 80.84 µs +0.8%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 387.89 µs 390.85 µs +0.8%
✅ matmul/medium_generic_f16_threads=1/32x512x512 30.54 µs 30.67 µs +0.4%
✅ tokenization/encode_tokens_per_second 378.39 µs 378.59 µs +0.1%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.39 ms 9.36 ms -0.3%
✅ sampling_latency/top_k_per_token 53.17 µs 52.62 µs -1.1%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 85.84 µs 84.90 µs -1.1%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 537.13 µs 527.22 µs -1.8%
✅ matmul/medium_generic_f16_threads=8/32x512x512 31.43 µs 30.65 µs -2.5%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.54 µs 15.11 µs -2.7%
✅ add/small_bf16_threads=1-internal/1024 13.40 µs 12.99 µs -3.0%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.05 ms 1.98 ms -3.6%
✅ add/medium_bf16_threads=1-internal/262144 2.78 ms 2.65 ms -4.9%
✅ add/large_bf16_threads=1-internal/4194304 45.93 ms 42.76 ms -6.9%
✅ add/large_f32_threads=1-internal/4194304 43.69 ms 40.52 ms -7.2%
✅ add/large_f16_threads=1-internal/4194304 44.41 ms 41.09 ms -7.5%
✅ sampling_latency/top_p_per_token 381.21 µs 348.74 µs -8.5%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.12 ms 1.02 ms -9.2%
✅ add/medium_f32_threads=1-internal/262144 2.89 ms 2.55 ms -12.0%
✅ add/medium_f16_threads=1-internal/262144 2.93 ms 2.56 ms -12.7%
✅ add/small_f32_threads=1-internal/1024 226.9 ns 193.3 ns -14.8%
🟢 gather/large_f32_threads=1-internal/131072 48.57 µs 39.67 µs -18.3%
🟢 logit_processing/seven_processor_chain_per_step 319.08 µs 251.07 µs -21.3%
🟢 add/small_f16_threads=1-internal/1024 16.41 µs 12.81 µs -21.9%
🟢 qwen3_sampling_processors/top_k_top_p_fast 661.63 µs 253.94 µs -61.6%
🟢 qwen3_sampling_processors/top_p_fast_after_top_k 516.16 µs 109.26 µs -78.8%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.66 3.75 5.82 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Byte-exact GPU e2e gate — PASSED with the flag ON

native_autoderive_io_cuda_e2e.rs :: stock_export_auto_derives_io_and_matches_cpu_oracle run with ONNX_GENAI_DECODE_INLINE_SCAN=1 (CUDA, --test-threads=1, ONNX_GENAI_REQUIRE_CUDA=1):

cuda tokens=[11751, 13, 271, 248068, 271, 248069, 271, 4639, 369, 4252, 13, 11751, 369, 279, 6511, 321] (16.9s, decode-inline path)
cpu  tokens=[11751, 13, 271, 248068, 271, 248069, 271, 4639, 369, 4252, 13, 11751, 369, 279, 6511, 321] (5589s, fp32 oracle)
test result: ok. 1 passed; 0 failed; finished in 5606.79s

GPU (flag ON, inlined-body plan) is byte-identical to the CPU fp32 oracle and to the expected sequence.

Measured decode perf — Qwen3.6-27B int4 hybrid (H200)

profile_native --ep cuda --backend native --steady --decode-skip 8 --warmups 2 --runs 3 --tokens 64, 4 OFF/ON pairs:

flag decode ms/tok (medians) tok/s
OFF ~153 (149.2 / 155.9 / 157.7 / 150.6) ~6.5
ON ~122 (117.7 / 124.0 / 121.9 / 126.7) ~8.2

~1.26× decode speedup, generated ids byte-identical OFF vs ON on every run.

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review of PR #588 (Inc-1b PR-2 — decode-specialized inlined-body Executor, dual-plan EAGER, flag default-OFF). Head 7ef5b6d. I re-ran every check first-hand in a detached worktree and did not rely on author claims.

Blast radius — CONFINED. Local main was stale; the true PR base is 70bac71 (branch is 3 commits). git diff and gh pr view both report 10 files, +678/-1: native_decode/{backend,cpu,cuda,load,mod,tests}.rs, executor/{build,tests}.rs, session lib.rs, and the decision note. ZERO changes to plan_capture_segments / run_plan_segmented / graph.rs capture / #443/#543 / onnx-runtime-ep-cuda.

Flag default-OFF = zero behavior change — PROVEN. enable_decode_inline (the only sibling builder) is reachable only via maybe_enable_decode_inline, which checks decode_inline_scan_enabled() first and latches Disabled when unset/falsy, so the sibling is never constructed. run_decode_inline_with_device_bindings runs only when route_decode_inline_decision returns true (Enabled + sibling_ready + token_count==1). OFF path is the pre-existing Scan child-session path, byte-identical to today.

Tests — all 8 named guards pass. onnx-runtime-session executor: 65 passed/0 failed. onnx-genai-engine native_decode: 58 passed/0 failed.

Mutation testing — every guard NON-VACUOUS (broke code, mapped test failed, reverted, re-verified green + git diff --quiet):

  • (a) state-continuity: bind body state formal to a FRESH value instead of the persistent parent state input -> byte-exact AND persistent-state-handoff tests FAIL. Critical continuity guard fires.
  • (b) state_pairs ordering: swap present-state/scan output targets -> preserves_state_output_order_and_resolves_shapes FAILS (present-state resolves [1,3] vs [3]).
  • (c) extent==1 routing: token_count==1 -> token_count>=1 -> decode_inline_routes_only_single_token_when_enabled FAILS.
  • (d) byte-exact parity: rewrite inlined body Mul->Add -> decode_inline_sibling_is_byte_exact_with_scan_and_preserves_state FAILS (output First milestone #1 diverged).

Semantic equivalence — verified by reading the transform and the CUDA/CPU branches: same persistent state/KV buffers bound by name across the prefill->decode boundary (graph input/output names+order preserved); greedy device-argmax uses the identical read_greedy_result kernel (no host round-trip); only single-token decode is diverted; the sibling shares only read-only weights + EP and does not mutate the main exec; inline branches replicate mask/input/logical-len bookkeeping and correctly omit the capture-error poll because they run eager.

Byte-exact GPU e2e with flag ON (H200, CUDA_VISIBLE_DEVICES=2, ONNX_GENAI_DECODE_INLINE_SCAN=1, --test-threads=1) — PASS:
cuda tokens=[11751,13,271,248068,271,248069,271,4639,369,4252,13,11751,369,279,6511,321] (21.96s, inline path)
cpu tokens=[11751,13,271,248068,271,248069,271,4639,369,4252,13,11751,369,279,6511,321] (5450.58s oracle)
test result: ok. 1 passed; 0 failed. CUDA inline path == CPU fp32 oracle == expected sequence.

Independent perf (27B int4 hybrid, release, CUDA_VISIBLE_DEVICES=3, --steady --decode-skip 8 --warmups 2 --runs 3 --tokens 64):
flag OFF: decode median 145.96 ms/tok (6.85 tok/s)
flag ON: decode median 121.36 ms/tok (8.24 tok/s)
=> 1.20x decode speedup (author claimed ~1.26x; within GPU/run variance). Generated token ids BYTE-IDENTICAL OFF vs ON.

fmt/clippy — clean: cargo fmt --all --check exit 0; clippy on onnx-runtime-session and onnx-genai-engine (features cuda,native-backend) both -D warnings exit 0.

Non-blocking recs for PR-3 (capture; greenlight + capture-team sign-off required): re-introduce the check_device_capture_error poll once the inlined body is captured; add the design section 4 capture-engagement (segment-count growth) test; confirm inlined interior shapes join capture_warm snapshots and body sync-ops stay quarantined; add an assertion that inputs_embeds/routed-port paths never route to the sibling to lock the stated scope.

Full evidence: .squad/decisions/inbox/harry-review-588.md.

@justinchuby
justinchuby merged commit eca4088 into main Aug 2, 2026
15 checks passed
@justinchuby
justinchuby deleted the squad/inc1b-pr2-decode-inline branch August 2, 2026 01:34
justinchuby added a commit that referenced this pull request Aug 2, 2026
…et-A) (#589)

# Inc-1b PR-3 — capture-fold the decode-inline sibling (flag-gated,
bucket-A)

Part of the Inc-1b 27B decode-perf lane. PR-1 (#580,
inline_single_trip_scan_bodies) and PR-2 (#588, the eager decode-inline
sibling behind ONNX_GENAI_DECODE_INLINE_SCAN, default OFF) are merged.
This is **PR-3, the capture step**: let the decode-inline sibling's
inlined body ops fold into the CUDA-graph capture.

Bucket-(A)-only per my accepted scope note
(.squad/decisions/inbox/cohaagen-inc1b-pr3-scope.md): **no change to the
shared #443/#543 capture surface**. The sibling is an ordinary Executor
whose plan has no Scan after inlining, so it reuses the segmenter /
warm-seed / quarantine machinery verbatim; PR-3 only drives it through
the existing capture state machine.

## Flag-gated, default-OFF (structural no-op)
Capture engages only when ONNX_GENAI_DECODE_INLINE_SCAN gates
route_inline. Flag off: the sibling is never built and the inline branch
is never taken — byte-identical to current main.

## Harry's 4 PR-3 invariants (each with a non-vacuous test)
1. **Re-introduce check_device_capture_error()** on the sibling capture
path, piggybacked on the single logits/greedy device-to-host sync
(detection-before-consumption). The latch lives on the shared EP, so the
poll observes the sibling's captured-replay result; a latched violation
rejects the token and invalidates the graph.
2. **Capture-engagement test**
decode_inline_sibling_folds_body_into_captured_graph_byte_exact: the
inlined body folds into >= 1 captured segment while staying byte-exact
with the eager sibling run.
3. **Inlined-interior shapes join the warm-seeded snapshots** — proven
by the engagement test capturing after an eager warmup (warm-seed is the
precondition for capture engaging).
4. **Scope-lock** route_decode_inline_decision refuses inputs_embeds /
Routed step-input decoders (new has_eager_step_inputs arg); covered by
decode_inline_never_routes_when_decoder_has_eager_step_inputs.

## Single-slot / single-latch EP safety
Prefill runs the main exec (multi-token, non-capturable); ALL
single-token decode routes to the sibling; the main capture machine
stays dormant. So the shared EP's one graph slot + one capture-error
latch are owned solely by the sibling — no double-capture, no
cross-latch bleed. invalidate_graph now also resets the sibling's
host-side capture schedule so KV-growth / shape-change re-warms instead
of replaying a dropped graph.

## Correctness gate (real 27B, H200, qwen3.6-27b-int4-cuda)
Byte-exact vs the CPU fp32 oracle with capture engaged (flag ON), greedy
ids:

[11751, 13, 271, 248068, 271, 248069, 271, 4639, 369, 4252, 13, 11751,
369, 279, 6511, 321]

native_autoderive_io_cuda_e2e passed with
ONNX_GENAI_DECODE_INLINE_SCAN=1 (CUDA tokens == CPU oracle tokens).

## Measured decode perf (27B, H200, ms/tok, prefill+load cancelled by
token-count delta)
- flag OFF (main eager path): 143.8 ms/tok
- flag ON + capture: 70.1 ms/tok  => **2.05x**

ORT-CUDA crashes on this hybrid linear-attention export (documented
stl_vector assertion), so there is no live ORT baseline for this model;
the trusted reference is our native CPU fp32 oracle.

## Mutation map
- crates/onnx-runtime-session/src/lib.rs — 5 additive sibling capture
wrappers on InferenceSession + the CUDA capture-engagement test.
- crates/onnx-genai-engine/src/native_decode/cuda.rs —
inline_graph_phase field; new run_one_token_inline capture state
machine; invalidate_graph resets the sibling too; both inline branches
drive capture + re-add the capture-error poll; has_eager_step_inputs
widened to pub(super).
- crates/onnx-genai-engine/src/native_decode/mod.rs —
route_decode_inline_decision gains has_eager_step_inputs.
- crates/onnx-genai-engine/src/native_decode/tests.rs — updated calls +
new scope-lock test.

## Not in scope
Does NOT flip the default to ON — that stays Justin's decision. Build
evidence: .squad/decisions/inbox/cohaagen-inc1b-pr3-build.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant