Skip to content

feat(engine): flag-gated single-trip Scan inline dual-path (27B Scan-capture slice 1a) - #564

Merged
justinchuby merged 1 commit into
mainfrom
feat/27b-scan-capture-1a
Jul 31, 2026
Merged

justinchuby merged 1 commit into
mainfrom
feat/27b-scan-capture-1a

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

27B decode: flag-gated single-trip Scan inline dual-path (slice 1a of Scan→CUDA-capture)

First slice of the Justin-approved 27B Scan→CUDA-capture workstream (root-cause: eager Scan/LinearAttention recurrence = 56.5% of 27B decode, structurally un-capturable; ~15-30× lever). Slice 1a is correctness-only — NO capture changes yet.

What

A runtime-conditional dual-path in exec_scan: when ONNX_GENAI_SCAN_INLINE_SINGLE_TRIP is ON and the runtime trip_count==1, the Scan body runs once straight-line via a shared run_scan_body_step helper; otherwise the unchanged loop. Selection is runtime-keyed (not a graph rewrite) because prefill (trip_count>1) and decode (trip_count==1) share one executor/plan — a static seq=1 inline would corrupt prefill.

  • Flag default OFF ⇒ zero behavior change (loop for all trip counts).
  • Loop and inline share the same body-step + finishing code → byte-exact with a one-iteration loop by construction. DRY, no model/op special-casing.
  • No capture-core changes (provider.rs/capture.rs untouched); Scan still declines capture in both paths. That's slice 1b.

Evidence

Review

Independent review by Melina (author Mary locked out on rejection) — APPROVE with full non-vacuity evidence.

1b handoff

Let the single-trip inlined body enter CUDA-graph capture (blast radius provider.rs:458 + executor/capture.rs); assert captures/replays counters rise and 27B tokens stay byte-identical to this 1a reference.

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

A Scan whose RUNTIME trip_count == 1 (a single decode step) runs its body
once straight-line instead of the generic exec_scan loop, while prefill
(trip_count = prompt_len > 1) keeps the unchanged loop. Selection is keyed
on the observed trip_count at execution time — NOT a graph rewrite — because
prefill and decode share one executor/plan, so a static single-trip bake
would corrupt prefill.

Gated by ONNX_GENAI_SCAN_INLINE_SINGLE_TRIP (default OFF): flag OFF is zero
behavior change. Both the loop and the inline path drive the body through a
shared run_scan_body_step helper and share the finishing code, so the inline
path is byte-exact with a one-iteration loop by construction (DRY, general —
no op/model special-casing).

Correctness-only foundation; no capture changes (slice 1b will let the
inlined body enter CUDA-graph capture). No changes to plan_capture_region /
node_capture_reason / StructuralCaptureDecline.

Evidence:
- CPU test scan_single_trip_inline_is_byte_exact_and_runtime_keyed: byte-exact
  vs loop over both outputs, engages only at trip_count==1, count==0 on
  prefill (runtime-keyed). Mutation-checked non-vacuous.
- CUDA-gated regression cuda_scan_single_trip_inline_is_byte_exact_and_runtime_keyed
  (device 4): same on real ORT-CUDA, via scan_inline_single_trip_count.
- On-model 27B (qwen3.6-27b int4, device 4, greedy 48 tok): token ids
  IDENTICAL flag OFF vs ON across prefill + 48 decode steps.
- Re-ran #554 (recurrent-state reset) and #544 (weight page-in WAR) green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review (Melina) — author Mary locked out on rejection; reviewer modified no product code (mutations reverted, tree clean).

  • Dual-path byte-exact BY CONSTRUCTION: both branches drive the body through shared run_scan_body_step and the same finishing code (finish_scan); inline reuses the loop's step-0 scan_slices (source_index 0 at trip_count==1 both directions). A one-iteration loop and the inline path cannot diverge.
  • Runtime-keyed, not static: guard enabled && trip_count==1 evaluated at exec time; prefill (trip_count>1) provably takes the loop.
  • Flag OFF = zero change; parser true only for 1/true/on.
  • No capture-core changes (no provider.rs/capture.rs in diff); Scan still declines capture on both paths.
  • DRY: no model/op special-casing; one shared helper.

Non-vacuity (mutations run, all FAILED as required):

  • CPU: Mutation A (trip_count==1 -> >=1) FAILED on prefill-count tripwire; Mutation B (drop state carry) FAILED on byte-exactness.
  • CUDA (device 0): Mutation A FAILED on device-level prefill-count tripwire.

Regressions (device 0) all PASS: session-reuse #554, prefetch-WAR #544, cuda_control_flow_safety, CPU lib 105 / control_flow 23. fmt --check clean, clippy clean.

@justinchuby
justinchuby enabled auto-merge (squash) July 31, 2026 15:40
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 76.28866% with 23 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.60%. Comparing base (6e25895) to head (cbf240a).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
.../onnx-runtime-session/src/executor/control_flow.rs 78.57% 13 Missing and 5 partials ⚠️
crates/onnx-runtime-session/src/lib.rs 0.00% 3 Missing ⚠️
crates/onnx-runtime-session/src/executor/state.rs 80.00% 2 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #564      +/-   ##
==========================================
- Coverage   81.17%   80.60%   -0.58%     
==========================================
  Files         315      315              
  Lines      123528   123574      +46     
  Branches   123528   123574      +46     
==========================================
- Hits       100276    99606     -670     
- Misses      19194    19916     +722     
+ Partials     4058     4052       -6     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.67% <ø> (-0.11%) ⬇️
mlas 77.91% <ø> (ø)
offline 80.52% <76.28%> (-0.61%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-session/src/executor/state.rs 83.33% <80.00%> (-0.31%) ⬇️
crates/onnx-runtime-session/src/lib.rs 73.20% <0.00%> (-0.30%) ⬇️
.../onnx-runtime-session/src/executor/control_flow.rs 64.94% <78.57%> (+0.75%) ⬆️

... and 4 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@justinchuby
justinchuby merged commit d598c5a into main Jul 31, 2026
15 checks passed
@justinchuby
justinchuby deleted the feat/27b-scan-capture-1a branch July 31, 2026 15:49
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 gather/large_f32_threads=1-internal/131072 34.68 µs 66.96 µs +93.1%
⚠️ grammar_masking/llguidance_compute_mask/32 78.19 µs 101.55 µs +29.9%
⚠️ gather/medium_bf16_threads=1-internal/32768 2.69 µs 3.36 µs +24.9%
⚠️ tokenization/decode_tokens_per_second 6.12 ms 7.52 ms +22.8%
⚠️ logit_processing/seven_processor_chain_per_step 1.16 ms 1.39 ms +20.2%
⚠️ gather/large_f16_threads=1-internal/131072 17.43 µs 20.84 µs +19.6%
⚠️ sampling_latency/min_p_per_token 330.97 µs 395.01 µs +19.3%
⚠️ matmul/large_generic_bf16_threads=1/32x1024x1024 2.35 ms 2.79 ms +18.7%
✅ tokenization/encode_tokens_per_second 453.20 µs 507.74 µs +12.0%
✅ kv_cache/alloc_dealloc_pages 36.56 µs 40.63 µs +11.1%
✅ sampling_latency/top_k_per_token 476.89 µs 495.96 µs +4.0%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 619.75 µs 641.70 µs +3.5%
✅ sampling_latency/top_p_per_token 1.04 ms 1.06 ms +2.5%
✅ sampling_latency/greedy_per_token 3.49 µs 3.39 µs -2.8%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.91 ms 10.52 ms -3.6%
✅ add/small_bf16_threads=1-internal/1024 19.07 µs 18.27 µs -4.2%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 2.12 ms 2.02 ms -4.6%
✅ gather/medium_f16_threads=1-internal/32768 3.21 µs 3.03 µs -5.6%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 499.08 µs 454.63 µs -8.9%
✅ matmul/small_generic_f16_threads=1/1x256x256 40.17 µs 36.21 µs -9.8%
✅ matmul/medium_generic_f16_threads=1/32x512x512 41.21 µs 36.88 µs -10.5%
✅ add/medium_f16_threads=1-internal/262144 3.37 ms 2.96 ms -12.0%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 97.20 µs 84.89 µs -12.7%
✅ add/medium_bf16_threads=1-internal/262144 3.51 ms 3.02 ms -14.1%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.74 ms 2.33 ms -14.8%
🟢 matmul/small_generic_bf16_threads=8/1x256x256 42.02 µs 35.70 µs -15.0%
🟢 matmul/small_generic_f16_threads=8/1x256x256 41.06 µs 34.71 µs -15.5%
🟢 matmul/large_generic_f16_threads=8/32x1024x1024 111.36 µs 93.72 µs -15.8%
🟢 matmul/large_generic_f32_threads=8/32x1024x1024 6.77 ms 5.57 ms -17.8%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 42.72 µs 34.65 µs -18.9%
🟢 add/large_f32_threads=1-internal/4194304 55.86 ms 44.56 ms -20.2%
🟢 add/medium_f32_threads=1-internal/262144 3.55 ms 2.83 ms -20.5%
🟢 matmul/small_generic_f32_threads=1/1x256x256 60.23 µs 47.72 µs -20.8%
🟢 gather/small_f32_threads=1-internal/4096 893.8 ns 705.9 ns -21.0%
🟢 add/large_f16_threads=1-internal/4194304 55.74 ms 43.91 ms -21.2%
🟢 matmul/small_generic_f32_threads=8/1x256x256 69.28 µs 53.60 µs -22.6%
🟢 add/small_f16_threads=1-internal/1024 16.29 µs 12.60 µs -22.6%
🟢 gather/large_bf16_threads=1-internal/131072 20.02 µs 15.27 µs -23.7%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 47.83 µs 36.43 µs -23.8%
🟢 gather/medium_f32_threads=1-internal/32768 5.59 µs 4.18 µs -25.3%
🟢 add/large_bf16_threads=1-internal/4194304 64.62 ms 47.06 ms -27.2%
🟢 reduce_mean/small_f32_threads=1-internal/4096 24.36 µs 17.52 µs -28.1%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.63 ms 1.17 ms -28.1%
🟢 gather/small_f16_threads=1-internal/4096 688.7 ns 481.1 ns -30.1%
🟢 gather/small_bf16_threads=1-internal/4096 731.2 ns 491.0 ns -32.9%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 384.08 µs 249.74 µs -35.0%
🟢 add/small_f32_threads=1-internal/1024 294.9 ns 188.6 ns -36.1%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.67 ms 1.06 ms -36.5%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.81 4.29 7.62 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Jul 31, 2026
#568) (#569)

Consolidates the accumulated GAP-3 native-paged-decode decision inbox (7
notes) into the canonical `.squad/decisions.md`, clearing the drop-box.

Folds in the decisions from the 4 merged GAP-3 PRs plus the Scan-capture
deferral:
- **#565** Inc-A — native multi-component pipeline construction
- **#566** Inc-C — host present-KV mirroring → paged native decode
- **#567** Inc-D — device-resident f32 CUDA present-KV read-out → paged
- **#568** Inc-D.1 — f16 device present-KV read-out → paged
- Inc-C test-rigor fix (byte-equality oracle pattern)
- Scan-capture slice-1b **deferral** (needs per-EP device-graph
handle-keyed registry; resumption trigger recorded)
- Scan slice-1a (#564)

decisions.md 20465 → 28227 bytes (under the 30KB healthy threshold; no
archiving needed). Inbox now README-only. Docs-only change under
`.squad/`.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant