Skip to content

test(pipeline): Inc3c real-model capture validation + general Bool value clone - #538

Closed
justinchuby wants to merge 1 commit into
mainfrom
squad/native-multi-component-pipeline-inc3c
Closed

justinchuby wants to merge 1 commit into
mainfrom
squad/native-multi-component-pipeline-inc3c

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Follow-up to #533 (Inc3c, merged). Validates the captured step-inputs decode path against the real multi-component inputs_embeds decoder class and banks the honest real-model finding. Does not touch #533's reviewed code (native_decode/*). Refs #384.

Real-model finding

  • qwen3-0.6b is the wrong class: single-component (input_ids), loads via Engine::from_dir, so it structurally cannot exercise the capture-step-inputs flag (that path only fires for a multi-component inputs_embeds/Routed decoder). The banked 612 captured / 220 eager / 443 ORT-CUDA numbers are qwen3-0.6b's single-graph CUDA-graph capture lever — a faithful launch-overhead proxy, not the flag path.
  • The real inputs_embeds decoders (qwen3.5-0.8b hybrid, gemma-3n-e2b) are the correct class. gemma-3n's decoder is GroupQueryAttention capacity-aware KV → graph_enabled=true, so it WOULD engage capture (same class as the synthetic fixture and the real 35B-A3B target). It was gated by two independent, non-Inc3c blockers:
    1. Bool audio-mask (FIXED here, general) — decode/values.rs::clone_value hard-errored on Bool. Added a general raw-byte fallback (to_raw_bytes→from_raw_bytes, bit-exact) covering Bool/Int32/etc., with a focused unit test. Advances gemma-3n past the Bool error. Quick win.
    2. Vision required-input (a slice, not fixed) — gemma-3n eagerly runs the vision encoder for text-only prompts; its pooler OneHot rejects synthetic patches (Depth is negative). Needs a real image or optional-modality skip.

Artifacts

  • src/decode/values.rs — general clone_value raw-byte fallback + unit test.
  • tests/gemma3n_native_cuda_capture_realmodel.rs — forward-looking real-model capture-engagement + token-parity harness (native decoder cuda:0, counter engagement, OFF==ON parity); skips gracefully on the vision-input gap (matches the qwen35_0_8b_hybrid precedent). Becomes a live proof once a real image is supplied / vision is optional.
  • .squad/decisions/inbox/mary-inc3c-realmodel-capture.md — full finding + default-on recommendation.

Default-on recommendation

Safe to default-on for engaging models (byte-identical when it declines; token-parity when it engages; graceful eager fallback everywhere). Blocker to recommending it now: no GREEN real-weights e2e capture number yet (fixture-only), due to the two unrelated loader/modality gaps — not the optimization. Keep default-off until one real inputs_embeds model runs the flag e2e.

Verify

  • cargo fmt --check clean; clippy ×4 (default/native-backend/cuda/cuda,native-backend) clean.
  • Full cargo test -p onnx-genai-engine --features cuda,native-backend --no-fail-fast: failing set 17, byte-identical to base, 0 regressions.
  • Synthetic engagement proof (native_cuda_captured_step_inputs_parity, tokens [0,5,6,7], captured 0→3) + Bool unit GREEN.

…ol value clone (inc3c follow-up)

Inc3c #533 proved the captured step-inputs decode path engages on the
synthetic tiny-gqa-embeds-cuda fixture. This follow-up validates it against
the real multi-component inputs_embeds decoder class and banks the honest
real-model finding:

- qwen3-0.6b is single-component (input_ids) and structurally cannot exercise
  the capture-step-inputs flag; the 612/220/443 numbers are its single-graph
  CUDA-graph capture lever (a faithful launch-overhead proxy), not the flag.
- The real inputs_embeds decoders (qwen3.5-0.8b hybrid, gemma-3n-e2b) are the
  correct class. gemma-3n's decoder is GroupQueryAttention capacity-aware KV, so
  it WOULD engage capture; it was blocked by (1) a Bool audio-mask clone gap and
  (2) required vision inputs for text-only decode.

Fixes (1) generally: decode/values.rs::clone_value gains a raw-byte fallback
arm (to_raw_bytes -> from_raw_bytes, bit-exact) covering Bool/Int32/etc., with
a focused unit test. This unblocks the multimodal pipeline value/cache path and
advances gemma-3n past the Bool error.

Adds tests/gemma3n_native_cuda_capture_realmodel.rs: a forward-looking
real-model capture-engagement + token-parity harness (native decoder cuda:0,
counter engagement, OFF==ON parity) that skips gracefully on the remaining
vision-required-input gap (2), matching the qwen35_0_8b_hybrid skip precedent.

Verify: fmt --check clean; clippy x4 (default/native-backend/cuda/
cuda,native-backend) clean; full cuda,native-backend suite failing set 17,
byte-identical to base, 0 regressions; synthetic engagement proof + Bool unit
GREEN. Does not touch #533's reviewed code (native_decode/*).

Refs #384.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.59%. Comparing base (34d184d) to head (ab49caa).

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #538      +/-   ##
==========================================
- Coverage   80.59%   80.59%   -0.01%     
==========================================
  Files         315      315              
  Lines      123259   123259              
  Branches   123259   123259              
==========================================
- Hits        99338    99335       -3     
- Misses      19890    19891       +1     
- Partials     4031     4033       +2     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.78% <ø> (ø)
mlas 77.91% <ø> (ø)
offline 80.51% <ø> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

✅ Benchmarks — No Regression

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.96 ms 2.09 ms +6.7%
✅ gather/large_f32_threads=1-internal/131072 28.91 µs 30.75 µs +6.4%
✅ sampling_latency/greedy_per_token 3.26 µs 3.38 µs +3.9%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.57 µs 16.15 µs +3.8%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 85.56 µs 87.76 µs +2.6%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.79 ms 3.85 ms +1.7%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 380.58 µs 385.18 µs +1.2%
✅ matmul/medium_generic_f16_threads=8/32x512x512 31.19 µs 31.48 µs +0.9%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.32 ms 2.34 ms +0.9%
✅ matmul/small_generic_f16_threads=1/1x256x256 30.62 µs 30.81 µs +0.6%
✅ logit_processing/seven_processor_chain_per_step 1.18 ms 1.19 ms +0.5%
✅ grammar_masking/llguidance_compute_mask/32 75.65 µs 75.93 µs +0.4%
✅ gather/medium_f32_threads=1-internal/32768 3.81 µs 3.82 µs +0.3%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.27 ms 1.27 ms +0.1%
✅ add/medium_f32_threads=1-internal/262144 2.70 ms 2.70 ms +0.0%
✅ matmul/small_generic_bf16_threads=1/1x256x256 31.90 µs 31.86 µs -0.1%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 80.32 µs 80.05 µs -0.3%
✅ gather/large_bf16_threads=1-internal/131072 12.14 µs 12.10 µs -0.4%
✅ matmul/small_generic_f32_threads=1/1x256x256 37.09 µs 36.86 µs -0.6%
✅ matmul/small_generic_bf16_threads=8/1x256x256 32.48 µs 32.23 µs -0.8%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 531.31 µs 527.13 µs -0.8%
✅ tokenization/decode_tokens_per_second 6.38 ms 6.33 ms -0.9%
✅ sampling_latency/top_k_per_token 478.63 µs 473.95 µs -1.0%
✅ matmul/medium_generic_f16_threads=1/32x512x512 30.75 µs 30.43 µs -1.0%
✅ gather/small_bf16_threads=1-internal/4096 473.2 ns 467.8 ns -1.1%
✅ gather/medium_f16_threads=1-internal/32768 2.35 µs 2.32 µs -1.2%
✅ tokenization/encode_tokens_per_second 382.97 µs 378.27 µs -1.2%
✅ kv_cache/alloc_dealloc_pages 37.21 µs 36.62 µs -1.6%
✅ gather/medium_bf16_threads=1-internal/32768 2.35 µs 2.30 µs -1.7%
✅ matmul/small_generic_f32_threads=8/1x256x256 34.35 µs 33.47 µs -2.5%
✅ add/large_bf16_threads=1-internal/4194304 43.53 ms 42.22 ms -3.0%
✅ gather/small_f16_threads=1-internal/4096 486.4 ns 469.4 ns -3.5%
✅ sampling_latency/top_p_per_token 1.03 ms 991.84 µs -3.9%
✅ gather/small_f32_threads=1-internal/4096 673.9 ns 647.3 ns -3.9%
✅ reduce_mean/medium_f32_threads=1-internal/65536 259.04 µs 247.84 µs -4.3%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.91 ms 9.42 ms -4.9%
✅ matmul/small_generic_f16_threads=8/1x256x256 31.53 µs 29.85 µs -5.3%
✅ matmul/medium_generic_f32_threads=8/32x512x512 970.98 µs 918.73 µs -5.4%
✅ sampling_latency/min_p_per_token 361.52 µs 340.42 µs -5.8%
✅ add/medium_bf16_threads=1-internal/262144 2.84 ms 2.65 ms -6.8%
✅ add/large_f16_threads=1-internal/4194304 44.47 ms 41.21 ms -7.3%
✅ add/small_bf16_threads=1-internal/1024 14.15 µs 12.86 µs -9.1%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.09 ms 991.47 µs -9.3%
✅ add/small_f16_threads=1-internal/1024 14.47 µs 12.96 µs -10.4%
✅ add/medium_f16_threads=1-internal/262144 2.88 ms 2.57 ms -10.7%
✅ add/large_f32_threads=1-internal/4194304 46.57 ms 40.39 ms -13.3%
🟢 add/small_f32_threads=1-internal/1024 234.5 ns 193.5 ns -17.5%
🟢 gather/large_f16_threads=1-internal/131072 16.29 µs 12.20 µs -25.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.50 3.67 4.33 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

Superseded by #541, which is the same validation work on a cleanly-named standalone branch (squad/inc3c-realmodel-capture-validation) off fresh main. This branch name collided with the merged #533 branch; closing to avoid confusion.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant