Skip to content

profile_native: resolve genai_config decoder io so capture counters print (#87 capture-observability) - #552

Merged
justinchuby merged 1 commit into
mainfrom
squad/profile-native-capture-obs
Jul 31, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/profile-native-capture-obs

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What

Threads model I/O resolution into profile_native's simple generate path so CUDA-graph capture observability works for onnxruntime-genai genai_config.json decoders.

Adds public NativeDecodeSession::load_with_resolved_io(path, device) which resolves the model directory's ModelIoSpec (native inference_metadata.{yaml,yml,json} sidecar first, else genai_config.json compatibility synthesis — same precedence as the engine directory loader) and threads it into the existing io-aware load path. profile_native's simple path now calls it.

Why

The simple generate() path is the only path that prints cuda_graph: captures/replays/fallbacks counters and supports --trace capture-reject reasons, but it loaded via NativeDecodeSession::load with io = None. For genai_config decoders whose input_ids/attention_mask/position_ids are all Int64 rank-2, shape-based I/O inference is ambiguous and the load failed outright — so capture measurements had to rely on byte-identical A/B inference instead of real counters.

Scope guardrails

Observability / plumbing only. No capture-semantics, default-on behavior, or kernel changes — graph_capture stays auto-decided and no defaults move. The new API just resolves and threads io.

Before / After (Qwen2.5-0.5B-Instruct int4 CUDA, genai_config decoder, device 2, ONNX_GENAI_DEVICE_KV=1)

Before (NativeDecodeSession::load, io=None):

Error: load native decoder .../model.onnx
Caused by:
    cannot resolve model.io.token_input from tensor shape because 2 ports match:
    [("input_ids", Int64, [-1, -1]), ("attention_mask", Int64, [-1, -1])];
    declare the exact graph port in model.io.token_input

After (load_with_resolved_io):

profile_native: model=.../model.onnx ep=Cuda layers=24 ...
throughput: 626.79 tok/s, 1.595 ms/step
cuda_graph: enabled=true captures=3 replays=87 fallbacks=0
cuda_graph_measured: captures=2 replays=58 fallbacks=0

--trace now emits per-op capture_status (all captured, zero ARG_CAPTURE_REJECTED = capture fully engaged) for the genai_config decoder.

Testing

  • cargo build --release -p onnx-genai-bench --features "bench-native bench-ort cuda" --bin profile_native ✅
  • cargo build -p onnx-genai-engine --features native-backend (no cuda) ✅ — new API is cuda-independent
  • cargo fmt --all --check ✅
  • Manual before/after runs above.

Refs #87 (capture observability).

…rint

The simple `profile_native` generate path (the only path that prints
`cuda_graph: captures/replays/fallbacks` and supports `--trace`
capture-reject reasons) loaded decoders via `NativeDecodeSession::load`
with `io = None`. For onnxruntime-genai `genai_config.json` decoders whose
token/attention/position inputs are all `Int64` rank-2, shape-based I/O
inference is ambiguous, so the load failed:

    cannot resolve model.io.token_input from tensor shape because 2 ports
    match: [("input_ids", ...), ("attention_mask", ...)]

and capture observability was unreachable for those models.

Add a public `NativeDecodeSession::load_with_resolved_io` that resolves the
model directory's `ModelIoSpec` from an adjacent
`inference_metadata.{yaml,yml,json}` sidecar, else onnxruntime-genai
`genai_config.json` compatibility synthesis (same precedence as the engine
directory loader), and threads it into the existing io-aware load path.
`profile_native`'s simple path now calls it.

Observability/plumbing only: capture semantics are unchanged
(`graph_capture` stays auto-decided, no defaults change, no kernel touched).

After: genai_config decoders load in the simple path and print real
counters, e.g. `cuda_graph: enabled=true captures=3 replays=87 fallbacks=0`,
and `--trace` surfaces per-op `capture_status` / ARG_CAPTURE_REJECTED reasons.

Refs #87 (capture observability).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.59%. Comparing base (080735a) to head (8736a06).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main     #552   +/-   ##
=======================================
  Coverage   80.59%   80.59%           
=======================================
  Files         315      315           
  Lines      123446   123446           
  Branches   123446   123446           
=======================================
  Hits        99489    99489           
  Misses      19907    19907           
  Partials     4050     4050           
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.78% <ø> (ø)
mlas 77.91% <ø> (ø)
offline 80.51% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 4.01 ms 7.32 ms +82.7%
⚠️ gather/large_f16_threads=1-internal/131072 12.47 µs 15.15 µs +21.5%
⚠️ gather/large_bf16_threads=1-internal/131072 12.83 µs 15.47 µs +20.5%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 82.31 µs 98.08 µs +19.2%
✅ add/small_bf16_threads=1-internal/1024 12.98 µs 14.66 µs +12.9%
✅ add/small_f16_threads=1-internal/1024 13.22 µs 14.85 µs +12.3%
✅ gather/large_f32_threads=1-internal/131072 35.05 µs 39.19 µs +11.8%
✅ add/medium_f16_threads=1-internal/262144 2.62 ms 2.89 ms +10.2%
✅ add/medium_bf16_threads=1-internal/262144 2.66 ms 2.90 ms +9.2%
✅ add/medium_f32_threads=1-internal/262144 2.60 ms 2.75 ms +5.9%
✅ matmul/medium_generic_f16_threads=8/32x512x512 41.80 µs 44.19 µs +5.7%
✅ add/large_f32_threads=1-internal/4194304 41.09 ms 43.40 ms +5.6%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.44 ms 2.57 ms +5.3%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 565.04 µs 592.09 µs +4.8%
✅ add/small_f32_threads=1-internal/1024 223.6 ns 233.1 ns +4.3%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.71 µs 16.37 µs +4.3%
✅ matmul/medium_generic_f16_threads=1/32x512x512 36.23 µs 37.53 µs +3.6%
✅ add/large_f16_threads=1-internal/4194304 41.89 ms 43.34 ms +3.5%
✅ matmul/small_generic_f32_threads=1/1x256x256 41.90 µs 43.33 µs +3.4%
✅ reduce_mean/medium_f32_threads=1-internal/65536 251.63 µs 259.42 µs +3.1%
✅ gather/small_f16_threads=1-internal/4096 488.5 ns 502.4 ns +2.8%
✅ gather/medium_f32_threads=1-internal/32768 3.91 µs 4.02 µs +2.6%
✅ gather/small_f32_threads=1-internal/4096 652.9 ns 665.6 ns +1.9%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.09 ms 2.12 ms +1.6%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.00 ms 1.01 ms +1.4%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 664.16 µs 672.76 µs +1.3%
✅ add/large_bf16_threads=1-internal/4194304 43.02 ms 43.47 ms +1.0%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 2.32 ms 2.34 ms +1.0%
✅ gather/medium_f16_threads=1-internal/32768 2.34 µs 2.35 µs +0.3%
✅ matmul/small_generic_bf16_threads=8/1x256x256 38.91 µs 38.98 µs +0.2%
✅ gather/small_bf16_threads=1-internal/4096 479.3 ns 479.7 ns +0.1%
✅ gather/medium_bf16_threads=1-internal/32768 2.36 µs 2.35 µs -0.5%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.79 ms 1.77 ms -0.8%
✅ kv_cache/alloc_dealloc_pages 40.95 µs 40.57 µs -0.9%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.70 ms 9.56 ms -1.5%
✅ matmul/small_generic_f16_threads=1/1x256x256 36.30 µs 35.72 µs -1.6%
✅ matmul/small_generic_f16_threads=8/1x256x256 38.30 µs 37.60 µs -1.8%
✅ logit_processing/seven_processor_chain_per_step 1.27 ms 1.24 ms -2.5%
✅ sampling_latency/min_p_per_token 373.05 µs 360.05 µs -3.5%
✅ tokenization/encode_tokens_per_second 405.80 µs 391.14 µs -3.6%
✅ sampling_latency/greedy_per_token 3.46 µs 3.33 µs -3.7%
✅ tokenization/decode_tokens_per_second 6.69 ms 6.36 ms -4.9%
✅ sampling_latency/top_p_per_token 1.09 ms 1.03 ms -6.0%
✅ sampling_latency/top_k_per_token 532.58 µs 498.99 µs -6.3%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 106.19 µs 98.78 µs -7.0%
✅ grammar_masking/llguidance_compute_mask/32 83.94 µs 77.87 µs -7.2%
✅ matmul/small_generic_f32_threads=8/1x256x256 51.96 µs 47.41 µs -8.8%
✅ matmul/small_generic_bf16_threads=1/1x256x256 39.94 µs 36.36 µs -9.0%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.87 3.67 4.05 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review by Lori (not the author). Scratch worktree off origin/squad/profile-native-capture-obs @ 8736a06, base origin/main. Device 2 (CUDA_VISIBLE_DEVICES=2).

Scope confirmed — observability/plumbing only

git diff origin/main...squad/profile-native-capture-obs --stat: exactly 2 files, +58/-1.

  • crates/onnx-genai-engine/src/native_decode/load.rs (+57): new pub load_with_resolved_io + private resolve_io_metadata_from_model_path.
  • crates/onnx-genai-bench/src/bin/profile_native.rs (1 line): load -> load_with_resolved_io.
    No pipeline/, ep-cuda offload, refactor/, decode_contract.rs, or model-package/LoRA files touched (checklist Future: Browser deployment (WASM + WebGPU) #6 PASS).

Capture semantics unchanged (checklist #1, #3 PASS)

The ONLY behavioral delta is the io argument. Both paths bottom out at load_with_cuda_options_and_io with NativeDecodeCudaOptions::default():

  • Old load (mod.rs:122-124) -> load_with_cuda_options -> load_with_cuda_options_and_io(.., default(), io=None).
  • New load_with_resolved_io (load.rs:62-73) -> load_with_cuda_options_and_io(.., default(), io=Some(spec)).
    graph_capture stays None (auto-decided); metadata_max_len is still filled by native_metadata_max_len_from_model_path (load.rs:223-224). No default-on/off forcing, no kernel change. Delta is io: None -> Some(resolved spec).

Precedence matches engine loader (checklist #2 PASS)

resolve_io_metadata_from_model_path (load.rs:22-48): inference_metadata.{yaml,yml,json} sidecar via load_metadata FIRST, else genai_config.json via genai_config_compat_metadata_from_model_path, else None (shape-based fallback). This mirrors the engine directory loader engine/load.rs:200-214 (metadata_path first, else genai_config_compat_metadata_from_model_path, else default). The model-file path threaded to the compat fn matches the engine's model_path usage, so profile_native resolves the same io spec as production.

Build + fmt (checklist #4 PASS)

  • cargo fmt --all --check: clean (exit 0).
  • cargo build --release -p onnx-genai-engine --no-default-features --features native-backend: Finished (no cuda).
  • cargo build --release -p onnx-genai-bench --bin profile_native --features bench-native,cuda: Finished.

Non-vacuity — it actually works (checklist #5 PASS)

Qwen2.5-0.5B-Instruct int4 genai_config decoder (genai_config.json present, no sidecar), device 2, --ep cuda --tokens 32.

BEFORE (reverted the one line back to load(), rebuilt):
Error: load native decoder .../model.onnx
Caused by: cannot resolve model.io.token_input from tensor shape because 2 ports match: [("input_ids", Int64, [-1, -1]), ("attention_mask", Int64, [-1, -1])]; declare the exact graph port in model.io.token_input

AFTER (load_with_resolved_io):
cuda_graph: enabled=true captures=3 replays=90 fallbacks=0
cuda_graph_measured: captures=1 replays=30 fallbacks=0
device_kv_measured: h2d_calls=0 h2d_bytes=0 d2h_calls=0 d2h_bytes=0
throughput: 937 tok/s
--trace timeline surfaces capture_status (179 "captured" spans) plus kernel_variant annotations.

Counters are nonzero and capture_status surfaces; the pre-existing genai_config load error is resolved. Clean, correctly-scoped observability fix.

@justinchuby
justinchuby marked this pull request as ready for review July 31, 2026 06:26
@justinchuby
justinchuby merged commit fa1afed into main Jul 31, 2026
14 checks passed
@justinchuby
justinchuby deleted the squad/profile-native-capture-obs branch July 31, 2026 06:26
justinchuby added a commit that referenced this pull request Jul 31, 2026
…dation (#555)

Consolidates 6 inbox decision notes into decisions.md (20458->20332 B,
under the 20480 gate) and archives two historical wave records. Updates
agent histories.

Wave summary (all merged):
- #544 — async fence-ordered CUDA weight page-in (#87 increment-1) +
deterministic anti-regression test
- #552 — profile_native capture-counter observability for genai_config
decoders
- #554 — native-CUDA session-reuse recurrent-state reset fix (closes
#553); 27B LinearAttention gen#2+ corruption
- 27B native offload A/B proof: 2.9x VRAM reduction, byte-exact output

State-only change (decisions/histories/archive). Logs are gitignored.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant