Repository navigation
profile_native: resolve genai_config decoder io so capture counters print (#87 capture-observability) - #552
Conversation
…rint
The simple `profile_native` generate path (the only path that prints
`cuda_graph: captures/replays/fallbacks` and supports `--trace`
capture-reject reasons) loaded decoders via `NativeDecodeSession::load`
with `io = None`. For onnxruntime-genai `genai_config.json` decoders whose
token/attention/position inputs are all `Int64` rank-2, shape-based I/O
inference is ambiguous, so the load failed:
cannot resolve model.io.token_input from tensor shape because 2 ports
match: [("input_ids", ...), ("attention_mask", ...)]
and capture observability was unreachable for those models.
Add a public `NativeDecodeSession::load_with_resolved_io` that resolves the
model directory's `ModelIoSpec` from an adjacent
`inference_metadata.{yaml,yml,json}` sidecar, else onnxruntime-genai
`genai_config.json` compatibility synthesis (same precedence as the engine
directory loader), and threads it into the existing io-aware load path.
`profile_native`'s simple path now calls it.
Observability/plumbing only: capture semantics are unchanged
(`graph_capture` stays auto-decided, no defaults change, no kernel touched).
After: genai_config decoders load in the simple path and print real
counters, e.g. `cuda_graph: enabled=true captures=3 replays=87 fallbacks=0`,
and `--trace` surfaces per-op `capture_status` / ARG_CAPTURE_REJECTED reasons.
Refs #87 (capture observability).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #552 +/- ##
=======================================
Coverage 80.59% 80.59%
=======================================
Files 315 315
Lines 123446 123446
Branches 123446 123446
=======================================
Hits 99489 99489
Misses 19907 19907
Partials 4050 4050
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
VERDICT: APPROVE Independent review by Lori (not the author). Scratch worktree off origin/squad/profile-native-capture-obs @ 8736a06, base origin/main. Device 2 (CUDA_VISIBLE_DEVICES=2). Scope confirmed — observability/plumbing onlygit diff origin/main...squad/profile-native-capture-obs --stat: exactly 2 files, +58/-1.
Capture semantics unchanged (checklist #1, #3 PASS)The ONLY behavioral delta is the io argument. Both paths bottom out at load_with_cuda_options_and_io with NativeDecodeCudaOptions::default():
Precedence matches engine loader (checklist #2 PASS)resolve_io_metadata_from_model_path (load.rs:22-48): inference_metadata.{yaml,yml,json} sidecar via load_metadata FIRST, else genai_config.json via genai_config_compat_metadata_from_model_path, else None (shape-based fallback). This mirrors the engine directory loader engine/load.rs:200-214 (metadata_path first, else genai_config_compat_metadata_from_model_path, else default). The model-file path threaded to the compat fn matches the engine's model_path usage, so profile_native resolves the same io spec as production. Build + fmt (checklist #4 PASS)
Non-vacuity — it actually works (checklist #5 PASS)Qwen2.5-0.5B-Instruct int4 genai_config decoder (genai_config.json present, no sidecar), device 2, --ep cuda --tokens 32. BEFORE (reverted the one line back to load(), rebuilt): AFTER (load_with_resolved_io): Counters are nonzero and capture_status surfaces; the pre-existing genai_config load error is resolved. Clean, correctly-scoped observability fix. |
…dation (#555) Consolidates 6 inbox decision notes into decisions.md (20458->20332 B, under the 20480 gate) and archives two historical wave records. Updates agent histories. Wave summary (all merged): - #544 — async fence-ordered CUDA weight page-in (#87 increment-1) + deterministic anti-regression test - #552 — profile_native capture-counter observability for genai_config decoders - #554 — native-CUDA session-reuse recurrent-state reset fix (closes #553); 27B LinearAttention gen#2+ corruption - 27B native offload A/B proof: 2.9x VRAM reduction, byte-exact output State-only change (decisions/histories/archive). Logs are gitignored. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
What
Threads model I/O resolution into
profile_native's simple generate path so CUDA-graph capture observability works for onnxruntime-genaigenai_config.jsondecoders.Adds public
NativeDecodeSession::load_with_resolved_io(path, device)which resolves the model directory'sModelIoSpec(nativeinference_metadata.{yaml,yml,json}sidecar first, elsegenai_config.jsoncompatibility synthesis — same precedence as the engine directory loader) and threads it into the existing io-aware load path.profile_native's simple path now calls it.Why
The simple
generate()path is the only path that printscuda_graph: captures/replays/fallbackscounters and supports--tracecapture-reject reasons, but it loaded viaNativeDecodeSession::loadwithio = None. For genai_config decoders whoseinput_ids/attention_mask/position_idsare allInt64rank-2, shape-based I/O inference is ambiguous and the load failed outright — so capture measurements had to rely on byte-identical A/B inference instead of real counters.Scope guardrails
Observability / plumbing only. No capture-semantics, default-on behavior, or kernel changes —
graph_capturestays auto-decided and no defaults move. The new API just resolves and threadsio.Before / After (Qwen2.5-0.5B-Instruct int4 CUDA, genai_config decoder, device 2,
ONNX_GENAI_DEVICE_KV=1)Before (
NativeDecodeSession::load,io=None):After (
load_with_resolved_io):--tracenow emits per-opcapture_status(allcaptured, zeroARG_CAPTURE_REJECTED= capture fully engaged) for the genai_config decoder.Testing
cargo build --release -p onnx-genai-bench --features "bench-native bench-ort cuda" --bin profile_native✅cargo build -p onnx-genai-engine --features native-backend(no cuda) ✅ — new API is cuda-independentcargo fmt --all --check✅Refs #87 (capture observability).