Repository navigation
perf(native-decode): default-on CUDA-graph capture for multi-component step-inputs - #571
Conversation
…t step-inputs Invert ONNX_GENAI_NATIVE_DECODER_CAPTURE_STEP_INPUTS from opt-in to opt-out so the Inc3c (#384) captured per-step-input decode path is enabled by default for capture-eligible multi-component / routed native CUDA decoders (inputs_embeds + routed ports: gemma-3n / gemma4-e2b, 35B-A3B GQA layers). The routed decode step now reuses the persistent-binding run_one_token graph instead of the eager owned uploads that forfeited CUDA-graph capture — the root cause of the ~6.5x gemma4-e2b native-vs-ORT decode gap. Only the step-inputs (multi-component) path is affected; monolithic single-component decoders already captured by default via the token-id path. No capture-core / plan_capture_region / standard_attention / GAP-3 edits. Structural eligibility gates (graph_enabled, non-empty captured_step_inputs) are unchanged, so ineligible decoders (growing-logical bindings, GLM-style mask-exposed indexers, recurrent Scan/LinearAttention hybrids) auto-decline to eager — no silent-wrong. `...CAPTURE_STEP_INPUTS=0` remains as an opt-out escape hatch. Correctness (GPU): native_cuda_captured_step_inputs_parity proves the default (no env) engages capture (captured_decodes=3) and is byte-identical to the env=0 eager baseline (tokens [0,5,6,7]); a new unit module covers the opt-out parse; the qwen3-0.6b decline test now guards default-on non-engagement for a single-component decoder. Refs #384 (Inc3c lineage). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #571 +/- ##
==========================================
- Coverage 81.31% 80.60% -0.72%
==========================================
Files 315 315
Lines 123574 123574
Branches 123574 123574
==========================================
- Hits 100489 99606 -883
- Misses 19026 19916 +890
+ Partials 4059 4052 -7
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
VERDICT: APPROVE Independent review by Harry (author Cohaagen locked out). Re-ran everything from PR #571 head e1422a2 on GPU 6; did not trust the author report. This is a regression-sensitive perf change, so the bar was: prove it cannot change decode output and cannot silently mis-capture. 1. Diff scope — CLEAN
2. Eligibility-gate trace — nothing can silently mis-capture
3. Tests re-run (results)
4. Mutation testing — BOTH mutations FAIL the suite (teeth confirmed)
5. Oracle-strength caveat (non-blocking)The parity oracle is full-token-stream argmax equality over the whole rollout (errors compound across steps), NOT raw per-logit byte equality — the position-offset no-op shows a sub-argmax-threshold perturbation would not be caught by tokens alone. For a CUDA-graph capture regression the realistic failure modes (stale/wrong persistent bindings, no-capture, bad replay) produce materially wrong logits that DO flip argmax (demonstrated by Mutation B) and are additionally caught by the device capture-error poll + non-finite guard. I judge the oracle non-vacuous and adequate. Recommendation for a follow-up (NOT a merge blocker): add a logit L-inf/allclose assertion between default and eager to harden against sub-threshold drift, and re-enable a real-model multi-component engagement test that doesn't depend on the gemma vision path. SummaryScope is exactly the default-flag inversion + eligibility-preserving plumbing + tests + docs. Structural eligibility gates are untouched, so default-on captures only already-capture-safe decoders and ineligible decoders still decline to the byte-identical eager fallback (proven on the real qwen3 model). Parity holds and both negative-control mutations fail the suite. Approving. |
## What Flip `ONNX_GENAI_WEIGHT_OFFLOAD_ASYNC_PAGEIN` from default-ON to **opt-IN**. Unset/falsey now uses the synchronous device page-in (new default); a truthy value (`1`/`true`/`yes`/`on`) opts into the asynchronous fence-ordered page-in. Only the default changes — the async path is fully preserved behind the flag. ## Why — measured A/B (#544 follow-up) Async page-in net-regresses in the eviction/thrash regime. qwen3-0.6b-int4, native CUDA, weight-offload engaged, 96 MiB device budget (every admit evicts): | Config | tok/s | |---|---| | async page-in ON | 12.16 | | async page-in OFF (sync) | **15.84** | Sync is ~1.30x faster. Per-page-in tax breakdown (96 MiB, async): materialize 791 ms + pinned-staging alloc/copy 792 ms co-dominate (~48% each); raw H2D 46 ms (~3%); eviction drain 15 ms; fence wait 7 ms. The transfer async tries to overlap is ~3% of the cost, and when every admit evicts the eviction compute-stream drain re-serializes — so async cannot hide anything and only adds a non-overlappable pinned-staging alloc. Async becomes a net win only once a warm-host materialize cache lands; it stays available via `=1`. ## Correctness (regression-sensitive) - **Byte-exact preserved** (weight_paging section 9): offloaded == resident token stream unchanged. Verified on `weight_offload_native_cuda_e2e` with the NEW sync default — tokens byte-identical to resident baseline, page_ins=12544, evictions=12541 (non-vacuous). - WAR / eviction-drain safety and fence-ordering primitives **untouched**. The async fence anti-regression GPU test (`async_pagein_fence_orders_weight_page_in_consumer`) still passes and still guards the async path. - No capture / GAP-3 interaction — dynamic page-in is outside any captured region. ## Tests - Unit `async_pagein_env_is_opt_in`: `None -> false`, truthy spellings -> true, falsey/garbage -> false (non-vacuous both directions). - `device_policy_defaults_to_disabled` extended to assert the default policy is sync. - e2e asserts the resolved `from_env()` policy is sync by default AND offloaded == resident on real int4 GPU. - `cargo fmt --all --check` clean. ## Blast radius Flag default + tests + docs only. Pager internals, capture (#571), and GAP-3 untouched. Escape hatch: `ONNX_GENAI_WEIGHT_OFFLOAD_ASYNC_PAGEIN=1` restores async. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
What & why
Flips the native decoder multi-component / routed step-inputs CUDA-graph capture from opt-in to default-on, the high-performance win authorized by Justin ("确保高性能").
The captured per-step-input decode path (Inc3c, #384) writes the one-token
inputs_embeds/routed tensors into persistent device bindings and reuses therun_one_tokencaptured graph — instead of the eager owned-input uploads that forfeited CUDA-graph capture on every routed decode step. That eager fallback was the root cause of the ~6.5× gemma4-e2b native-vs-ORT decode gap. It was already built and parity-tested but gated OFF; this PR simply makes it the default.ONNX_GENAI_NATIVE_DECODER_CAPTURE_STEP_INPUTSis inverted opt-in → opt-out:0/false/no/off→ force the eager owned path (escape hatch).Scope: this only affects the step-inputs (multi-component / routed) decode path. Monolithic single-component decoders (qwen2.5/qwen3) already captured by default via the token-id
run_one_tokenpath — unchanged. No capture-core /plan_capture_region/standard_attention/ GAP-3 KV edits.Regression safety (the point)
Structural eligibility gates are unchanged (
graph_enabled, non-emptycaptured_step_inputs), so ineligible decoders — growing-logical bindings, GLM-style mask-exposed indexers, recurrent Scan/LinearAttention hybrids — auto-decline to eager. No silent-wrong. Default-on output is byte-identical to the eager path.Tests (GPU, CUDA_VISIBLE_DEVICES=2)
native_cuda_captured_step_inputs_parity(--test-threads=1):tokens=[0,5,6,7] default_captured_decodes=3 opt_in_captured_decodes=3 opt_out_captured_decodes=0— the default (no env) engages capture and is byte-identical to the=0opt-out eager baseline; the counter makes it non-vacuous (a silent no-capture fallback fails).capture_step_inputs_gate_testsunit module: default-on, falsy set opts out, truthy/unknown stay on.qwen3_0_6b_capture_step_inputs_decline(--ignored): single-component decoder declines under the default (no env) — default-on never mis-engages an ineligible decoder.gemma3n_native_cuda_capture_realmodel(--ignored): default engages + token parity +=0opt-out declines; graceful-skips on the orthogonal vision-export blocker.cargo fmt --all --checkclean.Caveats
vision_encoder OneHot Depth is negativeeven text-only; the pipeline forces vision load). The real-model harness is ready the moment a clean export / text-only-skip exists.Decision note:
.squad/decisions/inbox/cohaagen-capture-default-on.md. Scope note:.squad/decisions/inbox/cohaagen-perfcap-scope.md.Refs #384 (Inc3c lineage).