Repository navigation
test(pipeline): Inc3c real-model capture validation + general Bool value clone - #538
Closed
justinchuby wants to merge 1 commit into
Closed
justinchuby wants to merge 1 commit into
justinchuby wants to merge 1 commit into
Conversation
…ol value clone (inc3c follow-up) Inc3c #533 proved the captured step-inputs decode path engages on the synthetic tiny-gqa-embeds-cuda fixture. This follow-up validates it against the real multi-component inputs_embeds decoder class and banks the honest real-model finding: - qwen3-0.6b is single-component (input_ids) and structurally cannot exercise the capture-step-inputs flag; the 612/220/443 numbers are its single-graph CUDA-graph capture lever (a faithful launch-overhead proxy), not the flag. - The real inputs_embeds decoders (qwen3.5-0.8b hybrid, gemma-3n-e2b) are the correct class. gemma-3n's decoder is GroupQueryAttention capacity-aware KV, so it WOULD engage capture; it was blocked by (1) a Bool audio-mask clone gap and (2) required vision inputs for text-only decode. Fixes (1) generally: decode/values.rs::clone_value gains a raw-byte fallback arm (to_raw_bytes -> from_raw_bytes, bit-exact) covering Bool/Int32/etc., with a focused unit test. This unblocks the multimodal pipeline value/cache path and advances gemma-3n past the Bool error. Adds tests/gemma3n_native_cuda_capture_realmodel.rs: a forward-looking real-model capture-engagement + token-parity harness (native decoder cuda:0, counter engagement, OFF==ON parity) that skips gracefully on the remaining vision-required-input gap (2), matching the qwen35_0_8b_hybrid skip precedent. Verify: fmt --check clean; clippy x4 (default/native-backend/cuda/ cuda,native-backend) clean; full cuda,native-backend suite failing set 17, byte-identical to base, 0 regressions; synthetic engagement proof + Bool unit GREEN. Does not touch #533's reviewed code (native_decode/*). Refs #384. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #538 +/- ##
==========================================
- Coverage 80.59% 80.59% -0.01%
==========================================
Files 315 315
Lines 123259 123259
Branches 123259 123259
==========================================
- Hits 99338 99335 -3
- Misses 19890 19891 +1
- Partials 4031 4033 +2
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
✅ Benchmarks — No RegressionComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Owner
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #533 (Inc3c, merged). Validates the captured step-inputs decode path against the real multi-component
inputs_embedsdecoder class and banks the honest real-model finding. Does not touch #533's reviewed code (native_decode/*). Refs #384.Real-model finding
input_ids), loads viaEngine::from_dir, so it structurally cannot exercise the capture-step-inputs flag (that path only fires for a multi-componentinputs_embeds/Routeddecoder). The banked 612 captured / 220 eager / 443 ORT-CUDA numbers are qwen3-0.6b's single-graph CUDA-graph capture lever — a faithful launch-overhead proxy, not the flag path.inputs_embedsdecoders (qwen3.5-0.8b hybrid, gemma-3n-e2b) are the correct class. gemma-3n's decoder isGroupQueryAttentioncapacity-aware KV →graph_enabled=true, so it WOULD engage capture (same class as the synthetic fixture and the real 35B-A3B target). It was gated by two independent, non-Inc3c blockers:decode/values.rs::clone_valuehard-errored onBool. Added a general raw-byte fallback (to_raw_bytes→from_raw_bytes, bit-exact) coveringBool/Int32/etc., with a focused unit test. Advances gemma-3n past the Bool error. Quick win.OneHotrejects synthetic patches (Depth is negative). Needs a real image or optional-modality skip.Artifacts
src/decode/values.rs— generalclone_valueraw-byte fallback + unit test.tests/gemma3n_native_cuda_capture_realmodel.rs— forward-looking real-model capture-engagement + token-parity harness (native decodercuda:0, counter engagement, OFF==ON parity); skips gracefully on the vision-input gap (matches theqwen35_0_8b_hybridprecedent). Becomes a live proof once a real image is supplied / vision is optional..squad/decisions/inbox/mary-inc3c-realmodel-capture.md— full finding + default-on recommendation.Default-on recommendation
Safe to default-on for engaging models (byte-identical when it declines; token-parity when it engages; graceful eager fallback everywhere). Blocker to recommending it now: no GREEN real-weights e2e capture number yet (fixture-only), due to the two unrelated loader/modality gaps — not the optimization. Keep default-off until one real
inputs_embedsmodel runs the flag e2e.Verify
cargo fmt --checkclean; clippy ×4 (default/native-backend/cuda/cuda,native-backend) clean.cargo test -p onnx-genai-engine --features cuda,native-backend --no-fail-fast: failing set 17, byte-identical to base, 0 regressions.native_cuda_captured_step_inputs_parity, tokens[0,5,6,7], captured 0→3) + Bool unit GREEN.