Repository navigation
feat(cuda-ep): CausalConvWithState + declare GatherBlockQuantized (#67) - #480
Conversation
…zed (#67) Data-driven CUDA op-coverage batch for issue #67. A placement audit over the real target decode models (Qwen2.5 0.5b/1.5b/7b, Phi-4-mini, Qwen3.6-27B, Qwen3.5-35B-A3B) showed the classic transformer decode path already places 100% of EP-placeable nodes on CUDA — the only uncovered types are the executor-handled control-flow ops (If/Loop/Scan). The genuine remaining gaps are the Qwen3.5 hybrid (Mamba + linear-attention) family. This batch: * Adds a new NVRTC CUDA kernel for `com.microsoft::CausalConvWithState` (depthwise causal 1-D short-conv with rolling state), covering fp32/fp16/bf16 with f32 accumulation to match the CPU EP / ORT numerics, an optional bias and past_state, optional present_state output, and none/silu(swish) activation. Registers it, adds a claim gate (ndim=1, activation, float dtypes), and a dedicated GPU parity suite vs the CPU EP oracle. Unblocks the Qwen3.5 hybrid short-conv on CUDA (18-24 nodes/model across qwen3.5 0.8b/2b/9b text decoders). * Declares the already-registered `com.microsoft::GatherBlockQuantized` in CUDA_COVERED_OPS and adds a dedicated GPU parity suite (bits 8/4, fp32+fp16 scales/output, int32/int64 indices), closing a registered-but-undeclared coverage-of-coverage hole. It is the whole compute of the int4 embedding component of the Qwen3.5 split models. Coverage-of-coverage guards pass (163 advertised op names, 168 registry pairs). Honest follow-ups: com.microsoft::LinearAttention (pairs with CausalConv to fully land Qwen3.5 hybrid decode), registering RotaryEmbedding for the com.microsoft domain, and a Bool-input NonZero path. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #480 +/- ##
==========================================
+ Coverage 80.56% 80.57% +0.01%
==========================================
Files 314 315 +1
Lines 122637 122793 +156
Branches 122637 122793 +156
==========================================
+ Hits 98804 98944 +140
- Misses 19814 19826 +12
- Partials 4019 4023 +4
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
|
| Status | Scenario | Base | PR | Change |
|---|---|---|---|---|
gather/large_f32_threads=1-internal/131072 |
34.35 µs | 42.42 µs | +23.5% | |
matmul/small_generic_f32_threads=8/1x256x256 |
39.29 µs | 48.22 µs | +22.7% | |
matmul/small_generic_f32_threads=1/1x256x256 |
36.78 µs | 42.63 µs | +15.9% | |
| ✅ | matmul/small_generic_f16_threads=8/1x256x256 |
32.59 µs | 37.03 µs | +13.6% |
| ✅ | gather/large_bf16_threads=1-internal/131072 |
12.16 µs | 13.56 µs | +11.5% |
| ✅ | gather/large_f16_threads=1-internal/131072 |
13.71 µs | 14.68 µs | +7.1% |
| ✅ | gather/medium_f32_threads=1-internal/32768 |
4.44 µs | 4.63 µs | +4.5% |
| ✅ | matmul/small_generic_f16_threads=1/1x256x256 |
36.10 µs | 37.16 µs | +3.0% |
| ✅ | gather/medium_bf16_threads=1-internal/32768 |
2.78 µs | 2.82 µs | +1.4% |
| ✅ | matmul/large_generic_f16_threads=1/32x1024x1024 |
81.07 µs | 80.91 µs | -0.2% |
| ✅ | matmul/medium_generic_f32_threads=8/32x512x512 |
951.23 µs | 948.60 µs | -0.3% |
| ✅ | matmul/medium_generic_f32_threads=1/32x512x512 |
2.33 ms | 2.30 ms | -1.2% |
| ✅ | matmul/small_generic_bf16_threads=8/1x256x256 |
32.48 µs | 31.96 µs | -1.6% |
| ✅ | matmul/large_generic_bf16_threads=1/32x1024x1024 |
1.95 ms | 1.90 ms | -2.8% |
| ✅ | reduce_mean/small_f32_threads=1-internal/4096 |
16.74 µs | 16.22 µs | -3.1% |
| ✅ | reduce_mean/medium_f32_threads=1-internal/65536 |
278.84 µs | 267.41 µs | -4.1% |
| ✅ | matmul/medium_generic_bf16_threads=1/32x512x512 |
523.22 µs | 500.79 µs | -4.3% |
| ✅ | matmul/medium_generic_f16_threads=1/32x512x512 |
30.84 µs | 29.50 µs | -4.3% |
| ✅ | matmul/large_generic_f16_threads=8/32x1024x1024 |
86.34 µs | 82.39 µs | -4.6% |
| ✅ | add/small_bf16_threads=1-internal/1024 |
14.82 µs | 14.12 µs | -4.7% |
| ✅ | matmul/large_generic_bf16_threads=8/32x1024x1024 |
1.39 ms | 1.32 ms | -5.1% |
| ✅ | add/large_bf16_threads=1-internal/4194304 |
47.73 ms | 45.28 ms | -5.1% |
| ✅ | matmul/medium_generic_f16_threads=8/32x512x512 |
31.01 µs | 29.29 µs | -5.5% |
| ✅ | gather/small_f16_threads=1-internal/4096 |
525.0 ns | 491.1 ns | -6.5% |
| ✅ | add/large_f32_threads=1-internal/4194304 |
47.61 ms | 44.40 ms | -6.7% |
| ✅ | gather/small_f32_threads=1-internal/4096 |
800.3 ns | 744.8 ns | -6.9% |
| ✅ | matmul/large_generic_f32_threads=8/32x1024x1024 |
4.10 ms | 3.79 ms | -7.6% |
| ✅ | gather/small_bf16_threads=1-internal/4096 |
543.0 ns | 500.7 ns | -7.8% |
| ✅ | sampling_latency/top_p_per_token |
1.06 ms | 963.62 µs | -8.7% |
| ✅ | sampling_latency/greedy_per_token |
3.42 µs | 3.10 µs | -9.3% |
| ✅ | add/small_f16_threads=1-internal/1024 |
15.54 µs | 14.06 µs | -9.5% |
| ✅ | tokenization/encode_tokens_per_second |
408.12 µs | 367.52 µs | -9.9% |
| ✅ | add/medium_bf16_threads=1-internal/262144 |
3.24 ms | 2.90 ms | -10.3% |
| ✅ | reduce_mean/large_f32_threads=1-internal/262144 |
1.18 ms | 1.06 ms | -10.4% |
| ✅ | kv_cache/alloc_dealloc_pages |
41.26 µs | 36.90 µs | -10.6% |
| ✅ | sampling_latency/min_p_per_token |
369.21 µs | 329.12 µs | -10.9% |
| ✅ | tokenization/decode_tokens_per_second |
6.71 ms | 5.98 ms | -10.9% |
| ✅ | matmul/small_generic_bf16_threads=1/1x256x256 |
34.29 µs | 30.45 µs | -11.2% |
| ✅ | matmul/large_generic_f32_threads=1/32x1024x1024 |
10.13 ms | 8.93 ms | -11.9% |
| ✅ | logit_processing/seven_processor_chain_per_step |
1.29 ms | 1.14 ms | -12.0% |
| ✅ | grammar_masking/llguidance_compute_mask/32 |
83.12 µs | 71.08 µs | -14.5% |
| 🟢 | add/large_f16_threads=1-internal/4194304 |
51.92 ms | 44.04 ms | -15.2% |
| 🟢 | add/small_f32_threads=1-internal/1024 |
245.5 ns | 204.9 ns | -16.6% |
| 🟢 | add/medium_f16_threads=1-internal/262144 |
3.56 ms | 2.80 ms | -21.4% |
| 🟢 | matmul/medium_generic_bf16_threads=8/32x512x512 |
486.61 µs | 378.30 µs | -22.3% |
| 🟢 | sampling_latency/top_k_per_token |
584.78 µs | 453.41 µs | -22.5% |
| 🟢 | gather/medium_f16_threads=1-internal/32768 |
3.45 µs | 2.66 µs | -22.9% |
| 🟢 | add/medium_f32_threads=1-internal/262144 |
4.01 ms | 2.72 ms | -32.3% |
Visual flags:
Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.99 4.49 7.64 }
What this cannot catch
- Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
- Sub-threshold regressions that compound over multiple PRs
- Performance changes that only manifest under GPU execution
- Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)
|
VERDICT: APPROVE Independent review by Melina (reviewer). Verified on an H200 (CUDA_VISIBLE_DEVICES=5), worktree at origin/squad/67-cuda-op-coverage @ 2ee8a9d. Author Cohaagen locked out; I re-derived every claim from source and ran the suites myself. 1. CausalConvWithState numerical correctness — FAITHFULf32 accumulation: confirmed. In the NVRTC kernel (causal_conv_with_state.rs:66-80) acc is float, every load widens via LOAD (CCWS_ID / __half2float / __bfloat162float), and store narrows round-to-nearest (__float2half_rn / __float2bfloat16_rn). fp16/bf16 arithmetic is done in f32, matching the CPU EP dtype policy. State-update causality: correct, no future leakage. Per (b,c) the kernel walks pos = t+k over the virtual seq = concat(past_state[pad], x[L]); pos<pad reads state, else x[pos-pad] (lines 70-77). The current input x[t] is the LAST window element (k=K-1), so output t depends only on past+current. present_state = seq[length..length+pad] = trailing K-1 frames (lines 82-92). This is byte-identical in structure to the CPU kernel (ep-cpu/src/kernels/causal_conv.rs:203-230, incl. the L=1 fast path), which its header documents as verified to fp32 epsilon against ORT 1.26 (tests/qwen35_ort_parity.rs). So the oracle chain is CUDA -> CPU EP -> ORT, a trusted reference — not a self-comparison. Padding / kernel-window indexing: no off-by-one. Max seq index = (L-1)+(K-1) = L+pad-1 = last element; min = 0. pad = K-1. Shapes fully validated (x rank-3, weight [C,1,K], bias [C], state/present [B,C,K-1]) and contiguity enforced (lines 213-312). Parity suite (causal_conv_with_state_gpu.rs) compares run_cuda vs run_cpu on non-trivial shapes (B1C3K4, B2C2K3), decode L=1 + past_state + present, prefill L=4, with/without bias, with/without state, none+silu, across fp32/fp16/bf16, tol 2e-5/4e-3/4e-2. Real reference parity. 2. Coverage-of-coverage — INTACT, COUNTS MATCHBoth ops added to CUDA_COVERED_OPS (kernels/mod.rs:330-331) and as dedicated conformance profile entries pointing at their real suites (cuda_conformance_gpu.rs). Guards: every_covered_op_has_a_conformance_entry (also asserts profile.len()==CUDA_COVERED_OPS.len(), so no phantom declarations), profile_has_no_duplicate_entries, dedicated_suites_exist_and_name_their_op (greps the suite for the quoted op name), plus the sweep. 3. GatherBlockQuantized bits=4 zero-point edge — REAL but SAFE-TO-DEFER (my call)The bug is real. CUDA unpacks the zero-point with GLOBAL nibble addressing: block_id = in_idx/block_size; zp[block_id/2], nibble block_id%2 (gather_block_quantized.rs:84-89). ORT/CPU uses PER-ROW addressing: zp[scale_row*ceil(bpr/2) + q_in_row/2] (ep-cpu/.../gather_block_quantized.rs:108-121). For EVEN blocks-per-row (bits=4, components=2) the two coincide exactly — the real int4 embedding.onnx layout. For ODD blocks-per-row they diverge. Why it is safe to defer:
Non-blocking follow-up I recommend: (a) fix the kernel to per-row nibble addressing for bits<8, OR add an EXPLICIT claim-gate/assert rejecting odd-blocks-per-row with bits<8 (today it only fails closed incidentally via the shape-size mismatch); and (b) soften the kernel doc header, which currently claims it "mirrors ORT numerics exactly" — true only for even blocks-per-row. 4. cfg-correctness / regressions
Test results (H200, device 5)
VerdictAPPROVE. The CausalConvWithState kernel is a faithful, causal, f32-accumulating port validated against an ORT-verified oracle; coverage-of-coverage integrity and counts are correct; the GBQ bits=4 zero-point edge is a pre-existing, inert-for-real-models, fail-closed latent issue that is safe to defer with the hardening follow-up noted above. Not merging; not using gh pr review. |
…CUDA-hybrid merge logs (#483) Scribe round 4 state consolidation (squad-internal, no code). Merges 4 design-note inbox drops into decisions/archive, distils 2 standing directives, logs #477/#478/#479/#480 merges, appends histories. decisions.md 25.4KB→28.5KB (older 07-29 entries flagged for next-round distillation). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…rnel Implement the recurrent gated delta-rule linear attention op used by the Qwen3.5 / Qwen3-Next hybrid family on the CUDA EP, the second-ranked coverage gap after CausalConvWithState (#480). A single NVRTC kernel (f32/f16/bf16 entry points) exploits the fact that each column of the per-head state matrix S[d_k, d_v] evolves independently, mapping one thread to each (batch, kv_head, d_v-column) and running the whole recurrent scan in f32 (state kept in a per-thread f32 register array) so the arithmetic matches ORT's float CPU kernel regardless of I/O dtype. Covers all four update_rule variants (linear/gated/delta/gated_delta), standard and inverse GQA, key-head sharing (n_k < H_kv), per-head and per-key-dim decay, per-head and shared beta, and step-to-step state carry via past/present_state. The claim gate fail-closes on unsupported dtypes and d_k > 256. Wiring: factory + unsupported_reason claim gate + registration + CUDA_COVERED_OPS entry + dedicated GPU parity suite (linear_attention_gpu.rs) vs the CPU EP oracle plus a dedicated() conformance profile entry (coverage-of-coverage green). Placement probe confirms all 18/18/24 LinearAttention nodes in qwen3.5-0.8b/2b/9b now place on CUDA (0 before). Pairs with CausalConvWithState to land the hybrid decode path. Refs #67, #384. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…en3.5 hybrid (#67) (#484) ## What Implements `com.microsoft::LinearAttention` (Gated DeltaNet / gated delta-rule linear attention) on the **CUDA EP** — the recurrent attention of the Qwen3.5 / Qwen3-Next **hybrid** family, and the second-ranked CUDA coverage gap after `CausalConvWithState` (#480). Together they land the hybrid (Mamba / linear-attention) decode path on CUDA. Refs #67, #384. **Draft — do not merge.** ## Design (see `.squad/decisions/inbox/cohaagen-linear-attention-design.md`) **Oracle:** the CPU EP kernel `onnx-runtime-ep-cpu/src/kernels/linear_attention.rs` (a faithful port of ORT `contrib_ops/cpu/bert/linear_attention.cc`). **Key structural fact:** each column `j` of the per-head state matrix `S[d_k, d_v]` evolves *independently* — retrieval `r[j]=Σ_i S[i,j]·k[i]`, the delta/linear update, and readout `o[j]=scale·Σ_i q[i]·S[i,j]` all touch only column `j`. So the op is embarrassingly parallel across `(b, h_kv, j)`; each is one sequential scan over `t`. **Mapping:** one CUDA thread per `(b, h_kv, j)`; each thread keeps its state column in a **per-thread f32 register array** and runs the whole recurrent scan in **f32**, so the arithmetic matches ORT's `float` CPU kernel regardless of I/O dtype (`f16`/`bf16` widened on read, narrowed on write). No shared memory, no cross-thread reduction, no scratch alloc. ## Coverage - All four `update_rule` variants: `linear`, `gated`, `delta`, `gated_delta`. - Standard GQA (`H_q ≥ H_kv`) **and** inverse GQA (`H_q < H_kv`, the 9b config). - Key-head sharing (`n_k < H_kv`), per-head **and** per-key-dim decay, per-head **and** shared beta. - Step-to-step state carry via `past_state` / `present_state`. - Dtypes `Float32` / `Float16` / `BFloat16`, f32 accumulation. - Claim gate fail-closes on unsupported dtypes and `d_k > 256`. ## Verification (GPU device 0) - **Parity vs CPU EP oracle** (`tests/linear_attention_gpu.rs`, 4 tests): every config's `output` **and** `present_state` compared to the CPU kernel (tol f32 2e-4 / f16 5e-2 / bf16 2e-1), plus a **chained-vs-full state-carry** proof (two 3-step halves chained through `past_state` == one 6-step run). All green. - **Coverage-of-coverage**: `LinearAttention` added to `CUDA_COVERED_OPS` + a `dedicated()` conformance entry; `every_covered_op_has_a_conformance_entry` and `dedicated_suites_exist_and_name_their_op` pass. Full `cuda_conformance_gpu` suite: 4/4. - **Lib**: 274/274. - **Placement proof** (probe over real Foundry models): LinearAttention nodes placed on CUDA went **0 → 18** (qwen3.5-0.8b), **0 → 18** (2b), **0 → 24** (9b). - `cargo fmt --all --check` clean; `cargo clippy` clean for `-p onnx-runtime-ep-cuda` (cuda / no-cuda) and `-p onnx-genai-engine --features cuda,native-backend` (no new warnings from these files). ## Deferred (kept out to keep this PR focused) `com.microsoft` domain `RotaryEmbedding` registration and `Bool`-input `NonZero` — independent of LinearAttention, each with its own semantic checks; tracked as small follow-ups. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#67) (#525) Focused #67 coverage-completion + robustness PR closing three small, independent CUDA EP gaps surfaced during the Qwen3.5 hybrid work (#480/#484). Each is proven against the CPU EP oracle and on real qwen3.5-0.8b nodes. Refs #67. ## 1. `RotaryEmbedding` for `com.microsoft` CUDA previously registered only `ai.onnx::RotaryEmbedding` (opset 23); the `com.microsoft` contrib variant used by qwen3.5 text decoders fell to CPU. The math is identical — only the input order differs (`X, position_ids, cos, sin` vs `X, cos, sin, position_ids?`). - Shared `contrib`-aware kernel + `RotaryEmbeddingContribFactory`; a single `resolve_input_order(contrib, n)` maps `(cos_i, sin_i, pos_i)`; all hardcoded `inputs[1]/[2]/[3]` refs (execute, dtype check, cache validation, CUDA-graph capture signature, launch) route through it (DRY). - Fixes a latent dtype bug: the claim/execute checks used `inputs[..3]`/`take(3)`, wrongly comparing the Int64 `position_ids` against the float dtype for the contrib ordering. - Claim gate + registration wired for `com.microsoft`. - **Parity:** new `tests/rope_contrib_gpu.rs`, fp32/fp16/bf16 × interleaved∈{0,1}, tol-exact vs CPU EP. - **Placement:** 12/12 com.microsoft RotaryEmbedding nodes now claim CUDA in qwen3.5-0.8b `text.onnx`. ## 2. `Bool`-input `NonZero` CUDA handled f32/f16/bf16 only, and the CPU EP oracle *also* rejected Bool — so real models with a Bool `NonZero` (qwen3.5 `embedding.onnx`) failed on both EPs. - CUDA: `DEFINE_NONZERO(unsigned char, bool_)` on the existing NVRTC macro + Bool dispatch arm + `DataType::Bool` in the claim gate. - CPU: `NonZeroKernel` reads a Bool mask via a new `to_dense_bool` helper (makes CPU fallback work AND a valid parity oracle). - **Parity:** `NonZero[bool]` conformance sweep case (ExactBytes vs CPU) + CPU `nonzero_accepts_bool` unit test. - **Placement:** 1/1 Bool NonZero node claims CUDA in qwen3.5-0.8b `embedding.onnx`. ## 3. GatherBlockQuantized odd-blocks-per-row gate + honest doc (from Melina's #480 review) The CUDA bits=4 zero-point unpack uses **global** nibble addressing; CPU/ORT pack zero points **per row**. They agree only for an **even** number of blocks per row (bits=8 always agrees). Odd-bpr-with-zp previously failed only incidentally via a confusing size mismatch, and GBQ had no claim gate (claim-then-fail). - Explicit **loud** execute bail + a static-shape **claim gate** `unsupported_reason(node, shapes)` (conservative on symbolic shapes / no zero points). - Doc softened: "matches ORT for the uint8 path" + a dedicated even-blocks-per-row precondition section. - Per-row addressing itself **unchanged** (even-bpr is the real int4 embedding layout). - **Test:** `claim_gates_gpu::gather_block_quantized_odd_blocks_per_row_with_zero_points_declines` (odd declines with a clear reason; even still claims). Existing bits4/bits8 parity green. ## Verification - `cargo test -p onnx-runtime-ep-cuda --features cuda`: lib **274 passed**; conformance 4 (sweep incl. Bool case + coverage-of-coverage green); rope_contrib 1, claim_gates 4, gather_block_quantized 2, rope_capture 1 — all green. - `cargo test -p onnx-runtime-ep-cpu --lib` NonZero incl. `nonzero_accepts_bool` green. - `cargo fmt --all --check` clean; clippy `-p onnx-runtime-ep-cuda` (cuda / no-cuda), `-p onnx-runtime-ep-cpu`, `-p onnx-genai-engine --features cuda,native-backend` — zero new warnings from touched files. Coverage-of-coverage: all three op names already covered (keys on name, not domain) — no phantom declarations. Draft; do not merge. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Composition proof for the Qwen3.5 hybrid recurrent op coverage: the whole qwen3.5-0.8b split decode graph (embedding.onnx 24 + text.onnx 1265 = 1289 nodes, control-flow bodies recursed) places 100% on the native CUDA EP with ZERO declines now that all coverage is merged (CausalConvWithState #480, GatherBlockQuantized #480, LinearAttention #484, com.microsoft RotaryEmbedding + Bool NonZero #525). - qwen35_0_8b_placement_lock (onnx-runtime-ep-cuda): drives every node through the native CUDA claim gate and asserts CausalConvWithState (18, #480), LinearAttention (18, #484) and GatherBlockQuantized (#480) all claim, and the whole 1289-node graph places with zero declines. Any decline (a regressed covered op) fails the test. - qwen35_0_8b_hybrid_native_cuda_e2e (onnx-genai-engine): end-to-end native-CUDA decoder-vs-ORT token-parity harness. Skips gracefully on the current loader blocker (Engine::from_dir rejects the 3-onnx split; from_pipeline_dir refuses the package during vision smart_resize preprocessing admission) and becomes a live parity lock once that loader/preprocessing gap closes. Full end-to-end native-CUDA decode is blocked only by loader/preprocessing plumbing outside the CUDA-EP op scope; documented in .squad/decisions/inbox/cohaagen-hybrid-e2e.md. Refs #67, #384. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…) (#529) ## Capstone: Qwen3.5-0.8B hybrid → native CUDA composition proof Refs #67, #384. **DRAFT — do not merge.** The per-op CUDA coverage for the Qwen3.5 hybrid recurrent path is done (`CausalConvWithState` #480, `LinearAttention` #484 — both **merged** —, plus `com.microsoft::RotaryEmbedding` + `Bool` `NonZero` + GBQ gate in **#525**, in review). This PR proves those per-op kernels **compose over a real model** and locks it against regression. ### Whole-graph CUDA placement: PROVEN ✅ `qwen35_0_8b_placement_lock` (onnx-runtime-ep-cuda) walks the real split decode graph — `embedding.onnx` (24 nodes) + `text.onnx` (1265 nodes) = **1289 nodes**, recursing control-flow bodies — through the **native CUDA claim gate** and asserts: - `com.microsoft::CausalConvWithState` = 18, all claim (#480) - `com.microsoft::LinearAttention` = 18, all claim (#484) - `com.microsoft::GatherBlockQuantized` claims (#480) - the **only** declines are the in-review **#525** ops (`com.microsoft::RotaryEmbedding` ×12 + `Bool` `NonZero` ×1). Any *other* decline (a regressed covered op) fails the test. On plain `origin/main`: 13 declines, all #525. **With #525 merged (verified locally): 0 declines, 100% placement.** This branch is based on plain main and carries **no** #525 code (to avoid duplicating Melina's in-review PR). ### End-to-end native-CUDA *decode*: blocked on loader plumbing (documented)⚠️ A token-parity decode run could **not** be executed: no public high-level engine entry loads this split hybrid package. - `Engine::from_dir` rejects the 3 sibling `.onnx` files (single-model loader). - `Engine::from_pipeline_dir` refuses the package during **vision** preprocessing admission (`Resize.attrs.smart_resize=true` unrepresentable) — even though a text decode never touches vision. - The decoder's sequence source is `inputs_embeds` (no token-id input); the native inputs_embeds step driver is `pub(crate)`. All three are loader / preprocessing-metadata / pipeline-plumbing gaps **outside the CUDA-EP op scope** (#67). Details + follow-up options in `.squad/decisions/inbox/cohaagen-hybrid-e2e.md`. `qwen35_0_8b_hybrid_native_cuda_e2e` (onnx-genai-engine) is the committed decoder-vs-ORT parity harness; it **skips gracefully** on that loader blocker and becomes a live parity lock the moment the gap closes. ### Verification - `qwen35_0_8b_placement_lock` — **1 passed** (1289 nodes, 13 declines all #525) - `qwen35_0_8b_hybrid_native_cuda_e2e` — **1 passed** (graceful skip, blocker logged) - `cuda_conformance_gpu` coverage-of-coverage — **4 passed** (no ops added) - `cargo fmt --all --check` clean - clippy `onnx-runtime-ep-cuda` (with/without `cuda`) + `onnx-genai-engine` (`cuda,native-backend`) — no new warnings Purely additive: 2 new `#[ignore]` GPU+model-gated tests + their Cargo entries + the decisions note. No existing code touched. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…packages (#67, #384) A split VLM package (vision+embedding+decoder) whose declared image preprocessing is not representable by the runtime (Qwen `smart_resize`) previously aborted the entire pipeline-metadata synthesis, so BOTH ORT and native `Engine::from_pipeline_dir` refused it and text decode could never run. This blocked the Qwen3.5-0.8B hybrid (Mamba/linear-attention) model whose per-op CUDA coverage already landed (#480/#484/#525). Text never touches vision, so admit such a package for text-only decode, driven purely by its declared modality shape (not a model name): - New `GenAiConfigError::UnrepresentablePreprocessing`, returned by the smart_resize branch, kept distinct from `IncompletePipeline` so genuinely incomplete packages still fail hard. - `to_strict_text_only_pipeline_metadata` synthesizes an embedding->decoder AR pipeline with no vision/image-preprocessing/image-dataflow. Rank-3 positions use `linear_increment` (every mrope axis advances with the sequence position, the correct pure-text coordinates); decoder declares `sequence_source: inputs_embeds`; the vision-fed `image_features` embedding input becomes optional with an empty (zero image-token) absent value. - `pipeline_inference_metadata_from_dir` falls back to it on the unrepresentable-preprocessing signal; representable VLMs are unchanged. Also resolve a symbolic leading (batch) axis when zero-initializing loop-carried fixed state (`conv_state`, `recurrent_state` export batch as -1), mirroring the empty-KV convention; non-batch symbolic dims are still refused loudly. This is in the shared ORT decode path (decode/values.rs, resolved_io.rs), not the native step driver. Result: the real qwen3.5-0.8b hybrid now loads and greedy-decodes coherently end-to-end via ORT ("The capital of France is" -> " Paris, and the capital of Germany is Berlin."), locked by an active regression test that skips gracefully when the model dir is absent. Native-CUDA decoder parity is a documented handoff: the native step driver hardcodes rank-2 position_ids and must build rank-3 mrope coordinates (Mary's Inc3c native_decode files) before native CUDA can drive it. Existing VLM synthesis and native/CPU/CUDA pipeline decode parity remain green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…tive runtime (#67, #384) (#535) ## Summary Closes the **loader/plumbing gap** so the Qwen3.5-0.8B hybrid (Mamba/linear-attention) split package actually **loads and decodes end-to-end**, composing the per-op CUDA coverage work (#480 CausalConvWithState, #484 LinearAttention, #525 RoPE-contrib + Bool NonZero) into a working model. ### The blocker The Foundry export is a 3-ONNX split package (`vision.onnx` + `embedding.onnx` + `text.onnx`) whose declared image preprocessing uses Qwen `smart_resize` — which has no lossless runtime encoding. That error aborted the **entire** pipeline-metadata synthesis before admission, so BOTH ORT and native `Engine::from_pipeline_dir` refused the package and text decode could never run. ### The fix (general, modality-driven — not a model-name special-case) Text decode never touches vision, so a split VLM package whose image path is unusable is admitted for **text-only decode**: 1. **`GenAiConfigError::UnrepresentablePreprocessing`** — new distinct variant from the `smart_resize` branch, kept separate from `IncompletePipeline` so genuinely-incomplete packages still fail hard. 2. **`to_strict_text_only_pipeline_metadata`** — synthesizes an embedding→decoder AR pipeline with no vision/image-preprocessing/dataflow. Rank-3 positions use `linear_increment` (every mrope axis advances with the sequence position → correct pure-text `[t,t,t]`); decoder declares `sequence_source: inputs_embeds`; the vision-fed `image_features` embedding input becomes optional with an empty (zero image-token) absent value. 3. **`pipeline_inference_metadata_from_dir`** falls back on the unrepresentable-preprocessing signal; representable VLMs are unchanged. 4. **Symbolic-batch loop-state init** (shared ORT decode path, `decode/values.rs` + `resolved_io.rs`): `conv_state`/`recurrent_state` export the leading batch axis as `-1`; it now resolves to the decode batch (1), mirroring the empty-KV convention. Non-batch symbolic dims still refused loudly. **Not** in Mary's native step driver. ### Result — runs & coherent (ORT reference) ``` prompt : "The capital of France is" output : " Paris, and the capital of Germany is Berlin.\nThe capital of France is" ``` Correct fact ("Paris") ⇒ positions / sequence-source / optional-image / loop-state are all correct. Locked by an **active** regression test `qwen35_0_8b_hybrid_text_decode_e2e.rs` (skips gracefully when the model dir is absent). Reference caveat (honest): the reference is an ORT decode of the same synthesized spec (ORT falls back to CPU for the `com.microsoft` hybrid ops). Per-op CUDA↔reference parity is proven separately (#480/#484/#525); no independent onnxruntime-genai oracle is wired, so coherence is the mitigating oracle for a shared-spec bug. ### Native-CUDA last mile — HANDOFF to Mary (Inc3c) Native decode can't yet drive this model: the native step driver hardcodes **rank-2** `position_ids` (`native_decode/cuda.rs:248`, `cpu.rs:203`) but the hybrid decoder declares **rank-3** mrope positions → `rank mismatch (graph declares rank 3, got 2)`. These are Mary's active Inc3c files, so per the collision rule they were **not** edited. Needed: build rank-3 mrope coordinates in the native step driver honoring the pipeline `positions` program (as `decode/step.rs` already does for ORT). Then flip the `qwen35_0_8b_hybrid_native_cuda_e2e` harness (#529) to native-vs-ORT parity. Details in `.squad/decisions/inbox/cohaagen-hybrid-loader.md`. ## Verification - `qwen35_0_8b_hybrid_text_decode_e2e` — **1 passed** (active, coherent + exact greedy lock). - `onnx-genai-genai-config` — **27 lib + 4 vlm_pipeline passed** (incl. new text-only synthesis test; existing VLM synthesis unchanged). - `onnx-genai-engine --lib` — **284 passed** (incl. 2 new symbolic-batch state cases). - Existing VLM pipeline decode parity — `native_cuda_pipeline_decoder_parity`, `native_pipeline_decoder_parity` **passed** (no regression from shared decode changes). - `every_covered_op_has_a_conformance_entry` (coverage-of-coverage) **passed** (CUDA_COVERED_OPS untouched). - `cargo fmt --all --check` clean; `cargo clippy` clean for `onnx-genai-genai-config`, `onnx-genai-engine` (default) and `onnx-genai-engine --features cuda,native-backend`. Refs #67, #384. Do not merge. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Closes part of #67 (CUDA EP operator coverage parity — additive decode-path batch).
Part 1 — Assessment (data-driven)
I probed the real target decode models through the production loader (shape
inference on), then ran the actual CUDA EP
supports_opper node (recursingIf/Loop/Scansubgraph bodies) to find both op-type gaps and dtype/shapeclaim-gate silent fallbacks.
Current covered set:
CUDA_COVERED_OPShad 161 entries; the registryalso had
GatherBlockQuantizedregistered but undeclared (a realcoverage-of-coverage hole).
Result — the classic transformer decode path is already 100% on CUDA:
IfScanScanZero dtype/claim-gate declines on these paths. The only "uncovered" types are
control-flow ops the session executor handles recursively (not EP ops).
Ranked genuinely-missing ops (block real LLM decode on CUDA) — the Qwen3.5
hybrid (Mamba + linear-attention) family:
embedding.onnxExcluded per scope: the 27 vision
Attentionfloat-mask nodes and MoE #82.Part 2 — Implemented (top batch)
CausalConvWithState— new NVRTC kernel (fp32/fp16/bf16, f32accumulation matching CPU EP/ORT), factory + claim gate + registration +
CUDA_COVERED_OPSentry + dedicated GPU parity suite(
causal_conv_with_state_gpu.rs, vs the CPU EP oracle: decode/prefill,with/without bias & state, none/silu). Verified the 18 previously-uncovered
nodes in qwen3.5-0.8b
text.onnxnow place on CUDA.GatherBlockQuantized— declared inCUDA_COVERED_OPS+ dedicated GPUparity suite (
gather_block_quantized_gpu.rs; bits 8/4, fp32+fp16 scales,int32/int64 indices), closing the registered-but-undeclared hole.
Verification
cargo test -p onnx-runtime-ep-cuda --features cuda: coverage-of-coverageguards (4/4), new suites (4 tests), lib (274), claim_gates (3), construction
(18) — all pass on GPU 0.
cargo fmt --all --checkclean; clippy clean for the new files under bothcuda and non-cuda feature sets.
Honest follow-ups
LinearAttention(rank 2) pairs with CausalConvWithState to fully landQwen3.5 hybrid decode on CUDA — deferred (larger CPU kernel, own parity PR).
com.microsoft::RotaryEmbedding(registered only forai.onnx; 12 nodes decline in qwen3.5 text) and a
Bool-inputNonZero(embedding) — both small follow-ups.
bits=4zero-point packing is globalvs the CPU/ORT per-row packing; they coincide only when blocks-per-row is
even (the real int4 embedding layout). Parity suite uses that real layout.
Draft — do not merge.