Skip to content

feat(cuda-ep): CausalConvWithState + declare GatherBlockQuantized (#67) - #480

Merged
justinchuby merged 1 commit into
mainfrom
squad/67-cuda-op-coverage
Jul 30, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/67-cuda-op-coverage

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Closes part of #67 (CUDA EP operator coverage parity — additive decode-path batch).

Part 1 — Assessment (data-driven)

I probed the real target decode models through the production loader (shape
inference on), then ran the actual CUDA EP supports_op per node (recursing
If/Loop/Scan subgraph bodies) to find both op-type gaps and dtype/shape
claim-gate silent fallbacks.

Current covered set: CUDA_COVERED_OPS had 161 entries; the registry
also had GatherBlockQuantized registered but undeclared (a real
coverage-of-coverage hole).

Result — the classic transformer decode path is already 100% on CUDA:

Model covered-type nodes placed on CUDA uncovered (executor-handled)
qwen2.5-0.5b/1.5b/7b-instruct-cuda all —
Phi-4-mini-instruct-cuda 624/624 1× If
Qwen3.6-27B 7065/7065 48× Scan
Qwen3.5-35B-A3B int4 decoder 127815/127815 30× Scan

Zero dtype/claim-gate declines on these paths. The only "uncovered" types are
control-flow ops the session executor handles recursively (not EP ops).

Ranked genuinely-missing ops (block real LLM decode on CUDA) — the Qwen3.5
hybrid (Mamba + linear-attention) family:

Rank Op (com.microsoft) Real models blocked Nodes/model
1 CausalConvWithState qwen3.5 9b / 2b-text / 0.8b 24 / 18 / 18
2 LinearAttention qwen3.5 9b / 2b-text / 0.8b 24 / 18 / 18
3 GatherBlockQuantized qwen3-0.6b, qwen3.5-2b, every qwen3.5 embedding.onnx 1 (whole embedding component); already registered on CUDA

Excluded per scope: the 27 vision Attention float-mask nodes and MoE #82.

Part 2 — Implemented (top batch)

  1. CausalConvWithState — new NVRTC kernel (fp32/fp16/bf16, f32
    accumulation matching CPU EP/ORT), factory + claim gate + registration +
    CUDA_COVERED_OPS entry + dedicated GPU parity suite
    (causal_conv_with_state_gpu.rs, vs the CPU EP oracle: decode/prefill,
    with/without bias & state, none/silu). Verified the 18 previously-uncovered
    nodes in qwen3.5-0.8b text.onnx now place on CUDA.
  2. GatherBlockQuantized — declared in CUDA_COVERED_OPS + dedicated GPU
    parity suite (gather_block_quantized_gpu.rs; bits 8/4, fp32+fp16 scales,
    int32/int64 indices), closing the registered-but-undeclared hole.

Verification

  • cargo test -p onnx-runtime-ep-cuda --features cuda: coverage-of-coverage
    guards (4/4), new suites (4 tests), lib (274), claim_gates (3), construction
    (18) — all pass on GPU 0.
  • cargo fmt --all --check clean; clippy clean for the new files under both
    cuda and non-cuda feature sets.

Honest follow-ups

  • LinearAttention (rank 2) pairs with CausalConvWithState to fully land
    Qwen3.5 hybrid decode on CUDA — deferred (larger CPU kernel, own parity PR).
  • Probe also surfaced com.microsoft::RotaryEmbedding (registered only for
    ai.onnx; 12 nodes decline in qwen3.5 text) and a Bool-input NonZero
    (embedding) — both small follow-ups.
  • Latent note: the existing GBQ kernel's bits=4 zero-point packing is global
    vs the CPU/ORT per-row packing; they coincide only when blocks-per-row is
    even (the real int4 embedding layout). Parity suite uses that real layout.

Draft — do not merge.

…zed (#67)

Data-driven CUDA op-coverage batch for issue #67. A placement audit over the
real target decode models (Qwen2.5 0.5b/1.5b/7b, Phi-4-mini, Qwen3.6-27B,
Qwen3.5-35B-A3B) showed the classic transformer decode path already places
100% of EP-placeable nodes on CUDA — the only uncovered types are the
executor-handled control-flow ops (If/Loop/Scan). The genuine remaining gaps
are the Qwen3.5 hybrid (Mamba + linear-attention) family.

This batch:

* Adds a new NVRTC CUDA kernel for `com.microsoft::CausalConvWithState`
  (depthwise causal 1-D short-conv with rolling state), covering fp32/fp16/bf16
  with f32 accumulation to match the CPU EP / ORT numerics, an optional bias
  and past_state, optional present_state output, and none/silu(swish)
  activation. Registers it, adds a claim gate (ndim=1, activation, float
  dtypes), and a dedicated GPU parity suite vs the CPU EP oracle. Unblocks the
  Qwen3.5 hybrid short-conv on CUDA (18-24 nodes/model across qwen3.5
  0.8b/2b/9b text decoders).

* Declares the already-registered `com.microsoft::GatherBlockQuantized` in
  CUDA_COVERED_OPS and adds a dedicated GPU parity suite (bits 8/4, fp32+fp16
  scales/output, int32/int64 indices), closing a registered-but-undeclared
  coverage-of-coverage hole. It is the whole compute of the int4 embedding
  component of the Qwen3.5 split models.

Coverage-of-coverage guards pass (163 advertised op names, 168 registry pairs).
Honest follow-ups: com.microsoft::LinearAttention (pairs with CausalConv to
fully land Qwen3.5 hybrid decode), registering RotaryEmbedding for the
com.microsoft domain, and a Bool-input NonZero path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Jul 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.57%. Comparing base (6e4257e) to head (2ee8a9d).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #480      +/-   ##
==========================================
+ Coverage   80.56%   80.57%   +0.01%     
==========================================
  Files         314      315       +1     
  Lines      122637   122793     +156     
  Branches   122637   122793     +156     
==========================================
+ Hits        98804    98944     +140     
- Misses      19814    19826      +12     
- Partials     4019     4023       +4     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.78% <ø> (+0.10%) ⬆️
mlas 77.91% <ø> (ø)
offline 80.50% <ø> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 6 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

⚠️ Benchmark Change Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
⚠️ gather/large_f32_threads=1-internal/131072 34.35 µs 42.42 µs +23.5%
⚠️ matmul/small_generic_f32_threads=8/1x256x256 39.29 µs 48.22 µs +22.7%
⚠️ matmul/small_generic_f32_threads=1/1x256x256 36.78 µs 42.63 µs +15.9%
✅ matmul/small_generic_f16_threads=8/1x256x256 32.59 µs 37.03 µs +13.6%
✅ gather/large_bf16_threads=1-internal/131072 12.16 µs 13.56 µs +11.5%
✅ gather/large_f16_threads=1-internal/131072 13.71 µs 14.68 µs +7.1%
✅ gather/medium_f32_threads=1-internal/32768 4.44 µs 4.63 µs +4.5%
✅ matmul/small_generic_f16_threads=1/1x256x256 36.10 µs 37.16 µs +3.0%
✅ gather/medium_bf16_threads=1-internal/32768 2.78 µs 2.82 µs +1.4%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 81.07 µs 80.91 µs -0.2%
✅ matmul/medium_generic_f32_threads=8/32x512x512 951.23 µs 948.60 µs -0.3%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.33 ms 2.30 ms -1.2%
✅ matmul/small_generic_bf16_threads=8/1x256x256 32.48 µs 31.96 µs -1.6%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.95 ms 1.90 ms -2.8%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.74 µs 16.22 µs -3.1%
✅ reduce_mean/medium_f32_threads=1-internal/65536 278.84 µs 267.41 µs -4.1%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 523.22 µs 500.79 µs -4.3%
✅ matmul/medium_generic_f16_threads=1/32x512x512 30.84 µs 29.50 µs -4.3%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 86.34 µs 82.39 µs -4.6%
✅ add/small_bf16_threads=1-internal/1024 14.82 µs 14.12 µs -4.7%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.39 ms 1.32 ms -5.1%
✅ add/large_bf16_threads=1-internal/4194304 47.73 ms 45.28 ms -5.1%
✅ matmul/medium_generic_f16_threads=8/32x512x512 31.01 µs 29.29 µs -5.5%
✅ gather/small_f16_threads=1-internal/4096 525.0 ns 491.1 ns -6.5%
✅ add/large_f32_threads=1-internal/4194304 47.61 ms 44.40 ms -6.7%
✅ gather/small_f32_threads=1-internal/4096 800.3 ns 744.8 ns -6.9%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 4.10 ms 3.79 ms -7.6%
✅ gather/small_bf16_threads=1-internal/4096 543.0 ns 500.7 ns -7.8%
✅ sampling_latency/top_p_per_token 1.06 ms 963.62 µs -8.7%
✅ sampling_latency/greedy_per_token 3.42 µs 3.10 µs -9.3%
✅ add/small_f16_threads=1-internal/1024 15.54 µs 14.06 µs -9.5%
✅ tokenization/encode_tokens_per_second 408.12 µs 367.52 µs -9.9%
✅ add/medium_bf16_threads=1-internal/262144 3.24 ms 2.90 ms -10.3%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.18 ms 1.06 ms -10.4%
✅ kv_cache/alloc_dealloc_pages 41.26 µs 36.90 µs -10.6%
✅ sampling_latency/min_p_per_token 369.21 µs 329.12 µs -10.9%
✅ tokenization/decode_tokens_per_second 6.71 ms 5.98 ms -10.9%
✅ matmul/small_generic_bf16_threads=1/1x256x256 34.29 µs 30.45 µs -11.2%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.13 ms 8.93 ms -11.9%
✅ logit_processing/seven_processor_chain_per_step 1.29 ms 1.14 ms -12.0%
✅ grammar_masking/llguidance_compute_mask/32 83.12 µs 71.08 µs -14.5%
🟢 add/large_f16_threads=1-internal/4194304 51.92 ms 44.04 ms -15.2%
🟢 add/small_f32_threads=1-internal/1024 245.5 ns 204.9 ns -16.6%
🟢 add/medium_f16_threads=1-internal/262144 3.56 ms 2.80 ms -21.4%
🟢 matmul/medium_generic_bf16_threads=8/32x512x512 486.61 µs 378.30 µs -22.3%
🟢 sampling_latency/top_k_per_token 584.78 µs 453.41 µs -22.5%
🟢 gather/medium_f16_threads=1-internal/32768 3.45 µs 2.66 µs -22.9%
🟢 add/medium_f32_threads=1-internal/262144 4.01 ms 2.72 ms -32.3%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.99 4.49 7.64 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review by Melina (reviewer). Verified on an H200 (CUDA_VISIBLE_DEVICES=5), worktree at origin/squad/67-cuda-op-coverage @ 2ee8a9d. Author Cohaagen locked out; I re-derived every claim from source and ran the suites myself.

1. CausalConvWithState numerical correctness — FAITHFUL

f32 accumulation: confirmed. In the NVRTC kernel (causal_conv_with_state.rs:66-80) acc is float, every load widens via LOAD (CCWS_ID / __half2float / __bfloat162float), and store narrows round-to-nearest (__float2half_rn / __float2bfloat16_rn). fp16/bf16 arithmetic is done in f32, matching the CPU EP dtype policy.

State-update causality: correct, no future leakage. Per (b,c) the kernel walks pos = t+k over the virtual seq = concat(past_state[pad], x[L]); pos<pad reads state, else x[pos-pad] (lines 70-77). The current input x[t] is the LAST window element (k=K-1), so output t depends only on past+current. present_state = seq[length..length+pad] = trailing K-1 frames (lines 82-92). This is byte-identical in structure to the CPU kernel (ep-cpu/src/kernels/causal_conv.rs:203-230, incl. the L=1 fast path), which its header documents as verified to fp32 epsilon against ORT 1.26 (tests/qwen35_ort_parity.rs). So the oracle chain is CUDA -> CPU EP -> ORT, a trusted reference — not a self-comparison.

Padding / kernel-window indexing: no off-by-one. Max seq index = (L-1)+(K-1) = L+pad-1 = last element; min = 0. pad = K-1. Shapes fully validated (x rank-3, weight [C,1,K], bias [C], state/present [B,C,K-1]) and contiguity enforced (lines 213-312).

Parity suite (causal_conv_with_state_gpu.rs) compares run_cuda vs run_cpu on non-trivial shapes (B1C3K4, B2C2K3), decode L=1 + past_state + present, prefill L=4, with/without bias, with/without state, none+silu, across fp32/fp16/bf16, tol 2e-5/4e-3/4e-2. Real reference parity.
RESULT: 2 passed / 0 failed.

2. Coverage-of-coverage — INTACT, COUNTS MATCH

Both ops added to CUDA_COVERED_OPS (kernels/mod.rs:330-331) and as dedicated conformance profile entries pointing at their real suites (cuda_conformance_gpu.rs). Guards: every_covered_op_has_a_conformance_entry (also asserts profile.len()==CUDA_COVERED_OPS.len(), so no phantom declarations), profile_has_no_duplicate_entries, dedicated_suites_exist_and_name_their_op (greps the suite for the quoted op name), plus the sweep.
RESULT: cuda_conformance_gpu 4 passed / 0 failed.
Counts consistent: 167->168 registry pairs (+1: CausalConvWithState newly registered), 161->163 advertised (+2: CausalConvWithState new; GatherBlockQuantized was registered-but-undeclared). docs/CUDA_COVERAGE.md updated to 168/163. No discrepancy.

3. GatherBlockQuantized bits=4 zero-point edge — REAL but SAFE-TO-DEFER (my call)

The bug is real. CUDA unpacks the zero-point with GLOBAL nibble addressing: block_id = in_idx/block_size; zp[block_id/2], nibble block_id%2 (gather_block_quantized.rs:84-89). ORT/CPU uses PER-ROW addressing: zp[scale_row*ceil(bpr/2) + q_in_row/2] (ep-cpu/.../gather_block_quantized.rs:108-121). For EVEN blocks-per-row (bits=4, components=2) the two coincide exactly — the real int4 embedding.onnx layout. For ODD blocks-per-row they diverge.

Why it is safe to defer:

  • It is INERT for the real target layout (even blocks-per-row), which is the whole point of the declaration; the parity suite covers bits=4 WITH a packed zero_point (gather_block_quantized_gpu.rs:97-130) and passes.
  • It is FAIL-CLOSED, never silently wrong, for the odd multi-row case: the CUDA shape check expects global packing size block_count.div_ceil(components) (line 308), which for odd bpr is strictly smaller than ORT's per-row size rows*ceil(bpr/2), so a real per-row-packed zero_points tensor is REJECTED at execute (KernelFailed) rather than producing wrong embeddings. (vocab=1 single-row odd bpr coincides and is correct.)
  • The kernel is PRE-EXISTING and UNCHANGED by this PR (not among the 7 changed files); this PR only declares it + adds coverage. The limitation is documented in the test (lines 100-102) and Cohaagen's decision note.

Non-blocking follow-up I recommend: (a) fix the kernel to per-row nibble addressing for bits<8, OR add an EXPLICIT claim-gate/assert rejecting odd-blocks-per-row with bits<8 (today it only fails closed incidentally via the shape-size mismatch); and (b) soften the kernel doc header, which currently claims it "mirrors ORT numerics exactly" — true only for even blocks-per-row.

4. cfg-correctness / regressions

  • cargo fmt --all --check: clean.
  • Builds: EP crate compiles under default (no-cuda) and cuda; the new module is not cuda-gated and does not break the non-cuda build. (native-backend is an engine-level feature, not on this EP crate; this PR touches only the EP crate + docs.)
  • Clippy (cuda, --all-targets, -D warnings): the PR's OWN changed files are clean — zero findings in causal_conv_with_state.rs, kernels/mod.rs, provider.rs, or the two new suites. The crate-wide clippy failures that appear are all in files this PR does not touch (normalization.rs, standard_attention.rs, matmul_nbits.rs, conv_gpu.rs, pooling_gpu.rs, etc.) and are pre-existing lints surfaced by a local toolchain (rustc 1.97.0, 2026-07-07) that is ahead of CI stable — not a regression introduced here.

Test results (H200, device 5)

  • lib: 274 passed / 0 failed
  • causal_conv_with_state_gpu: 2 passed
  • gather_block_quantized_gpu: 2 passed
  • cuda_conformance_gpu (coverage-of-coverage): 4 passed
  • claim_gates_gpu: 3 passed
  • construction_gpu: 18 passed

Verdict

APPROVE. The CausalConvWithState kernel is a faithful, causal, f32-accumulating port validated against an ORT-verified oracle; coverage-of-coverage integrity and counts are correct; the GBQ bits=4 zero-point edge is a pre-existing, inert-for-real-models, fail-closed latent issue that is safe to defer with the hardening follow-up noted above. Not merging; not using gh pr review.

@justinchuby
justinchuby marked this pull request as ready for review July 30, 2026 17:24
@justinchuby
justinchuby merged commit e2b7565 into main Jul 30, 2026
14 checks passed
@justinchuby
justinchuby deleted the squad/67-cuda-op-coverage branch July 30, 2026 17:24
justinchuby added a commit that referenced this pull request Jul 30, 2026
…CUDA-hybrid merge logs (#483)

Scribe round 4 state consolidation (squad-internal, no code). Merges 4
design-note inbox drops into decisions/archive, distils 2 standing
directives, logs #477/#478/#479/#480 merges, appends histories.
decisions.md 25.4KB→28.5KB (older 07-29 entries flagged for next-round
distillation).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
…rnel

Implement the recurrent gated delta-rule linear attention op used by the
Qwen3.5 / Qwen3-Next hybrid family on the CUDA EP, the second-ranked coverage
gap after CausalConvWithState (#480). A single NVRTC kernel (f32/f16/bf16
entry points) exploits the fact that each column of the per-head state matrix
S[d_k, d_v] evolves independently, mapping one thread to each
(batch, kv_head, d_v-column) and running the whole recurrent scan in f32 (state
kept in a per-thread f32 register array) so the arithmetic matches ORT's float
CPU kernel regardless of I/O dtype.

Covers all four update_rule variants (linear/gated/delta/gated_delta), standard
and inverse GQA, key-head sharing (n_k < H_kv), per-head and per-key-dim decay,
per-head and shared beta, and step-to-step state carry via past/present_state.
The claim gate fail-closes on unsupported dtypes and d_k > 256.

Wiring: factory + unsupported_reason claim gate + registration + CUDA_COVERED_OPS
entry + dedicated GPU parity suite (linear_attention_gpu.rs) vs the CPU EP oracle
plus a dedicated() conformance profile entry (coverage-of-coverage green).

Placement probe confirms all 18/18/24 LinearAttention nodes in
qwen3.5-0.8b/2b/9b now place on CUDA (0 before). Pairs with CausalConvWithState
to land the hybrid decode path.

Refs #67, #384.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
…en3.5 hybrid (#67) (#484)

## What

Implements `com.microsoft::LinearAttention` (Gated DeltaNet / gated
delta-rule linear attention) on the **CUDA EP** — the recurrent
attention of the Qwen3.5 / Qwen3-Next **hybrid** family, and the
second-ranked CUDA coverage gap after `CausalConvWithState` (#480).
Together they land the hybrid (Mamba / linear-attention) decode path on
CUDA.

Refs #67, #384. **Draft — do not merge.**

## Design (see
`.squad/decisions/inbox/cohaagen-linear-attention-design.md`)

**Oracle:** the CPU EP kernel
`onnx-runtime-ep-cpu/src/kernels/linear_attention.rs` (a faithful port
of ORT `contrib_ops/cpu/bert/linear_attention.cc`).

**Key structural fact:** each column `j` of the per-head state matrix
`S[d_k, d_v]` evolves *independently* — retrieval `r[j]=Σ_i
S[i,j]·k[i]`, the delta/linear update, and readout `o[j]=scale·Σ_i
q[i]·S[i,j]` all touch only column `j`. So the op is embarrassingly
parallel across `(b, h_kv, j)`; each is one sequential scan over `t`.

**Mapping:** one CUDA thread per `(b, h_kv, j)`; each thread keeps its
state column in a **per-thread f32 register array** and runs the whole
recurrent scan in **f32**, so the arithmetic matches ORT's `float` CPU
kernel regardless of I/O dtype (`f16`/`bf16` widened on read, narrowed
on write). No shared memory, no cross-thread reduction, no scratch
alloc.

## Coverage

- All four `update_rule` variants: `linear`, `gated`, `delta`,
`gated_delta`.
- Standard GQA (`H_q ≥ H_kv`) **and** inverse GQA (`H_q < H_kv`, the 9b
config).
- Key-head sharing (`n_k < H_kv`), per-head **and** per-key-dim decay,
per-head **and** shared beta.
- Step-to-step state carry via `past_state` / `present_state`.
- Dtypes `Float32` / `Float16` / `BFloat16`, f32 accumulation.
- Claim gate fail-closes on unsupported dtypes and `d_k > 256`.

## Verification (GPU device 0)

- **Parity vs CPU EP oracle** (`tests/linear_attention_gpu.rs`, 4
tests): every config's `output` **and** `present_state` compared to the
CPU kernel (tol f32 2e-4 / f16 5e-2 / bf16 2e-1), plus a
**chained-vs-full state-carry** proof (two 3-step halves chained through
`past_state` == one 6-step run). All green.
- **Coverage-of-coverage**: `LinearAttention` added to
`CUDA_COVERED_OPS` + a `dedicated()` conformance entry;
`every_covered_op_has_a_conformance_entry` and
`dedicated_suites_exist_and_name_their_op` pass. Full
`cuda_conformance_gpu` suite: 4/4.
- **Lib**: 274/274.
- **Placement proof** (probe over real Foundry models): LinearAttention
nodes placed on CUDA went **0 → 18** (qwen3.5-0.8b), **0 → 18** (2b),
**0 → 24** (9b).
- `cargo fmt --all --check` clean; `cargo clippy` clean for `-p
onnx-runtime-ep-cuda` (cuda / no-cuda) and `-p onnx-genai-engine
--features cuda,native-backend` (no new warnings from these files).

## Deferred (kept out to keep this PR focused)

`com.microsoft` domain `RotaryEmbedding` registration and `Bool`-input
`NonZero` — independent of LinearAttention, each with its own semantic
checks; tracked as small follow-ups.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
#67) (#525)

Focused #67 coverage-completion + robustness PR closing three small,
independent CUDA EP gaps surfaced during the Qwen3.5 hybrid work
(#480/#484). Each is proven against the CPU EP oracle and on real
qwen3.5-0.8b nodes. Refs #67.

## 1. `RotaryEmbedding` for `com.microsoft`
CUDA previously registered only `ai.onnx::RotaryEmbedding` (opset 23);
the `com.microsoft` contrib variant used by qwen3.5 text decoders fell
to CPU. The math is identical — only the input order differs (`X,
position_ids, cos, sin` vs `X, cos, sin, position_ids?`).
- Shared `contrib`-aware kernel + `RotaryEmbeddingContribFactory`; a
single `resolve_input_order(contrib, n)` maps `(cos_i, sin_i, pos_i)`;
all hardcoded `inputs[1]/[2]/[3]` refs (execute, dtype check, cache
validation, CUDA-graph capture signature, launch) route through it
(DRY).
- Fixes a latent dtype bug: the claim/execute checks used
`inputs[..3]`/`take(3)`, wrongly comparing the Int64 `position_ids`
against the float dtype for the contrib ordering.
- Claim gate + registration wired for `com.microsoft`.
- **Parity:** new `tests/rope_contrib_gpu.rs`, fp32/fp16/bf16 ×
interleaved∈{0,1}, tol-exact vs CPU EP.
- **Placement:** 12/12 com.microsoft RotaryEmbedding nodes now claim
CUDA in qwen3.5-0.8b `text.onnx`.

## 2. `Bool`-input `NonZero`
CUDA handled f32/f16/bf16 only, and the CPU EP oracle *also* rejected
Bool — so real models with a Bool `NonZero` (qwen3.5 `embedding.onnx`)
failed on both EPs.
- CUDA: `DEFINE_NONZERO(unsigned char, bool_)` on the existing NVRTC
macro + Bool dispatch arm + `DataType::Bool` in the claim gate.
- CPU: `NonZeroKernel` reads a Bool mask via a new `to_dense_bool`
helper (makes CPU fallback work AND a valid parity oracle).
- **Parity:** `NonZero[bool]` conformance sweep case (ExactBytes vs CPU)
+ CPU `nonzero_accepts_bool` unit test.
- **Placement:** 1/1 Bool NonZero node claims CUDA in qwen3.5-0.8b
`embedding.onnx`.

## 3. GatherBlockQuantized odd-blocks-per-row gate + honest doc
(from Melina's #480 review) The CUDA bits=4 zero-point unpack uses
**global** nibble addressing; CPU/ORT pack zero points **per row**. They
agree only for an **even** number of blocks per row (bits=8 always
agrees). Odd-bpr-with-zp previously failed only incidentally via a
confusing size mismatch, and GBQ had no claim gate (claim-then-fail).
- Explicit **loud** execute bail + a static-shape **claim gate**
`unsupported_reason(node, shapes)` (conservative on symbolic shapes / no
zero points).
- Doc softened: "matches ORT for the uint8 path" + a dedicated
even-blocks-per-row precondition section.
- Per-row addressing itself **unchanged** (even-bpr is the real int4
embedding layout).
- **Test:**
`claim_gates_gpu::gather_block_quantized_odd_blocks_per_row_with_zero_points_declines`
(odd declines with a clear reason; even still claims). Existing
bits4/bits8 parity green.

## Verification
- `cargo test -p onnx-runtime-ep-cuda --features cuda`: lib **274
passed**; conformance 4 (sweep incl. Bool case + coverage-of-coverage
green); rope_contrib 1, claim_gates 4, gather_block_quantized 2,
rope_capture 1 — all green.
- `cargo test -p onnx-runtime-ep-cpu --lib` NonZero incl.
`nonzero_accepts_bool` green.
- `cargo fmt --all --check` clean; clippy `-p onnx-runtime-ep-cuda`
(cuda / no-cuda), `-p onnx-runtime-ep-cpu`, `-p onnx-genai-engine
--features cuda,native-backend` — zero new warnings from touched files.

Coverage-of-coverage: all three op names already covered (keys on name,
not domain) — no phantom declarations. Draft; do not merge.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
Composition proof for the Qwen3.5 hybrid recurrent op coverage: the whole
qwen3.5-0.8b split decode graph (embedding.onnx 24 + text.onnx 1265 = 1289
nodes, control-flow bodies recursed) places 100% on the native CUDA EP with
ZERO declines now that all coverage is merged (CausalConvWithState #480,
GatherBlockQuantized #480, LinearAttention #484, com.microsoft RotaryEmbedding
+ Bool NonZero #525).

- qwen35_0_8b_placement_lock (onnx-runtime-ep-cuda): drives every node through
  the native CUDA claim gate and asserts CausalConvWithState (18, #480),
  LinearAttention (18, #484) and GatherBlockQuantized (#480) all claim, and the
  whole 1289-node graph places with zero declines. Any decline (a regressed
  covered op) fails the test.
- qwen35_0_8b_hybrid_native_cuda_e2e (onnx-genai-engine): end-to-end native-CUDA
  decoder-vs-ORT token-parity harness. Skips gracefully on the current loader
  blocker (Engine::from_dir rejects the 3-onnx split; from_pipeline_dir refuses
  the package during vision smart_resize preprocessing admission) and becomes a
  live parity lock once that loader/preprocessing gap closes.

Full end-to-end native-CUDA decode is blocked only by loader/preprocessing
plumbing outside the CUDA-EP op scope; documented in
.squad/decisions/inbox/cohaagen-hybrid-e2e.md. Refs #67, #384.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 30, 2026
…) (#529)

## Capstone: Qwen3.5-0.8B hybrid → native CUDA composition proof

Refs #67, #384. **DRAFT — do not merge.**

The per-op CUDA coverage for the Qwen3.5 hybrid recurrent path is done
(`CausalConvWithState` #480, `LinearAttention` #484 — both **merged** —,
plus
`com.microsoft::RotaryEmbedding` + `Bool` `NonZero` + GBQ gate in
**#525**, in
review). This PR proves those per-op kernels **compose over a real
model** and
locks it against regression.

### Whole-graph CUDA placement: PROVEN ✅
`qwen35_0_8b_placement_lock` (onnx-runtime-ep-cuda) walks the real split
decode
graph — `embedding.onnx` (24 nodes) + `text.onnx` (1265 nodes) = **1289
nodes**,
recursing control-flow bodies — through the **native CUDA claim gate**
and asserts:
- `com.microsoft::CausalConvWithState` = 18, all claim (#480)
- `com.microsoft::LinearAttention` = 18, all claim (#484)
- `com.microsoft::GatherBlockQuantized` claims (#480)
- the **only** declines are the in-review **#525** ops
(`com.microsoft::RotaryEmbedding` ×12 + `Bool` `NonZero` ×1). Any
*other*
  decline (a regressed covered op) fails the test.

On plain `origin/main`: 13 declines, all #525. **With #525 merged
(verified
locally): 0 declines, 100% placement.** This branch is based on plain
main and
carries **no** #525 code (to avoid duplicating Melina's in-review PR).

### End-to-end native-CUDA *decode*: blocked on loader plumbing
(documented) ⚠️
A token-parity decode run could **not** be executed: no public
high-level engine
entry loads this split hybrid package.
- `Engine::from_dir` rejects the 3 sibling `.onnx` files (single-model
loader).
- `Engine::from_pipeline_dir` refuses the package during **vision**
preprocessing
admission (`Resize.attrs.smart_resize=true` unrepresentable) — even
though a
  text decode never touches vision.
- The decoder's sequence source is `inputs_embeds` (no token-id input);
the
  native inputs_embeds step driver is `pub(crate)`.

All three are loader / preprocessing-metadata / pipeline-plumbing gaps
**outside
the CUDA-EP op scope** (#67). Details + follow-up options in
`.squad/decisions/inbox/cohaagen-hybrid-e2e.md`.
`qwen35_0_8b_hybrid_native_cuda_e2e` (onnx-genai-engine) is the
committed
decoder-vs-ORT parity harness; it **skips gracefully** on that loader
blocker and
becomes a live parity lock the moment the gap closes.

### Verification
- `qwen35_0_8b_placement_lock` — **1 passed** (1289 nodes, 13 declines
all #525)
- `qwen35_0_8b_hybrid_native_cuda_e2e` — **1 passed** (graceful skip,
blocker logged)
- `cuda_conformance_gpu` coverage-of-coverage — **4 passed** (no ops
added)
- `cargo fmt --all --check` clean
- clippy `onnx-runtime-ep-cuda` (with/without `cuda`) +
`onnx-genai-engine`
  (`cuda,native-backend`) — no new warnings

Purely additive: 2 new `#[ignore]` GPU+model-gated tests + their Cargo
entries +
the decisions note. No existing code touched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 31, 2026
…packages (#67, #384)

A split VLM package (vision+embedding+decoder) whose declared image
preprocessing is not representable by the runtime (Qwen `smart_resize`)
previously aborted the entire pipeline-metadata synthesis, so BOTH ORT and
native `Engine::from_pipeline_dir` refused it and text decode could never run.
This blocked the Qwen3.5-0.8B hybrid (Mamba/linear-attention) model whose per-op
CUDA coverage already landed (#480/#484/#525).

Text never touches vision, so admit such a package for text-only decode, driven
purely by its declared modality shape (not a model name):

- New `GenAiConfigError::UnrepresentablePreprocessing`, returned by the
  smart_resize branch, kept distinct from `IncompletePipeline` so genuinely
  incomplete packages still fail hard.
- `to_strict_text_only_pipeline_metadata` synthesizes an embedding->decoder AR
  pipeline with no vision/image-preprocessing/image-dataflow. Rank-3 positions
  use `linear_increment` (every mrope axis advances with the sequence position,
  the correct pure-text coordinates); decoder declares `sequence_source:
  inputs_embeds`; the vision-fed `image_features` embedding input becomes
  optional with an empty (zero image-token) absent value.
- `pipeline_inference_metadata_from_dir` falls back to it on the
  unrepresentable-preprocessing signal; representable VLMs are unchanged.

Also resolve a symbolic leading (batch) axis when zero-initializing loop-carried
fixed state (`conv_state`, `recurrent_state` export batch as -1), mirroring the
empty-KV convention; non-batch symbolic dims are still refused loudly. This is in
the shared ORT decode path (decode/values.rs, resolved_io.rs), not the native
step driver.

Result: the real qwen3.5-0.8b hybrid now loads and greedy-decodes coherently
end-to-end via ORT ("The capital of France is" -> " Paris, and the capital of
Germany is Berlin."), locked by an active regression test that skips gracefully
when the model dir is absent. Native-CUDA decoder parity is a documented handoff:
the native step driver hardcodes rank-2 position_ids and must build rank-3 mrope
coordinates (Mary's Inc3c native_decode files) before native CUDA can drive it.

Existing VLM synthesis and native/CPU/CUDA pipeline decode parity remain green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Jul 31, 2026
…tive runtime (#67, #384) (#535)

## Summary

Closes the **loader/plumbing gap** so the Qwen3.5-0.8B hybrid
(Mamba/linear-attention) split package actually **loads and decodes
end-to-end**, composing the per-op CUDA coverage work (#480
CausalConvWithState, #484 LinearAttention, #525 RoPE-contrib + Bool
NonZero) into a working model.

### The blocker
The Foundry export is a 3-ONNX split package (`vision.onnx` +
`embedding.onnx` + `text.onnx`) whose declared image preprocessing uses
Qwen `smart_resize` — which has no lossless runtime encoding. That error
aborted the **entire** pipeline-metadata synthesis before admission, so
BOTH ORT and native `Engine::from_pipeline_dir` refused the package and
text decode could never run.

### The fix (general, modality-driven — not a model-name special-case)
Text decode never touches vision, so a split VLM package whose image
path is unusable is admitted for **text-only decode**:

1. **`GenAiConfigError::UnrepresentablePreprocessing`** — new distinct
variant from the `smart_resize` branch, kept separate from
`IncompletePipeline` so genuinely-incomplete packages still fail hard.
2. **`to_strict_text_only_pipeline_metadata`** — synthesizes an
embedding→decoder AR pipeline with no
vision/image-preprocessing/dataflow. Rank-3 positions use
`linear_increment` (every mrope axis advances with the sequence position
→ correct pure-text `[t,t,t]`); decoder declares `sequence_source:
inputs_embeds`; the vision-fed `image_features` embedding input becomes
optional with an empty (zero image-token) absent value.
3. **`pipeline_inference_metadata_from_dir`** falls back on the
unrepresentable-preprocessing signal; representable VLMs are unchanged.
4. **Symbolic-batch loop-state init** (shared ORT decode path,
`decode/values.rs` + `resolved_io.rs`): `conv_state`/`recurrent_state`
export the leading batch axis as `-1`; it now resolves to the decode
batch (1), mirroring the empty-KV convention. Non-batch symbolic dims
still refused loudly. **Not** in Mary's native step driver.

### Result — runs & coherent (ORT reference)
```
prompt : "The capital of France is"
output : " Paris, and the capital of Germany is Berlin.\nThe capital of France is"
```
Correct fact ("Paris") ⇒ positions / sequence-source / optional-image /
loop-state are all correct. Locked by an **active** regression test
`qwen35_0_8b_hybrid_text_decode_e2e.rs` (skips gracefully when the model
dir is absent).

Reference caveat (honest): the reference is an ORT decode of the same
synthesized spec (ORT falls back to CPU for the `com.microsoft` hybrid
ops). Per-op CUDA↔reference parity is proven separately
(#480/#484/#525); no independent onnxruntime-genai oracle is wired, so
coherence is the mitigating oracle for a shared-spec bug.

### Native-CUDA last mile — HANDOFF to Mary (Inc3c)
Native decode can't yet drive this model: the native step driver
hardcodes **rank-2** `position_ids` (`native_decode/cuda.rs:248`,
`cpu.rs:203`) but the hybrid decoder declares **rank-3** mrope positions
→ `rank mismatch (graph declares rank 3, got 2)`. These are Mary's
active Inc3c files, so per the collision rule they were **not** edited.
Needed: build rank-3 mrope coordinates in the native step driver
honoring the pipeline `positions` program (as `decode/step.rs` already
does for ORT). Then flip the `qwen35_0_8b_hybrid_native_cuda_e2e`
harness (#529) to native-vs-ORT parity. Details in
`.squad/decisions/inbox/cohaagen-hybrid-loader.md`.

## Verification
- `qwen35_0_8b_hybrid_text_decode_e2e` — **1 passed** (active, coherent
+ exact greedy lock).
- `onnx-genai-genai-config` — **27 lib + 4 vlm_pipeline passed** (incl.
new text-only synthesis test; existing VLM synthesis unchanged).
- `onnx-genai-engine --lib` — **284 passed** (incl. 2 new symbolic-batch
state cases).
- Existing VLM pipeline decode parity —
`native_cuda_pipeline_decoder_parity`, `native_pipeline_decoder_parity`
**passed** (no regression from shared decode changes).
- `every_covered_op_has_a_conformance_entry` (coverage-of-coverage)
**passed** (CUDA_COVERED_OPS untouched).
- `cargo fmt --all --check` clean; `cargo clippy` clean for
`onnx-genai-genai-config`, `onnx-genai-engine` (default) and
`onnx-genai-engine --features cuda,native-backend`.

Refs #67, #384. Do not merge.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant