Skip to content

feat(loader): text-only decode pipeline unblocks Qwen3.5 hybrid on native runtime (#67, #384) - #535

Merged
justinchuby merged 2 commits into
mainfrom
squad/qwen35-hybrid-loader-unblock
Jul 31, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/qwen35-hybrid-loader-unblock

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

Closes the loader/plumbing gap so the Qwen3.5-0.8B hybrid (Mamba/linear-attention) split package actually loads and decodes end-to-end, composing the per-op CUDA coverage work (#480 CausalConvWithState, #484 LinearAttention, #525 RoPE-contrib + Bool NonZero) into a working model.

The blocker

The Foundry export is a 3-ONNX split package (vision.onnx + embedding.onnx + text.onnx) whose declared image preprocessing uses Qwen smart_resize — which has no lossless runtime encoding. That error aborted the entire pipeline-metadata synthesis before admission, so BOTH ORT and native Engine::from_pipeline_dir refused the package and text decode could never run.

The fix (general, modality-driven — not a model-name special-case)

Text decode never touches vision, so a split VLM package whose image path is unusable is admitted for text-only decode:

  1. GenAiConfigError::UnrepresentablePreprocessing — new distinct variant from the smart_resize branch, kept separate from IncompletePipeline so genuinely-incomplete packages still fail hard.
  2. to_strict_text_only_pipeline_metadata — synthesizes an embedding→decoder AR pipeline with no vision/image-preprocessing/dataflow. Rank-3 positions use linear_increment (every mrope axis advances with the sequence position → correct pure-text [t,t,t]); decoder declares sequence_source: inputs_embeds; the vision-fed image_features embedding input becomes optional with an empty (zero image-token) absent value.
  3. pipeline_inference_metadata_from_dir falls back on the unrepresentable-preprocessing signal; representable VLMs are unchanged.
  4. Symbolic-batch loop-state init (shared ORT decode path, decode/values.rs + resolved_io.rs): conv_state/recurrent_state export the leading batch axis as -1; it now resolves to the decode batch (1), mirroring the empty-KV convention. Non-batch symbolic dims still refused loudly. Not in Mary's native step driver.

Result — runs & coherent (ORT reference)

prompt : "The capital of France is"
output : " Paris, and the capital of Germany is Berlin.\nThe capital of France is"

Correct fact ("Paris") ⇒ positions / sequence-source / optional-image / loop-state are all correct. Locked by an active regression test qwen35_0_8b_hybrid_text_decode_e2e.rs (skips gracefully when the model dir is absent).

Reference caveat (honest): the reference is an ORT decode of the same synthesized spec (ORT falls back to CPU for the com.microsoft hybrid ops). Per-op CUDA↔reference parity is proven separately (#480/#484/#525); no independent onnxruntime-genai oracle is wired, so coherence is the mitigating oracle for a shared-spec bug.

Native-CUDA last mile — HANDOFF to Mary (Inc3c)

Native decode can't yet drive this model: the native step driver hardcodes rank-2 position_ids (native_decode/cuda.rs:248, cpu.rs:203) but the hybrid decoder declares rank-3 mrope positions → rank mismatch (graph declares rank 3, got 2). These are Mary's active Inc3c files, so per the collision rule they were not edited. Needed: build rank-3 mrope coordinates in the native step driver honoring the pipeline positions program (as decode/step.rs already does for ORT). Then flip the qwen35_0_8b_hybrid_native_cuda_e2e harness (#529) to native-vs-ORT parity. Details in .squad/decisions/inbox/cohaagen-hybrid-loader.md.

Verification

  • qwen35_0_8b_hybrid_text_decode_e2e — 1 passed (active, coherent + exact greedy lock).
  • onnx-genai-genai-config — 27 lib + 4 vlm_pipeline passed (incl. new text-only synthesis test; existing VLM synthesis unchanged).
  • onnx-genai-engine --lib — 284 passed (incl. 2 new symbolic-batch state cases).
  • Existing VLM pipeline decode parity — native_cuda_pipeline_decoder_parity, native_pipeline_decoder_parity passed (no regression from shared decode changes).
  • every_covered_op_has_a_conformance_entry (coverage-of-coverage) passed (CUDA_COVERED_OPS untouched).
  • cargo fmt --all --check clean; cargo clippy clean for onnx-genai-genai-config, onnx-genai-engine (default) and onnx-genai-engine --features cuda,native-backend.

Refs #67, #384. Do not merge.

@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 81.05263% with 36 lines in your changes missing coverage. Please review.
✅ Project coverage is 81.30%. Comparing base (0c66ff0) to head (989cb85).
⚠️ Report is 4 commits behind head on main.

Files with missing lines Patch % Lines
...rates/onnx-genai-genai-config/src/compatibility.rs 80.97% 17 Missing and 18 partials ⚠️
crates/onnx-genai-genai-config/src/loading.rs 75.00% 0 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #535      +/-   ##
==========================================
+ Coverage   80.59%   81.30%   +0.71%     
==========================================
  Files         315      315              
  Lines      123259   123446     +187     
  Branches   123259   123446     +187     
==========================================
+ Hits        99337   100372    +1035     
+ Misses      19890    19017     -873     
- Partials     4032     4057      +25     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.78% <ø> (+0.10%) ⬆️
mlas 77.91% <ø> (ø)
offline 81.27% <81.05%> (+0.75%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...rates/onnx-genai-genai-config/src/json_builders.rs 86.64% <100.00%> (ø)
crates/onnx-genai-genai-config/src/loading.rs 34.24% <75.00%> (+1.38%) ⬆️
...rates/onnx-genai-genai-config/src/compatibility.rs 80.32% <80.97%> (+0.21%) ⬆️

... and 6 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

…packages (#67, #384)

A split VLM package (vision+embedding+decoder) whose declared image
preprocessing is not representable by the runtime (Qwen `smart_resize`)
previously aborted the entire pipeline-metadata synthesis, so BOTH ORT and
native `Engine::from_pipeline_dir` refused it and text decode could never run.
This blocked the Qwen3.5-0.8B hybrid (Mamba/linear-attention) model whose per-op
CUDA coverage already landed (#480/#484/#525).

Text never touches vision, so admit such a package for text-only decode, driven
purely by its declared modality shape (not a model name):

- New `GenAiConfigError::UnrepresentablePreprocessing`, returned by the
  smart_resize branch, kept distinct from `IncompletePipeline` so genuinely
  incomplete packages still fail hard.
- `to_strict_text_only_pipeline_metadata` synthesizes an embedding->decoder AR
  pipeline with no vision/image-preprocessing/image-dataflow. Rank-3 positions
  use `linear_increment` (every mrope axis advances with the sequence position,
  the correct pure-text coordinates); decoder declares `sequence_source:
  inputs_embeds`; the vision-fed `image_features` embedding input becomes
  optional with an empty (zero image-token) absent value.
- `pipeline_inference_metadata_from_dir` falls back to it on the
  unrepresentable-preprocessing signal; representable VLMs are unchanged.

Also resolve a symbolic leading (batch) axis when zero-initializing loop-carried
fixed state (`conv_state`, `recurrent_state` export batch as -1), mirroring the
empty-KV convention; non-batch symbolic dims are still refused loudly. This is in
the shared ORT decode path (decode/values.rs, resolved_io.rs), not the native
step driver.

Result: the real qwen3.5-0.8b hybrid now loads and greedy-decodes coherently
end-to-end via ORT ("The capital of France is" -> " Paris, and the capital of
Germany is Berlin."), locked by an active regression test that skips gracefully
when the model dir is absent. Native-CUDA decoder parity is a documented handoff:
the native step driver hardcodes rank-2 position_ids and must build rank-3 mrope
coordinates (Mary's Inc3c native_decode files) before native CUDA can drive it.

Existing VLM synthesis and native/CPU/CUDA pipeline decode parity remain green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

github-actions Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

✅ Benchmarks — No Regression

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
✅ tokenization/decode_tokens_per_second 6.11 ms 6.07 ms -0.8%
✅ grammar_masking/llguidance_compute_mask/32 75.34 µs 74.33 µs -1.3%
✅ sampling_latency/min_p_per_token 332.85 µs 327.66 µs -1.6%
✅ logit_processing/seven_processor_chain_per_step 1.17 ms 1.15 ms -1.8%
✅ sampling_latency/greedy_per_token 3.34 µs 3.27 µs -2.2%
✅ sampling_latency/top_k_per_token 478.08 µs 461.57 µs -3.5%
✅ tokenization/encode_tokens_per_second 378.03 µs 364.86 µs -3.5%
✅ matmul/small_generic_bf16_threads=8/1x256x256 37.30 µs 35.52 µs -4.8%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 84.30 µs 78.37 µs -7.0%
✅ kv_cache/alloc_dealloc_pages 39.45 µs 36.29 µs -8.0%
✅ matmul/small_generic_f16_threads=1/1x256x256 32.83 µs 30.00 µs -8.6%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.12 ms 9.16 ms -9.5%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.15 ms 1.91 ms -11.0%
✅ add/medium_f16_threads=1-internal/262144 2.88 ms 2.55 ms -11.2%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 589.01 µs 518.46 µs -12.0%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.58 ms 2.26 ms -12.3%
✅ sampling_latency/top_p_per_token 1.11 ms 966.18 µs -13.1%
🟢 add/small_f16_threads=1-internal/1024 14.81 µs 12.52 µs -15.5%
🟢 matmul/large_generic_f16_threads=8/32x1024x1024 100.70 µs 83.74 µs -16.8%
🟢 matmul/medium_generic_f16_threads=1/32x512x512 36.89 µs 30.59 µs -17.1%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 296.14 µs 241.41 µs -18.5%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 1.58 ms 1.27 ms -19.5%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 39.70 µs 31.50 µs -20.6%
🟢 reduce_mean/small_f32_threads=1-internal/4096 18.80 µs 14.82 µs -21.2%
🟢 matmul/large_generic_f32_threads=8/32x1024x1024 4.81 ms 3.77 ms -21.6%
🟢 add/medium_bf16_threads=1-internal/262144 3.29 ms 2.56 ms -22.1%
🟢 add/medium_f32_threads=1-internal/262144 3.21 ms 2.50 ms -22.1%
🟢 add/large_bf16_threads=1-internal/4194304 52.68 ms 40.87 ms -22.4%
🟢 add/large_f32_threads=1-internal/4194304 52.94 ms 39.21 ms -25.9%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.34 ms 991.28 µs -26.3%
🟢 add/large_f16_threads=1-internal/4194304 54.49 ms 40.13 ms -26.4%
🟢 matmul/medium_generic_bf16_threads=8/32x512x512 522.82 µs 376.72 µs -27.9%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 41.24 µs 29.57 µs -28.3%
🟢 add/small_bf16_threads=1-internal/1024 18.19 µs 12.91 µs -29.0%
🟢 add/small_f32_threads=1-internal/1024 308.4 ns 193.3 ns -37.3%
🟢 matmul/small_generic_f32_threads=1/1x256x256 58.13 µs 36.38 µs -37.4%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.46 ms 903.62 µs -38.2%
🟢 gather/medium_bf16_threads=1-internal/32768 3.76 µs 2.32 µs -38.4%
🟢 gather/small_f32_threads=1-internal/4096 1.03 µs 623.5 ns -39.3%
🟢 gather/medium_f16_threads=1-internal/32768 4.06 µs 2.28 µs -43.8%
🟢 gather/large_bf16_threads=1-internal/131072 22.25 µs 12.18 µs -45.2%
🟢 gather/small_f16_threads=1-internal/4096 931.3 ns 457.5 ns -50.9%
🟢 gather/large_f32_threads=1-internal/131072 72.59 µs 29.56 µs -59.3%
🟢 gather/large_f16_threads=1-internal/131072 27.43 µs 11.12 µs -59.5%
🟢 gather/small_bf16_threads=1-internal/4096 1.18 µs 461.8 ns -60.9%
🟢 matmul/small_generic_f16_threads=8/1x256x256 86.47 µs 31.71 µs -63.3%
🟢 gather/medium_f32_threads=1-internal/32768 11.16 µs 3.68 µs -67.0%
🟢 matmul/small_generic_f32_threads=8/1x256x256 956.81 µs 34.66 µs -96.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.20 4.30 7.25 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

…unblock (#67, #384)

With the loader-unblock fix the qwen3.5-0.8b hybrid ORT reference now decodes,
so the #529 native-CUDA e2e harness's skip-on-reference-error guard is stale:
it would drive into the native forward and fail on the rank-3 mrope position_ids
gap (`graph declares rank 3, got 2`), which lives in the native decode step
driver (native_decode/{load,cuda,cpu}.rs) — a separate owner's active files.

Flip the harness to auto-activation instead: take the working ORT reference,
run native, and gracefully skip on exactly the sanctioned native rank-3
position_ids gap (is_native_rank3_position_gap); every other native error
propagates. It enforces native-CUDA<->ORT token parity the instant the native
step driver constructs rank-3 positions, with zero further edits. Documents the
precise handoff in the decision note.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the squad/qwen35-hybrid-loader-unblock branch from 5741c39 to 989cb85 Compare July 31, 2026 00:42
@justinchuby

Copy link
Copy Markdown
Owner Author

Update — rebased onto merged #529 + #529 harness flipped to AUTO-ACTIVATION

Rebased this branch onto fresh main (now includes #529's qwen35_0_8b_hybrid_native_cuda_e2e harness + qwen35_0_8b_placement_lock, and #531). Clean rebase.

Empirical GPU run (device 0) of the #529 native-CUDA harness on this branch:

qwen3.5-0.8b hybrid ORT reference: 16 tokens = [11751, 11, 321, 279, 6511, 314, 9564, 369, 19241, 13, 198, 760, 6511, 314, 9338, 369]
native CUDA decoder forward pass failed: input position_ids: rank mismatch (graph declares rank 3, got 2)

So the loader fix works end-to-end (ORT reference now decodes), but the native decoder step driver supplies rank-2 positions while this hybrid declares rank-3 mrope position_ids. That rank-3 construction lives in native_decode/{load,cuda,cpu}.rs — a separate owner's active files (collision boundary), so I did not edit them and did not flip the harness to unconditional-active (it would be RED).

What I did instead: flipped the #529 harness to auto-activation — it takes the working ORT reference, runs native, and gracefully skips on exactly the sanctioned native rank-3 position_ids gap (is_native_rank3_position_gap); every other native error propagates. The harness enforces native-CUDA↔ORT token-for-token parity the instant the native step driver builds rank-3 positions, with zero further edits.

Verification (all green): qwen35_0_8b_hybrid_text_decode_e2e (active ORT lock) ✅; qwen35_0_8b_hybrid_native_cuda_e2e auto-skips green ✅; fmt --all --check ✅; clippy cuda / cuda,native-backend / genai-config default ✅; genai-config 4 pipeline tests ✅.

Native rank-3 mrope position construction is handed off (see .squad/decisions/inbox/cohaagen-hybrid-loader.md).

@justinchuby

Copy link
Copy Markdown
Owner Author

VERDICT: APPROVE

Independent review by Melina (reviewer; author Cohaagen locked out per reviewer-protocol). Reviewed at HEAD 989cb85 on a detached worktree off origin/squad/qwen35-hybrid-loader-unblock, device 5, ORT 1.27.0. Every claim was verified empirically against the real qwen3.5-0.8b hybrid model, not taken from the description.

B. Skip-gate narrowness [most important] — PROVEN TIGHT

is_native_rank3_position_gap requires BOTH substrings: message.contains("position_ids") AND message.contains("rank mismatch"). I did not just read it — I perturbed the native bind path to prove narrowness:

  • Injected a DIFFERENT error CLASS on the SAME input: forced a DtypeMismatch on position_ids (message: "input position_ids: dtype mismatch ...", contains "position_ids" but NOT "rank mismatch"). Result: the harness HARD-FAILED (exit 101, test FAILED) — the different error propagated, was NOT swallowed.
  • Reverted: the genuine rank-3 gap ("input position_ids: rank mismatch (graph declares rank 3, got 2)", origin onnx-runtime-session/executor/bindings.rs:46) skips gracefully as intended.

So the gate self-enforces native<->ORT parity the instant the driver builds rank-3 positions; a real native regression cannot hide behind it. Only a genuine RankMismatch on the position_ids input matches (a rank mismatch on any other input reads "input : ..." with no "position_ids" and would propagate). This satisfies the "don't report skipped-as-passing" principle (cf. #492). Note: the substrings are literal, so an unrelated error whose text happens to contain both phrases would match — not a realistic native failure mode, acceptable.

A. Loader admission safety — no vision regression

Fallback fires only on GenAiConfigError::UnrepresentablePreprocessing (the smart_resize branch). A representable VLM still admits through to_strict_pipeline_metadata FIRST and is unchanged: complete_config_synthesizes_typed_vlm_pipeline still passes. New vlm-smart-resize fixture is additive; the existing test was renamed (processor_signals_unrepresentable...) but its assertions were strengthened, not weakened. onnx-genai-genai-config: 27 lib + 4 vlm_pipeline PASS.

C. ORT reference lock is real

qwen35_0_8b_hybrid_text_decode_is_coherent_and_locked runs a real ~4s decode: output " Paris, and the capital of Germany is Berlin.\nThe capital of France is", coherence oracle (contains "Paris") + exact 16-token greedy lock both assert on real decoded tokens. Honest, disclosed caveat: ORT falls back to CPU for the com.microsoft hybrid ops; coherence is the mitigating oracle since no independent genai oracle is wired. Acceptable.

D. decode/values.rs — no clone_value overlap

The change is symbolic-batch loop-state init (new concrete_fixed_state_shape resolves a symbolic leading batch axis to 1, mirroring empty-KV; non-batch symbolic dims still refused loudly). clone_value (values.rs:185) is NOT touched by this PR. Only same-FILE proximity with the other agent's clone_value generalization branch — a possible textual merge conflict but no logical overlap. Sequence the merges; no code concern.

E. Regressions / hygiene

  • Mary's native driver (native_decode/{cuda,cpu,load}.rs) is UNTOUCHED — collision correctly avoided; last mile handed off via .squad/decisions/inbox/cohaagen-hybrid-loader.md (committed).
  • decode lib tests: 27/27 incl. new fixed_state_zero_initialization_resolves_symbolic_batch_axis (symbolic batch -> 1; symbolic non-batch refused).
  • qwen35_0_8b_hybrid_native_cuda harness: ORT reference decodes, native skips on the sanctioned rank-3 gap.
  • cargo fmt --all --check: clean.
  • cargo clippy clean on all three claimed sets: onnx-genai-genai-config, onnx-genai-engine (default), onnx-genai-engine --features cuda,native-backend.
  • No pre-existing failures encountered.

Merge note

The PR body says "Do not merge" (refs #67/#384 still in progress; native-CUDA last mile is Mary's follow-up). Code is correct and approved on its merits; respect the author/coordinator do-not-merge hold until the native rank-3 driver work lands and this harness auto-activates into a hard parity lock. Native rank-3 mrope position follow-up: Mary (already handed off) — not a revision of this PR.

@justinchuby
justinchuby marked this pull request as ready for review July 31, 2026 01:26
@justinchuby
justinchuby merged commit 15d2745 into main Jul 31, 2026
14 checks passed
@justinchuby
justinchuby deleted the squad/qwen35-hybrid-loader-unblock branch July 31, 2026 01:26
justinchuby added a commit that referenced this pull request Jul 31, 2026
Scribe round 7 bookkeeping (docs-only, no code):

- Merged **30** decision-inbox notes into `.squad/decisions.md` (20389 →
20458 B, under the 20480 gate); round-7 per-PR narrative archived
verbatim to `decisions-archive/2026-07.md`.
- Logged the #535 / #540 / #541 / #543 wave into
mary/harry/cohaagen/melina history.
- Cleared the processed inbox (README kept).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 3, 2026
…+ GAP-3 decomposition (increment 1/N) (#613)

## What & why

Scope pass on **GAP-3 (native pipeline decode)**. The finding is that
GAP-3's core is
**already implemented and conformance-locked on `origin/main`** — the
task was authored
against a local checkout (`1ba215ee`) that is ~35+ merged PRs behind
`origin/main`
(`be6d4e34`, #612). So this PR does **not** add new decode
functionality. It is a
bounded, **zero-behavior-change** truth-up + a design/decomposition
drop.

This is **GAP-3 housekeeping increment 1 of N** (see the design note in
this diff:

`.squad/decisions/inbox/cohaagen-gap3-native-pipeline-decode-design.md`).
The
substantive increments already landed:

- Backend-neutral component ownership seam — #546
- `NativePipelineDecoder` via `PipelineDecoderComponent` (Inc2b) — #479
- Pure-native multi-component decode wiring (Inc-A) — #565
- Native present-KV mirroring, paged (Inc-C) — #566
- Device-resident present-KV read-out (Inc-D / D.1) — #567/#568
- rank-3 mrope native positions — #543 · text-only decode pipeline —
#535 · fp16 TopK MoE router — #612

The pipeline decode loop (`PipelineDecodeLoopBackend`) already owns
`Box<dyn PipelineDecoderComponent>` + `Box<dyn ComponentSession>`, not
an ORT `Session`;
`DecodeState`/ORT `Value` are confined to `OrtPipelineDecoder`. The
Qwen3.5-0.8B hybrid
(same class as Qwen3.6-35B-A3B) decodes natively with **token-for-token
parity vs ORT**
under `tests/qwen35_0_8b_hybrid_native_cuda_e2e.rs`.

## The change

`native_component.rs`'s module doc still claimed wiring native sessions
into *"the
ORT-owned pipeline decode loop is the remaining GAP 3 work"* — false
since #546/#565.
Corrected to describe the now-backend-neutral loop and the merged
Inc-A/C/D, and to name
the genuinely-remaining feature-sized gaps (non-flat plans, native
cross-attn/vision KV).

**Behavior-preserving:** comment-only in `native_component.rs` + a
tracked decision drop.
No code path, signature, or data change.

## Remaining decomposition (in the design note)

Each is feature-sized (needs op/attention support and/or fixtures —
**not** zero-behavior),
none blocks the text-only 35B-A3B native number: R2 native
sliding-window paged mirror ·
R3 Inc-D.2 discontinuous prefix reuse · R4 native cross-attn/vision KV
(Inc3) · R5 non-flat
plans native · R6 (optional) neutral host tensor in the shared pool.

## Verification for the first native pipeline model

Byte/token-exact differential vs an ORT-backend decode of the **same
artifact** (ORT
front-end for both arms, decoder EP isolated), greedy, token-for-token —
already in place
for the 0.8B hybrid; same harness pattern applied to the real 35B-A3B is
the next step.

## Checks

- `cargo fmt --all` clean
- `cargo clippy -p onnx-genai-engine --features "native-backend cuda" --
-D warnings` clean

Left **open for Harry review**; do not merge without it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant