Repository navigation
feat(loader): text-only decode pipeline unblocks Qwen3.5 hybrid on native runtime (#67, #384) - #535
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #535 +/- ##
==========================================
+ Coverage 80.59% 81.30% +0.71%
==========================================
Files 315 315
Lines 123259 123446 +187
Branches 123259 123446 +187
==========================================
+ Hits 99337 100372 +1035
+ Misses 19890 19017 -873
- Partials 4032 4057 +25
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
…packages (#67, #384) A split VLM package (vision+embedding+decoder) whose declared image preprocessing is not representable by the runtime (Qwen `smart_resize`) previously aborted the entire pipeline-metadata synthesis, so BOTH ORT and native `Engine::from_pipeline_dir` refused it and text decode could never run. This blocked the Qwen3.5-0.8B hybrid (Mamba/linear-attention) model whose per-op CUDA coverage already landed (#480/#484/#525). Text never touches vision, so admit such a package for text-only decode, driven purely by its declared modality shape (not a model name): - New `GenAiConfigError::UnrepresentablePreprocessing`, returned by the smart_resize branch, kept distinct from `IncompletePipeline` so genuinely incomplete packages still fail hard. - `to_strict_text_only_pipeline_metadata` synthesizes an embedding->decoder AR pipeline with no vision/image-preprocessing/image-dataflow. Rank-3 positions use `linear_increment` (every mrope axis advances with the sequence position, the correct pure-text coordinates); decoder declares `sequence_source: inputs_embeds`; the vision-fed `image_features` embedding input becomes optional with an empty (zero image-token) absent value. - `pipeline_inference_metadata_from_dir` falls back to it on the unrepresentable-preprocessing signal; representable VLMs are unchanged. Also resolve a symbolic leading (batch) axis when zero-initializing loop-carried fixed state (`conv_state`, `recurrent_state` export batch as -1), mirroring the empty-KV convention; non-batch symbolic dims are still refused loudly. This is in the shared ORT decode path (decode/values.rs, resolved_io.rs), not the native step driver. Result: the real qwen3.5-0.8b hybrid now loads and greedy-decodes coherently end-to-end via ORT ("The capital of France is" -> " Paris, and the capital of Germany is Berlin."), locked by an active regression test that skips gracefully when the model dir is absent. Native-CUDA decoder parity is a documented handoff: the native step driver hardcodes rank-2 position_ids and must build rank-3 mrope coordinates (Mary's Inc3c native_decode files) before native CUDA can drive it. Existing VLM synthesis and native/CPU/CUDA pipeline decode parity remain green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
✅ Benchmarks — No RegressionComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
…unblock (#67, #384) With the loader-unblock fix the qwen3.5-0.8b hybrid ORT reference now decodes, so the #529 native-CUDA e2e harness's skip-on-reference-error guard is stale: it would drive into the native forward and fail on the rank-3 mrope position_ids gap (`graph declares rank 3, got 2`), which lives in the native decode step driver (native_decode/{load,cuda,cpu}.rs) — a separate owner's active files. Flip the harness to auto-activation instead: take the working ORT reference, run native, and gracefully skip on exactly the sanctioned native rank-3 position_ids gap (is_native_rank3_position_gap); every other native error propagates. It enforces native-CUDA<->ORT token parity the instant the native step driver constructs rank-3 positions, with zero further edits. Documents the precise handoff in the decision note. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
5741c39 to
989cb85
Compare
Update — rebased onto merged #529 + #529 harness flipped to AUTO-ACTIVATIONRebased this branch onto fresh Empirical GPU run (device 0) of the #529 native-CUDA harness on this branch: So the loader fix works end-to-end (ORT reference now decodes), but the native decoder step driver supplies rank-2 positions while this hybrid declares rank-3 mrope What I did instead: flipped the #529 harness to auto-activation — it takes the working ORT reference, runs native, and gracefully skips on exactly the sanctioned native rank-3 Verification (all green): Native rank-3 mrope position construction is handed off (see |
|
VERDICT: APPROVE Independent review by Melina (reviewer; author Cohaagen locked out per reviewer-protocol). Reviewed at HEAD 989cb85 on a detached worktree off origin/squad/qwen35-hybrid-loader-unblock, device 5, ORT 1.27.0. Every claim was verified empirically against the real qwen3.5-0.8b hybrid model, not taken from the description. B. Skip-gate narrowness [most important] — PROVEN TIGHTis_native_rank3_position_gap requires BOTH substrings: message.contains("position_ids") AND message.contains("rank mismatch"). I did not just read it — I perturbed the native bind path to prove narrowness:
So the gate self-enforces native<->ORT parity the instant the driver builds rank-3 positions; a real native regression cannot hide behind it. Only a genuine RankMismatch on the position_ids input matches (a rank mismatch on any other input reads "input : ..." with no "position_ids" and would propagate). This satisfies the "don't report skipped-as-passing" principle (cf. #492). Note: the substrings are literal, so an unrelated error whose text happens to contain both phrases would match — not a realistic native failure mode, acceptable. A. Loader admission safety — no vision regressionFallback fires only on GenAiConfigError::UnrepresentablePreprocessing (the smart_resize branch). A representable VLM still admits through to_strict_pipeline_metadata FIRST and is unchanged: complete_config_synthesizes_typed_vlm_pipeline still passes. New vlm-smart-resize fixture is additive; the existing test was renamed (processor_signals_unrepresentable...) but its assertions were strengthened, not weakened. onnx-genai-genai-config: 27 lib + 4 vlm_pipeline PASS. C. ORT reference lock is realqwen35_0_8b_hybrid_text_decode_is_coherent_and_locked runs a real ~4s decode: output " Paris, and the capital of Germany is Berlin.\nThe capital of France is", coherence oracle (contains "Paris") + exact 16-token greedy lock both assert on real decoded tokens. Honest, disclosed caveat: ORT falls back to CPU for the com.microsoft hybrid ops; coherence is the mitigating oracle since no independent genai oracle is wired. Acceptable. D. decode/values.rs — no clone_value overlapThe change is symbolic-batch loop-state init (new concrete_fixed_state_shape resolves a symbolic leading batch axis to 1, mirroring empty-KV; non-batch symbolic dims still refused loudly). clone_value (values.rs:185) is NOT touched by this PR. Only same-FILE proximity with the other agent's clone_value generalization branch — a possible textual merge conflict but no logical overlap. Sequence the merges; no code concern. E. Regressions / hygiene
Merge noteThe PR body says "Do not merge" (refs #67/#384 still in progress; native-CUDA last mile is Mary's follow-up). Code is correct and approved on its merits; respect the author/coordinator do-not-merge hold until the native rank-3 driver work lands and this harness auto-activates into a hard parity lock. Native rank-3 mrope position follow-up: Mary (already handed off) — not a revision of this PR. |
Scribe round 7 bookkeeping (docs-only, no code): - Merged **30** decision-inbox notes into `.squad/decisions.md` (20389 → 20458 B, under the 20480 gate); round-7 per-PR narrative archived verbatim to `decisions-archive/2026-07.md`. - Logged the #535 / #540 / #541 / #543 wave into mary/harry/cohaagen/melina history. - Cleared the processed inbox (README kept). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…+ GAP-3 decomposition (increment 1/N) (#613) ## What & why Scope pass on **GAP-3 (native pipeline decode)**. The finding is that GAP-3's core is **already implemented and conformance-locked on `origin/main`** — the task was authored against a local checkout (`1ba215ee`) that is ~35+ merged PRs behind `origin/main` (`be6d4e34`, #612). So this PR does **not** add new decode functionality. It is a bounded, **zero-behavior-change** truth-up + a design/decomposition drop. This is **GAP-3 housekeeping increment 1 of N** (see the design note in this diff: `.squad/decisions/inbox/cohaagen-gap3-native-pipeline-decode-design.md`). The substantive increments already landed: - Backend-neutral component ownership seam — #546 - `NativePipelineDecoder` via `PipelineDecoderComponent` (Inc2b) — #479 - Pure-native multi-component decode wiring (Inc-A) — #565 - Native present-KV mirroring, paged (Inc-C) — #566 - Device-resident present-KV read-out (Inc-D / D.1) — #567/#568 - rank-3 mrope native positions — #543 · text-only decode pipeline — #535 · fp16 TopK MoE router — #612 The pipeline decode loop (`PipelineDecodeLoopBackend`) already owns `Box<dyn PipelineDecoderComponent>` + `Box<dyn ComponentSession>`, not an ORT `Session`; `DecodeState`/ORT `Value` are confined to `OrtPipelineDecoder`. The Qwen3.5-0.8B hybrid (same class as Qwen3.6-35B-A3B) decodes natively with **token-for-token parity vs ORT** under `tests/qwen35_0_8b_hybrid_native_cuda_e2e.rs`. ## The change `native_component.rs`'s module doc still claimed wiring native sessions into *"the ORT-owned pipeline decode loop is the remaining GAP 3 work"* — false since #546/#565. Corrected to describe the now-backend-neutral loop and the merged Inc-A/C/D, and to name the genuinely-remaining feature-sized gaps (non-flat plans, native cross-attn/vision KV). **Behavior-preserving:** comment-only in `native_component.rs` + a tracked decision drop. No code path, signature, or data change. ## Remaining decomposition (in the design note) Each is feature-sized (needs op/attention support and/or fixtures — **not** zero-behavior), none blocks the text-only 35B-A3B native number: R2 native sliding-window paged mirror · R3 Inc-D.2 discontinuous prefix reuse · R4 native cross-attn/vision KV (Inc3) · R5 non-flat plans native · R6 (optional) neutral host tensor in the shared pool. ## Verification for the first native pipeline model Byte/token-exact differential vs an ORT-backend decode of the **same artifact** (ORT front-end for both arms, decoder EP isolated), greedy, token-for-token — already in place for the 0.8B hybrid; same harness pattern applied to the real 35B-A3B is the next step. ## Checks - `cargo fmt --all` clean - `cargo clippy -p onnx-genai-engine --features "native-backend cuda" -- -D warnings` clean Left **open for Harry review**; do not merge without it. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Summary
Closes the loader/plumbing gap so the Qwen3.5-0.8B hybrid (Mamba/linear-attention) split package actually loads and decodes end-to-end, composing the per-op CUDA coverage work (#480 CausalConvWithState, #484 LinearAttention, #525 RoPE-contrib + Bool NonZero) into a working model.
The blocker
The Foundry export is a 3-ONNX split package (
vision.onnx+embedding.onnx+text.onnx) whose declared image preprocessing uses Qwensmart_resize— which has no lossless runtime encoding. That error aborted the entire pipeline-metadata synthesis before admission, so BOTH ORT and nativeEngine::from_pipeline_dirrefused the package and text decode could never run.The fix (general, modality-driven — not a model-name special-case)
Text decode never touches vision, so a split VLM package whose image path is unusable is admitted for text-only decode:
GenAiConfigError::UnrepresentablePreprocessing— new distinct variant from thesmart_resizebranch, kept separate fromIncompletePipelineso genuinely-incomplete packages still fail hard.to_strict_text_only_pipeline_metadata— synthesizes an embedding→decoder AR pipeline with no vision/image-preprocessing/dataflow. Rank-3 positions uselinear_increment(every mrope axis advances with the sequence position → correct pure-text[t,t,t]); decoder declaressequence_source: inputs_embeds; the vision-fedimage_featuresembedding input becomes optional with an empty (zero image-token) absent value.pipeline_inference_metadata_from_dirfalls back on the unrepresentable-preprocessing signal; representable VLMs are unchanged.decode/values.rs+resolved_io.rs):conv_state/recurrent_stateexport the leading batch axis as-1; it now resolves to the decode batch (1), mirroring the empty-KV convention. Non-batch symbolic dims still refused loudly. Not in Mary's native step driver.Result — runs & coherent (ORT reference)
Correct fact ("Paris") ⇒ positions / sequence-source / optional-image / loop-state are all correct. Locked by an active regression test
qwen35_0_8b_hybrid_text_decode_e2e.rs(skips gracefully when the model dir is absent).Reference caveat (honest): the reference is an ORT decode of the same synthesized spec (ORT falls back to CPU for the
com.microsofthybrid ops). Per-op CUDA↔reference parity is proven separately (#480/#484/#525); no independent onnxruntime-genai oracle is wired, so coherence is the mitigating oracle for a shared-spec bug.Native-CUDA last mile — HANDOFF to Mary (Inc3c)
Native decode can't yet drive this model: the native step driver hardcodes rank-2
position_ids(native_decode/cuda.rs:248,cpu.rs:203) but the hybrid decoder declares rank-3 mrope positions →rank mismatch (graph declares rank 3, got 2). These are Mary's active Inc3c files, so per the collision rule they were not edited. Needed: build rank-3 mrope coordinates in the native step driver honoring the pipelinepositionsprogram (asdecode/step.rsalready does for ORT). Then flip theqwen35_0_8b_hybrid_native_cuda_e2eharness (#529) to native-vs-ORT parity. Details in.squad/decisions/inbox/cohaagen-hybrid-loader.md.Verification
qwen35_0_8b_hybrid_text_decode_e2e— 1 passed (active, coherent + exact greedy lock).onnx-genai-genai-config— 27 lib + 4 vlm_pipeline passed (incl. new text-only synthesis test; existing VLM synthesis unchanged).onnx-genai-engine --lib— 284 passed (incl. 2 new symbolic-batch state cases).native_cuda_pipeline_decoder_parity,native_pipeline_decoder_paritypassed (no regression from shared decode changes).every_covered_op_has_a_conformance_entry(coverage-of-coverage) passed (CUDA_COVERED_OPS untouched).cargo fmt --all --checkclean;cargo clippyclean foronnx-genai-genai-config,onnx-genai-engine(default) andonnx-genai-engine --features cuda,native-backend.Refs #67, #384. Do not merge.