Repository navigation
test(engine): qwen3.5-0.8b text export — native whole-graph capture + byte-identity lock (hybrid moat holds at the small end) - #1456
Merged
Conversation
… byte-identity lock (hybrid moat holds at the small end) Completes the qwen3.5 hybrid context-scaling moat family lock trio (0.8b + 2b #1444 + 9b #1449). Native loads a composed text-only qwen3.5-0.8b graph, captures the whole hybrid into one CUDA graph (fallbacks=0, context-flat), and decodes byte-identical to the ORT reference — while ORT-CUDA structurally cannot build a session (25 Memcpy nodes, CUDA graph disabled), the same graph-block confirmed at the arch level. The moat holds at the smallest family scale: 0.8b is not too small for the graph-block to dominate. Model: hybrid decoder = 6 periodic full-attention GroupQueryAttention layers (head_dim 256, 8 q-heads / 2 kv-heads) + 18 gated linear-attention / short-conv recurrent layers (conv_state + recurrent_state). Per-layer KV-vs-recurrent roles resolved structurally by the loader from the graph port inventory (RULES.md §2, never model-name-gated). Export (persisted, self-contained): QWEN35_0_8B_TEXT_DIR=/home/justinchu/qwen35-0.8b-text-cuda — a standalone input_ids->logits graph composed by pruning the embedding subgraph to its GatherBlockQuantized text gather (dropping the image-token Equal/NonZero/ScatterND merge) and fusing it into the text decoder via inputs_embeds. Same playbook as #1449 / #1442. A/B (H200, native CUDA, --steady medians): short(~64) 121.6 tok/s / mid(~256) 114.5 / deep(~1024) 113.6 — near context-flat. ORT-CUDA: no runnable whole-graph GPU path (session-init fails). Adds crates/onnx-genai-engine/tests/qwen35_0_8b_text_decode_lock.rs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
…#1457) Records #1456 (Deckard's 0.8b text-only export, merged @ 169febb) and marks the hybrid graph-block moat family lock trio {0.8b, 2b, 9b} complete. **Headline:** the moat is CATEGORICAL at the small end — ORT-CUDA cannot even initialize a session on the composed 0.8b export (25 Memcpy → CPU EP → session-init hard-fail); native whole-graph-captures byte-identical (16/16 vs golden). All three scales now golden-locked — a RULES §2 generality statement that the moat is an architectural family property. Docs-only. Co-authored-by: Copilot. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
`cargo fmt --all -- --check` runs in *both* required jobs (Fast (Linux x86_64) and Rust quality), so any surviving fmt diff on main blocks every PR. #1456 landed two trailing blank lines at the end of this file after this branch was written; without this hunk the merge result of this PR still fails fmt and the branch cannot go green. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1456 +/- ##
==========================================
+ Coverage 82.10% 82.64% +0.54%
==========================================
Files 12 12
Lines 5471 5475 +4
Branches 5471 5475 +4
==========================================
+ Hits 4492 4525 +33
+ Misses 780 757 -23
+ Partials 199 193 -6
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Completes the qwen3.5 hybrid context-scaling moat family lock trio — 0.8B (this PR) + 2B (#1444, 3.03× deep) + 9B (#1449, 1.32× deep) — a strong RULES §2 generality statement: the moat is a family property, proven from 0.8B up.
The generality question answered: the moat holds at the small end. 0.8B is not too small for the graph-block to dominate — ORT-CUDA is graph-blocked identically to 2B/9B (same 25 Memcpy nodes), while native captures the whole hybrid graph with
fallbacks=0.Adds
crates/onnx-genai-engine/tests/qwen35_0_8b_text_decode_lock.rs.Model
Hybrid decoder, 24 layers = 6 periodic full-attention
GroupQueryAttentionlayers (indices 3,7,11,15,19,23;head_dim 256, 8 q-heads / 2 kv-heads) + 18 gated linear-attention / short-conv recurrent layers (conv_state+recurrent_state). Per-layer KV-vs-recurrent roles resolved structurally by the loader from the graph port inventory (RULES.md §2, never model-name-gated).Export (persisted, self-contained)
QWEN35_0_8B_TEXT_DIR=/home/justinchu/qwen35-0.8b-text-cudaFoundry ships qwen3.5-0.8b as a multimodal split package (
embedding.onnx+text.onnx+vision.onnx) the native single-model loader rejects (and that segfaultsonnxruntime-genaion load). Composed a standaloneinput_ids→logitsgraph by pruning the embedding subgraph to itsGatherBlockQuantizedtext gather (dropping the image-tokenEqual/NonZero/ScatterNDmerge — a no-op with no image tokens, and whose dynamicNonZeronative shape-inference cannot bound) and fusing it into the text decoder viainputs_embeds. Same playbook as #1449 / #1442. Real weight copies in-dir (0.13 GB + 0.73 GB). No Mobius change needed. Reproducible viaqwen35-0.8b-text-cuda/export_qwen35_0_8b_text.py.Result
"The capital of France is"→" Paris, and the capital of Germany is Berlin.\nThe capital of France is". Whole-graph CUDA capture (captures=7 replays=433 fallbacks=0).qwen35_0_8b_hybrid_text_decode_e2elock (ORT places standard attention on its EP, CPU-falls-back the hybrid ops) decodes the exact same stream.Moat (graph-block, confirmed at 0.8B)
ORT-CUDA on this exact export inserts 25 Memcpy nodes (
"25 Memcpy nodes are added ... including unable to run CUDA graph") and cannot even initialize a session — nodes are forced to the default CPU EP and session-init hard-fails. ORT has no runnable whole-graph GPU path; the recurrentLinearAttention/CausalConvWithStateops break CUDA-graph placement exactly as at the arch level (1037 CUDA / 56 CPU, 25 Memcpy). Native captures the whole hybrid graph and stays context-flat.A/B (H200, GPU1 pinned,
--steady, native--ep cuda, medians)Native decode is near context-flat (~6–7% short→deep, then plateaus). ORT-CUDA's tok/s is undefined (no GPU-runnable path) — the moat at 0.8B is categorical (graph-blocked before it can decode a single token on GPU), the strongest form.
(An ORT-CPU baseline was not pursued: the composed export's minimal
inference_metadata.yamldoesn't declare the rank-3 mropepipeline.positionsthe ORT-genai loader requires — that spec only lives in the full split-package genai_config and is irrelevant to the native golden. A native-GPU-vs-ORT-CPU ratio is not the fair moat comparison anyway.)Verdict
GO. Do NOT self-merge — awaiting coordinator validation.
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com