Skip to content

test(engine): qwen3.5-0.8b text export — native whole-graph capture + byte-identity lock (hybrid moat holds at the small end) - #1456

Merged
justinchuby merged 1 commit into
mainfrom
squad/qwen35-0.8b-text-export
Aug 19, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/qwen35-0.8b-text-export

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

Completes the qwen3.5 hybrid context-scaling moat family lock trio — 0.8B (this PR) + 2B (#1444, 3.03× deep) + 9B (#1449, 1.32× deep) — a strong RULES §2 generality statement: the moat is a family property, proven from 0.8B up.

The generality question answered: the moat holds at the small end. 0.8B is not too small for the graph-block to dominate — ORT-CUDA is graph-blocked identically to 2B/9B (same 25 Memcpy nodes), while native captures the whole hybrid graph with fallbacks=0.

Adds crates/onnx-genai-engine/tests/qwen35_0_8b_text_decode_lock.rs.

Model

Hybrid decoder, 24 layers = 6 periodic full-attention GroupQueryAttention layers (indices 3,7,11,15,19,23; head_dim 256, 8 q-heads / 2 kv-heads) + 18 gated linear-attention / short-conv recurrent layers (conv_state + recurrent_state). Per-layer KV-vs-recurrent roles resolved structurally by the loader from the graph port inventory (RULES.md §2, never model-name-gated).

Export (persisted, self-contained)

QWEN35_0_8B_TEXT_DIR=/home/justinchu/qwen35-0.8b-text-cuda

Foundry ships qwen3.5-0.8b as a multimodal split package (embedding.onnx + text.onnx + vision.onnx) the native single-model loader rejects (and that segfaults onnxruntime-genai on load). Composed a standalone input_ids→logits graph by pruning the embedding subgraph to its GatherBlockQuantized text gather (dropping the image-token Equal/NonZero/ScatterND merge — a no-op with no image tokens, and whose dynamic NonZero native shape-inference cannot bound) and fusing it into the text decoder via inputs_embeds. Same playbook as #1449 / #1442. Real weight copies in-dir (0.13 GB + 0.73 GB). No Mobius change needed. Reproducible via qwen35-0.8b-text-cuda/export_qwen35_0_8b_text.py.

Result

  • Native e2e (Y): "The capital of France is" → " Paris, and the capital of Germany is Berlin.\nThe capital of France is". Whole-graph CUDA capture (captures=7 replays=433 fallbacks=0).
  • Byte-identity: all 16 tokens identical to the independently-validated ORT-driven reference — the split-package qwen35_0_8b_hybrid_text_decode_e2e lock (ORT places standard attention on its EP, CPU-falls-back the hybrid ops) decodes the exact same stream.
  • Golden lock PASS (GPU1 24.2s; re-verified GPU2 24.2s).

Moat (graph-block, confirmed at 0.8B)

ORT-CUDA on this exact export inserts 25 Memcpy nodes ("25 Memcpy nodes are added ... including unable to run CUDA graph") and cannot even initialize a session — nodes are forced to the default CPU EP and session-init hard-fails. ORT has no runnable whole-graph GPU path; the recurrent LinearAttention/CausalConvWithState ops break CUDA-graph placement exactly as at the arch level (1037 CUDA / 56 CPU, 25 Memcpy). Native captures the whole hybrid graph and stays context-flat.

A/B (H200, GPU1 pinned, --steady, native --ep cuda, medians)

context depth native-CUDA (graph=1) ORT-CUDA (best GPU path)
short (~64 tok) 121.6 tok/s (8.22 ms/tok, 5 runs) cannot run — 25 Memcpy, CUDA graph disabled, session-init fails
mid (~256–320) 114.5 tok/s (8.74 ms/tok) cannot run
deep (~1024) 113.6 tok/s (8.81 ms/tok, 5 runs) cannot run

Native decode is near context-flat (~6–7% short→deep, then plateaus). ORT-CUDA's tok/s is undefined (no GPU-runnable path) — the moat at 0.8B is categorical (graph-blocked before it can decode a single token on GPU), the strongest form.

(An ORT-CPU baseline was not pursued: the composed export's minimal inference_metadata.yaml doesn't declare the rank-3 mrope pipeline.positions the ORT-genai loader requires — that spec only lives in the full split-package genai_config and is irrelevant to the native golden. A native-GPU-vs-ORT-CPU ratio is not the fair moat comparison anyway.)

Verdict

GO. Do NOT self-merge — awaiting coordinator validation.

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

… byte-identity lock (hybrid moat holds at the small end)

Completes the qwen3.5 hybrid context-scaling moat family lock trio
(0.8b + 2b #1444 + 9b #1449). Native loads a composed text-only
qwen3.5-0.8b graph, captures the whole hybrid into one CUDA graph
(fallbacks=0, context-flat), and decodes byte-identical to the ORT
reference — while ORT-CUDA structurally cannot build a session (25
Memcpy nodes, CUDA graph disabled), the same graph-block confirmed at
the arch level. The moat holds at the smallest family scale: 0.8b is
not too small for the graph-block to dominate.

Model: hybrid decoder = 6 periodic full-attention GroupQueryAttention
layers (head_dim 256, 8 q-heads / 2 kv-heads) + 18 gated
linear-attention / short-conv recurrent layers (conv_state +
recurrent_state). Per-layer KV-vs-recurrent roles resolved structurally
by the loader from the graph port inventory (RULES.md §2, never
model-name-gated).

Export (persisted, self-contained):
QWEN35_0_8B_TEXT_DIR=/home/justinchu/qwen35-0.8b-text-cuda — a standalone
input_ids->logits graph composed by pruning the embedding subgraph to
its GatherBlockQuantized text gather (dropping the image-token
Equal/NonZero/ScatterND merge) and fusing it into the text decoder via
inputs_embeds. Same playbook as #1449 / #1442.

A/B (H200, native CUDA, --steady medians): short(~64) 121.6 tok/s /
mid(~256) 114.5 / deep(~1024) 113.6 — near context-flat. ORT-CUDA: no
runnable whole-graph GPU path (session-init fails).

Adds crates/onnx-genai-engine/tests/qwen35_0_8b_text_decode_lock.rs.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 169febb into main Aug 19, 2026
4 checks passed
@justinchuby
justinchuby deleted the squad/qwen35-0.8b-text-export branch August 19, 2026 11:22
justinchuby added a commit that referenced this pull request Aug 19, 2026
…#1457)

Records #1456 (Deckard's 0.8b text-only export, merged @ 169febb) and
marks the hybrid graph-block moat family lock trio {0.8b, 2b, 9b}
complete.

**Headline:** the moat is CATEGORICAL at the small end — ORT-CUDA cannot
even initialize a session on the composed 0.8b export (25 Memcpy → CPU
EP → session-init hard-fail); native whole-graph-captures byte-identical
(16/16 vs golden). All three scales now golden-locked — a RULES §2
generality statement that the moat is an architectural family property.

Docs-only. Co-authored-by: Copilot.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
`cargo fmt --all -- --check` runs in *both* required jobs (Fast (Linux
x86_64) and Rust quality), so any surviving fmt diff on main blocks every
PR. #1456 landed two trailing blank lines at the end of this file after
this branch was written; without this hunk the merge result of this PR
still fails fmt and the branch cannot go green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 56.56 µs 128.04 µs +126.4%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 29.71 µs 54.64 µs +83.9%
🔴 tokenization/decode_tokens_per_second 6.22 ms 11.16 ms +79.4%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 46.44 µs 76.76 µs +65.3%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 360.30 µs 567.37 µs +57.5%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 43.15 µs 66.81 µs +54.9%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.33 ms 1.98 ms +48.1%
🔴 matmul/medium_generic_bf16_threads=1/32x512x512 536.19 µs 786.40 µs +46.7%
🔴 qwen3_sampling_processors/top_k_partial_selection 161.74 µs 229.18 µs +41.7%
🔴 grammar_masking/llguidance_compute_mask/32 84.64 µs 119.34 µs +41.0%
🔴 matmul/medium_generic_f16_threads=1/32x512x512 30.95 µs 43.44 µs +40.4%
🔴 gather/large_f16_threads=1-internal/131072 15.02 µs 19.94 µs +32.8%
🔴 sampling_latency/greedy_per_token 3.29 µs 4.36 µs +32.3%
⚠️ logit_processing/seven_processor_chain_per_step 343.45 µs 439.56 µs +28.0%
⚠️ matmul/large_generic_f32_threads=1/32x1024x1024 9.37 ms 11.56 ms +23.4%
⚠️ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 544.10 µs 665.59 µs +22.3%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.24 ms 5.16 ms +21.8%
⚠️ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 405.51 µs 492.18 µs +21.4%
⚠️ qwen3_sampling_processors/top_k_top_p_fast 670.83 µs 805.15 µs +20.0%
⚠️ qwen3_sampling_processors/top_p_fast_after_top_k 557.32 µs 644.16 µs +15.6%
✅ add/small_f16_threads=1-internal/1024 478.5 ns 549.5 ns +14.8%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.01 ms 1.15 ms +14.2%
✅ matmul/small_generic_bf16_threads=8/1x256x256 31.20 µs 35.55 µs +14.0%
✅ kv_cache/alloc_dealloc_pages 46.01 µs 52.44 µs +14.0%
✅ tokenization/encode_tokens_per_second 434.80 µs 492.96 µs +13.4%
✅ gather/large_bf16_threads=1-internal/131072 12.98 µs 14.47 µs +11.5%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.63 ms 7.33 ms +10.6%
✅ add/small_bf16_threads=1-internal/1024 479.2 ns 527.1 ns +10.0%
✅ matmul/small_generic_f16_threads=1/1x256x256 32.58 µs 35.32 µs +8.4%
✅ matmul/small_generic_bf16_threads=1/1x256x256 33.00 µs 35.75 µs +8.3%
✅ matmul/small_generic_f16_threads=8/1x256x256 34.09 µs 36.58 µs +7.3%
✅ gather/small_f16_threads=1-internal/4096 488.8 ns 521.7 ns +6.7%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.30 ms 2.44 ms +6.5%
✅ sampling_latency/top_p_per_token 493.82 µs 523.84 µs +6.1%
✅ matmul/small_generic_f32_threads=8/1x256x256 38.65 µs 40.80 µs +5.6%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.01 ms 2.08 ms +3.8%
✅ sampling_latency/top_k_per_token 61.56 µs 63.43 µs +3.0%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 86.65 µs 88.59 µs +2.2%
✅ gather/small_bf16_threads=1-internal/4096 542.5 ns 529.7 ns -2.4%
✅ gather/small_f32_threads=1-internal/4096 697.3 ns 678.1 ns -2.7%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 92.40 µs 89.40 µs -3.3%
✅ gather/large_f32_threads=1-internal/131072 37.97 µs 36.59 µs -3.6%
✅ add/medium_f32_threads=1-internal/262144 28.35 µs 27.26 µs -3.8%
✅ matmul/small_generic_f32_threads=1/1x256x256 43.48 µs 41.76 µs -4.0%
✅ gather/medium_f16_threads=1-internal/32768 2.59 µs 2.47 µs -4.4%
✅ gather/medium_f32_threads=1-internal/32768 4.39 µs 4.13 µs -6.1%
✅ add/large_bf16_threads=1-internal/4194304 1.95 ms 1.81 ms -6.8%
✅ gather/medium_bf16_threads=1-internal/32768 2.74 µs 2.53 µs -7.6%
✅ reduce_mean/small_f32_threads=1-internal/4096 17.16 µs 15.72 µs -8.4%
✅ sampling_latency/min_p_per_token 326.33 µs 294.76 µs -9.7%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 6.16 ms 5.38 ms -12.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 3.12 ms 2.70 ms -13.7%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.21 ms 1.04 ms -14.0%
✅ add/small_f32_threads=1-internal/1024 234.5 ns 201.2 ns -14.2%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 303.18 µs 256.84 µs -15.3%
🟢 add/medium_f16_threads=1-internal/262144 142.94 µs 111.41 µs -22.1%
🟢 add/large_f16_threads=1-internal/4194304 2.67 ms 1.68 ms -37.0%
🟢 add/medium_bf16_threads=1-internal/262144 173.76 µs 108.05 µs -37.8%
🟢 add/large_f32_threads=1-internal/4194304 1.35 ms 731.47 µs -45.7%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 7.67 5.18 6.21 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.64%. Comparing base (ced0085) to head (b6f5b55).
⚠️ Report is 78 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1456      +/-   ##
==========================================
+ Coverage   82.10%   82.64%   +0.54%     
==========================================
  Files          12       12              
  Lines        5471     5475       +4     
  Branches     5471     5475       +4     
==========================================
+ Hits         4492     4525      +33     
+ Misses        780      757      -23     
+ Partials      199      193       -6     
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (?)
cli-ort-windows 82.10% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant