Repository navigation
test(engine): gemma4-e2b dual head_dim (256+512) e2e golden lock — closes #1438 512 arm end-to-end - #1442
Merged
Conversation
…ng + 512 full-attn) Add a native CUDA greedy-decode coherence lock for the Gemma-3n / gemma4-e2b text decoder, which interleaves head_dim=256 sliding layers (x28) with head_dim=512 full-attention layers (x7). This is the end-to-end companion to the unit-lock in #1438 (MAX_HEAD_DIM 256->512): it proves the 512 full-attn layers run on the fused split-K decode kernel and stay coherent, decoding a short greedy stream that is byte-identical to the ORT CPU EP reference for the same export (16/16 tokens). Per-layer head sizes are resolved structurally by the engine KV bridge from each present.N.key output shape (RULES.md $2: head size is a fully runtime per-attention-op parameter), so this exercises mixed head sizes generically. The lock runs with CUDA-graph capture OFF: the composed text export's merged present-KV sequence axis is an opaque symbol the prefill workspace planner cannot yet upper-bound for capture (it trips on a head_dim=256 sliding layer, not the 512 layers). Eager runs the identical fused kernels, so the 512 fast-path coherence is fully locked; capture is a documented follow-up. Export (persisted, self-contained): GEMMA4_E2B_512_DIR=/home/justinchu/gemma4-e2b-it-text-cuda Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
Recovers coordinator `now.md` campaign-brain entries destroyed when an agent ran `git reset --hard origin/main` in the main checkout (uncommitted working-tree edits lost). Records the merged arc since #1383: **#1435** grid-fill narrow-N GEMV, **#1438** MAX_HEAD_DIM 256→512 (general), **#1442** gemma4-e2b dual head-size (256+512) e2e golden lock, **#1444** qwen3.5-2b-text context-scaling moat lock (3.03× deep), **#1445** capability-driven RMSNorm-fold gate + banked GLM structural verdict. Also updates the moat table (qwen3.5-2b recharacterized as a context-SCALING graph-block moat, not a fixed 1.65×) and the GLM verdict (structurally ORT-ahead at depth; stop forcing GLM levers). Doc-only. Co-authored-by: Copilot. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 19, 2026
… byte-identity lock (hybrid moat holds at the small end) (#1456) ## Summary Completes the qwen3.5 **hybrid context-scaling moat** family lock **trio** — **0.8B (this PR) + 2B (#1444, 3.03× deep) + 9B (#1449, 1.32× deep)** — a strong RULES §2 generality statement: the moat is a family property, proven from 0.8B up. **The generality question answered:** the moat **holds at the small end**. 0.8B is *not* too small for the graph-block to dominate — ORT-CUDA is graph-blocked *identically* to 2B/9B (same 25 Memcpy nodes), while native captures the whole hybrid graph with `fallbacks=0`. Adds `crates/onnx-genai-engine/tests/qwen35_0_8b_text_decode_lock.rs`. ## Model Hybrid decoder, 24 layers = **6 periodic full-attention `GroupQueryAttention`** layers (indices 3,7,11,15,19,23; `head_dim 256`, 8 q-heads / 2 kv-heads) + **18 gated linear-attention / short-conv** recurrent layers (`conv_state` + `recurrent_state`). Per-layer KV-vs-recurrent roles resolved structurally by the loader from the graph port inventory (RULES.md §2, never model-name-gated). ## Export (persisted, self-contained) `QWEN35_0_8B_TEXT_DIR=/home/justinchu/qwen35-0.8b-text-cuda` Foundry ships qwen3.5-0.8b as a multimodal **split** package (`embedding.onnx` + `text.onnx` + `vision.onnx`) the native single-model loader rejects (and that segfaults `onnxruntime-genai` on load). Composed a standalone `input_ids→logits` graph by **pruning the embedding subgraph to its `GatherBlockQuantized` text gather** (dropping the image-token `Equal`/`NonZero`/`ScatterND` merge — a no-op with no image tokens, and whose dynamic `NonZero` native shape-inference cannot bound) and fusing it into the text decoder via `inputs_embeds`. Same playbook as #1449 / #1442. Real weight copies in-dir (0.13 GB + 0.73 GB). No Mobius change needed. Reproducible via `qwen35-0.8b-text-cuda/export_qwen35_0_8b_text.py`. ## Result - **Native e2e (Y):** `"The capital of France is"` → `" Paris, and the capital of Germany is Berlin.\nThe capital of France is"`. Whole-graph CUDA capture (`captures=7 replays=433 fallbacks=0`). - **Byte-identity:** all 16 tokens **identical** to the independently-validated ORT-driven reference — the split-package `qwen35_0_8b_hybrid_text_decode_e2e` lock (ORT places standard attention on its EP, CPU-falls-back the hybrid ops) decodes the exact same stream. - **Golden lock PASS** (GPU1 24.2s; re-verified GPU2 24.2s). ## Moat (graph-block, confirmed at 0.8B) ORT-CUDA on this exact export inserts **25 Memcpy nodes** (`"25 Memcpy nodes are added ... including unable to run CUDA graph"`) and **cannot even initialize a session** — nodes are forced to the default CPU EP and session-init hard-fails. ORT has **no runnable whole-graph GPU path**; the recurrent `LinearAttention`/`CausalConvWithState` ops break CUDA-graph placement exactly as at the arch level (1037 CUDA / 56 CPU, 25 Memcpy). Native captures the whole hybrid graph and stays context-flat. ## A/B (H200, GPU1 pinned, `--steady`, native `--ep cuda`, medians) | context depth | native-CUDA (graph=1) | ORT-CUDA (best GPU path) | |---|---|---| | short (~64 tok) | **121.6 tok/s** (8.22 ms/tok, 5 runs) | **cannot run** — 25 Memcpy, CUDA graph disabled, session-init fails | | mid (~256–320) | **114.5 tok/s** (8.74 ms/tok) | cannot run | | deep (~1024) | **113.6 tok/s** (8.81 ms/tok, 5 runs) | cannot run | Native decode is **near context-flat** (~6–7% short→deep, then plateaus). ORT-CUDA's tok/s is undefined (no GPU-runnable path) — the moat at 0.8B is **categorical** (graph-blocked before it can decode a single token on GPU), the strongest form. *(An ORT-CPU baseline was not pursued: the composed export's minimal `inference_metadata.yaml` doesn't declare the rank-3 mrope `pipeline.positions` the ORT-genai loader requires — that spec only lives in the full split-package genai_config and is irrelevant to the native golden. A native-GPU-vs-ORT-CPU ratio is not the fair moat comparison anyway.)* ## Verdict **GO.** Do NOT self-merge — awaiting coordinator validation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1442 +/- ##
===========================================
+ Coverage 80.12% 82.19% +2.06%
===========================================
Files 364 12 -352
Lines 160709 5471 -155238
Branches 160709 5471 -155238
===========================================
- Hits 128775 4497 -124278
+ Misses 27283 775 -26508
+ Partials 4651 199 -4452
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes the last mile of #1438 (MAX_HEAD_DIM 256→512): proves the
head_dim=512full-attention arm runs end-to-end on a real dual-head-size model, not just unit parity.Gemma-3n / gemma4-e2b interleaves 28 sliding layers at head_dim=256 with 7 full-attention layers at head_dim=512 (KV layers 4, 9, 14, 19, 24, 29, 34). This adds a native-CUDA greedy-decode coherence golden lock that exercises all 35 layers every step, including the 512 layers on the fused split-K decode kernel raised in #1438.
crates/onnx-genai-engine/tests/gemma4_e2b_head_dim_512_decode_lock.rs— the lock.crates/onnx-genai-engine/tests/common/decode_lock.rs—assert_native_matches_golden_eagerhelper.Result
"Hello"→"! How can I help you today?"(+<end_of_turn>/<eos>).--test-threads=1, ~33 s).Export (persisted, self-contained)
GEMMA4_E2B_512_DIR=/home/justinchu/gemma4-e2b-it-text-cudaEvery on-disk gemma4 export was broken (dangling symlink target / speculative draft / multimodal-mandatory audio+vision). I composed a standalone text-only single graph (
input_ids+attention_mask+past KV →logits+present KV) by fusing the official pipeline'sembedding(producesinputs_embeds+routedper_layer_inputs) into thedecoder, bakingimage_features/audio_featuresto empty constants, with real weight copies in-dir and a<bos>-prepending tokenizer patch (Gemma degenerates without BOS). No Mobius change was needed.RULES.md §2 (general head size)
Per-layer head sizes are read structurally by the engine KV bridge from each
present.N.keyshape (kv_bridge.rs::layer_configs_from_key_outputs) — never from a model name or a fixed value. No value-keyed 256/512 branch; a future 3-distinct-head-size model needs only a new export + golden.Capture caveat (documented; not the 512 head)
The lock runs with CUDA-graph capture OFF. The composed graph's merged present-KV sequence axis is an opaque symbol the prefill workspace planner can't yet upper-bound for capture — and it trips on a head_dim=256 sliding layer (
present.13.key), not a 512 layer. Eager runs the identical fused kernels, so the 512 fast-path coherence is fully locked. Root cause + minimal-fix directions are in.squad/decisions/inbox/deckard-gemma4-512-e2e.md(cleanest: a non-merged native Gemma-3n text export).Do not self-merge — coordinator validation requested.
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com