Repository navigation
Redesign inference metadata as a generic control-flow IR - #828
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #828 +/- ##
==========================================
- Coverage 81.54% 80.54% -1.00%
==========================================
Files 384 394 +10
Lines 180801 184785 +3984
Branches 180801 184785 +3984
==========================================
+ Hits 147427 148839 +1412
- Misses 28431 30820 +2389
- Partials 4943 5126 +183
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
|
Branch join blocker fixed in Producer contract:
Updated paths: Validation: metadata 20+49+2, engine 362 passed/1 ignored, workflow policy E2E 6, ORT admission 16. PR remains draft for Mobius generated-package validation. |
|
Pushed preprocessing-adapter SSA support in Key contract:
Paths: Validation: metadata 20 unit + 50 fixtures + 2 schema; preprocess 54; ORT admission 16; engine lib 362 passed/1 ignored; workflow policy E2E 7; formatting/schema/diff checks pass. Read-only review found no significant issues. PR remains draft pending cross-repo generated package validation. |
|
Additional state soundness fix pushed in |
|
Pushed generic loop induction SSA in Contract: kind: loop
iteration:
value: loop.i
contract: { dtype: int64, rank: 1, shape: [batch] }The value is zero-based and materialized before each body execution; rank-0 scalar and explicit rank-1 broadcast contracts are supported. It is visible in the body/condition, does not escape the loop, and nested loops must use distinct lexical names. Reverse/remaining indices stay ordinary ONNX-derived values. Validation rejects wrong dtype/rank/shape, shadowing, and post-loop references. Runtime E2E now binds Tests: metadata 20 unit + 51 fixtures + 2 schema; workflow E2E 8; engine lib 362 passed/1 ignored; formatting/schema/diff checks pass. Read-only review found no significant issues. |
|
Cross-repo Mobius workflow execution update (
Validation:
One producer-side blocker remains for normal decoders: full setup logits |
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
Legacy-removal milestone pushed in
Validation: metadata 24/24; workflow policy E2E 11/11; engine 308 passed (1 ignored); server 146 passed (2 ignored); ORT 116/116; CLI 105/105; all affected targets checked. Full workspace check remains blocked by the pre-existing missing |
|
Pushed producer/runtime contract unblocker in
Validation:
Known unrelated baseline failures remain in removed legacy fixture tests: CLI image REPL tests reference missing |
|
Pushed the approved grammar/adaptive-K contracts in Producer targets for #478:
Runtime now executes grammar clone/lookahead/commit and telemetry generically. The conformance workflow combines speculative verification, grammar-valid-prefix truncation, committed semantic grammar state, forced-token ONNX sampling, and loop-carried advisory adaptive K. Exact producer documentation: Validation: metadata 20+4+2, workflow policy E2E 14, engine 308/1 ignored, ORT 116, server lib 146/2 ignored, config, formatting, schema sync, and affected checks passed. Read-only review found no significant issues. PR remains draft for cross-repo generated-package validation. |
|
World-model v1 acceptance fixture is pushed in Coverage:
The checkpoint type records state contracts and rejects incompatible restores. Review identified a possible partial mutation if cloning failed during restore; restoration now stages every clone before touching live session state. Validation: workflow policy E2E 15 passed; engine 308 passed/1 ignored; server library 146 passed/2 ignored; formatting and affected checks passed. The PR remains draft only for the separate acceptance requirement that every existing supported family be demonstrated with its real producer graph/E2E; the synthetic world-model requirement itself is complete. |
|
Execution-island milestone pushed in e0bc7f3.
Validation: metadata suite 28 passed; workflow E2E 19 passed; engine lib 309 passed/1 ignored; ORT lib passed; config passed; CUDA feature check passed; formatting/schema sync passed. Read-only review finding on sanitized SSA-name collisions was fixed and covered by E2E. |
e0bc7f3 to
b2157a2
Compare
|
Trailer-only amend: execution-island milestone head is now |
|
Mobius producer update One exact runtime blocker remains for the required decoder policy island: Required generic lowering: resolve application overrides before island partitioning, then test the resolved implementation's device, ONNX purity, ports, and effects. The package default and a selected pure same-device ONNX replacement should remain fusible; a stateful/host replacement should delimit the island. This preserves the versioned override ABI without sacrificing optimizer visibility or CUDA Graph capture. Please add conformance coverage for both the no-override and pure-ONNX-override cases. |
eadb31f to
8bacf8c
Compare
|
Performance conformance update (
The PR remains draft. Real producer-package/KV/per-row serving benchmarks and release ORT kernel/provider profiles remain readiness gates. |
|
Producer review against current
Also, workflow components/state currently have no explicit generic KV grouping/past-present alias/sequence-axis semantic fields beyond tensor names and shape recurrence. Please identify the intended existing contract or add those fields; Mobius will keep ordinary ONNX ports artifact-inferred and emit only this non-inferable KV service metadata. |
|
Producer blocker contract is finalized and pushed at Exact surfaces:
Migration rules:
Conformance now includes B=2 speculative per-row acceptance/ragged emit, B=1 stable ragged namespace, mixed ragged+dense appends, active-row carry, KV logical-length validation, zero-trip, and valid/missing package artifacts. Validation: engine 313 passed / 1 ignored; workflow E2E 19 passed; metadata/schema/CLI/config suites all passed; formatting and diff checks passed. |
|
Pushed
Validation: engine suite passed (313 passed, 1 ignored); workflow policy E2E 19 passed; ORT suite passed (116 unit plus integration tests); metadata/config/schema/validator passed; CUDA feature compilation passed; formatting/diff checks passed. CPU native-matched median ratios: decoder 0.973, min-p 0.978. H200 benchmark could not run because the linked downloaded ORT exposes CPU only and lacks |
5bd4783 to
76a37d8
Compare
|
Trailer-order-only amend: the milestone commit is now |
|
North Star follow-up pushed:
Validation: metadata 32/32 passed; engine 312 passed, 1 ignored plus all engine integration targets (workflow policy 19/19). Workspace-wide check remains blocked by the pre-existing uninitialized H200 native-matched benchmark (ORT 1.27 CUDA 13, identical linked composite/session I/O/capture): decoder workflow/native 362.85/371.06 step/s = 0.978; min-p 205.57/207.01 = 0.993. Both captured once and replayed 1003 times with no fallback. Decoder workflow TTFT remains 20.60 ms vs 2.78 ms native due planner/session startup; steady-state graph execution meets the preliminary bar. Request upload/final materialization are included in aggregate transfer diagnostics; linked internal component boundaries do not round-trip through host. |
|
Additional atomic cleanup pushed as |
|
Explicit conformance status: implemented and executed, not schema-only.
Validation: metadata all passed; ORT 116 unit plus integration targets passed; engine 310 passed/1 ignored plus all integration targets; workflow policy 19/19; native-backend compilation passed. PR remains draft for checked-in Mobius fixtures and the remaining explicit native speculative path/KV top-level migration. |
|
Further legacy migration pushed as |
|
Mobius checked-in producer fixtures are now available at They are reproducibly generated by Please point unconditional ONNX GenAI fixture CI at the commit/path above. |
|
KV legacy-path migration pushed in
Validation: metadata all-targets, genai-config all-targets, engine all-targets (294 passed, 1 ignored plus integration targets; workflow policy 19/19), ORT all-targets (including 116 unit tests), CLI/server checks, native-backend check, formatting, and diff check. Independent serving/KV review found no high-confidence defects. PR remains draft only for checked-in Mobius fixture ingestion. |
Replace binary graph fixtures with reviewable ONNX TextFormat across workflow, package, ORT, and native-runtime tests. Keep external QMoE weights mmap-capable, update all artifact references, and add a guard against checked-in binary ONNX graphs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Add durable audio, video, image, and protein outputs with upstream/runtime parity and measured performance where available. Preserve honest gaps for runs that did not retain runtime timing or a separate runtime render. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Add nineteen validated pipeline.workflow examples covering text, multimodal, speech, diffusion, scientific, recurrent-state, adapter, speculative, cache, and operator ABI cases. Document the config-only proof level and graph-visible versus runtime-private attention distinctions, and validate every catalogue YAML without requiring graph artifacts. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Reject sequence_axis on replace-updated state and remove stale binary-ONNX fixture exemptions now that all checked-in graphs use TextFormat, including external-weight QMoE fixtures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Add a validated Qwen3.5 decoder configuration that combines attention KV with linear-attention and causal-convolution replacement state. Document atomic rollback, snapshot-or-replay handling for partially accepted proposals, and persistent draft-model requirements. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Put onnx-genai-comfyui-config back into every offline build/clippy lane and the dependency-ordered crates.io release list so restoring the importer also restores its supported distribution surface. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Use schema_version as the sole workflow syntax version and read ONNX opset imports from each component artifact. Keep only semantic port roles and state aliases in metadata, and demonstrate independently shaped K/V and per-layer KV geometry in the Gemma 4 catalogue example. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Document that state role and layer provide semantic pairing and ordering while each ONNX port remains authoritative for its own geometry, including different per-layer and K-versus-V head counts. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…composite-metadata
Apply the repository Rust formatter after merging main so required quality and fast checks evaluate the combined tree cleanly. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
## Summary Produce the canonical ONNX GenAI `pipeline.workflow` ABI across Mobius exports and validate it against [onnx-genai#828](justinchuby/onnx-genai#828). - remove serialized `model.io`, legacy phase/strategy duplication, and runtime-private scheduling/allocation policy; - emit typed SSA workflows for decoder generation, static cache, speculative decoding, adapters, VLM, diffusion/image edit, video, speech, codec/TTS, masked diffusion, ESM-2, and ProtBert; - describe graph-visible state semantics including append, indexed-scatter, and fixed-size recurrent replacement; - emit canonical encoder embedding profiles and application-sourced opaque auxiliary inputs; - assign deterministic symbols to anonymous dynamic ONNX dimensions instead of invalid YAML `null`; - stop emitting duplicated workflow `ir_version` and `onnx_opsets`, leaving schema and ONNX artifacts authoritative; - pin cross-repository CI to ONNX GenAI commit `509cd4e9c4471f4cbc59fe44b47168f0ae128fe3`; - merge current `main` into the branch without rebasing. ## Validation - 164 metadata producer tests passed, 4 skipped; - all 13 generated package families match committed fixtures and pass the authoritative `validate_metadata` parser/semantic validator; - 11 ONNX GenAI engine conformance tests passed using real generated Mobius packages; - changed files pass lintrunner formatting. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Signed-off-by: justinchuby <justinchuby@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com> Signed-off-by: Justin Chu <justinchu@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Justin Chu <justinchu@users.noreply.github.com>
…n main) (#1686) `Rust quality` is red on `main` and has been since #828: `cargo fmt --all -- --check` reports five hunks in `crates/onnx-runtime-session/src/executor/geometry.rs` and `.../tests.rs`. - CI job `96899749480` on `0f40538b2` fails with `Diff in .../executor/geometry.rs:582`. - Reproduced locally on `4c15a64b3`; both files are byte-identical to `origin/main`, so this is not branch-local drift. This PR is **`cargo fmt --all` output and nothing else**. Verified locally on `4c15a64b3`: | gate | result | |---|---| | `cargo fmt --all -- --check` | clean | | `cargo check --locked -p onnx-runtime-session` | ok | | `cargo test --locked -p onnx-runtime-session` | ok, 0 failed | Split out of my MatMul work rather than folded into it, so the unblock is reviewable on its own and lands for everyone. Refs #1600. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…oncat it hid #1685 reported an intermittent 8-row hole of exact `0.0` in SDPA output under `--features mlas`, plus a SIGSEGV and `identity specialization diverged`. The defect is already fixed on main — incidentally, by PR #828's unrelated IR redesign, which added `wait_for_workers` to the MLAS work-stealing pool — and nothing guarded it. So this closes the coverage hole rather than the kernel. Root cause, now documented on `wait_for_workers`: `wait_for_completion` returns when the last *block* has run, but the worker that ran it is still inside `run_job` holding a by-value copy of the old `Job`. Pre-#828 the dispatcher released `dispatch_lock` there, so the next dispatch republished the loop bounds while that straggler was live; it then decremented the new job's `remaining` without running the new closure (a partition never executes, and a `beta = 0` SGEMM leaves rows of `C` unwritten) and invoked the *old* closure with an index from the *new* range (the SIGSEGV). The race needs two dispatches to overlap, so it is ~3% per process, not per call: 6000 `sdpa_f32` invocations and 3000 nested rayon x sgemm iterations reproduced nothing. `crates/mlas-sys/tests/concurrent_dispatch.rs` instead falsifies the fix — 4 dispatchers on an 8-thread pool. It passes in 0.10s on main and, with `wait_for_workers` removed, one test SIGSEGVs and the other reports the resulting deadlock through a watchdog. Coverage closed: - `AB_COVERED` gains `AttentionTranspose`, whose `PLAN` entry already claimed `Graduation::Partial` while the graduation rule reads `AB_COVERED`. - `sdpa_f32_native` / `sdpa_f32_mlas` name the two routes, since `sdpa_f32` short-circuits to MLAS and cannot be the native half of an A/B. - `native_vs_mlas_differential` gains 8 SDPA shapes x causal, a NaN-prefill fail-loud check that decodes a hole to (tile, row, column), a convex- combination oracle that still holds in a default MLAS-free build, and a concurrent-sessions test. - `identity_hook_specialization_matches_the_general_epilogue` prefills NaN. Falsified: dropping 8 rows/tile (6144 of 98304 elements, #1685's exact signature) leaves the old zero-prefill version *passing*. Two real defects the new coverage found: 1. `concat_cache` looped head-dim outside sequence, striding 512B per store through a contiguous `[b][h][s][d]` buffer and traversing it `dim` times — ~20ms of a ~28ms decode node. It is now two `copy_from_slice` calls per plane with the same fan-out threshold the sibling transforms use. `llama_decode_past1023` 27.2x -> 13.8x ORT, `chunk8` 24.8x -> 13.1x, `chunk32` 27.7x -> 18.5x, parity preserved. 2. No golden exercised the causal offset: every self-attention case has `past == 0` and the one past-KV case has `q_seq == 1`, so `causal = unidirectional && q_seq > 1` is false. Hard-coding `past_seq = 0` leaves all 12 pre-existing goldens green; the new `past_kv_chunked_prefill_causal` is the only one that fails (2.22 abs). `gen_mha.py` gains causal, decode and past-KV rows, and now *rejects* `unidirectional` with `q_seq != kv_seq` and no past-KV: a NumPy oracle swept over every offset in `0..=kv_seq` matches ORT at none of them, so that cell has no defined answer and its 24/384 parity failures were not a kernel bug. Ledger section 45 records the audit, the falsifier method, the 18-cell x 4-thread production-default matrix (432/432 parity PASS, A/A null control per cell), and the finding it exposed: `native_p50` is flat across 1..16 threads on every cell because `sdpa_f32_simd` — the route a default build actually takes — has no rayon fan-out at all, while the MLAS research route does. That is filed separately rather than folded in here. Gates: fmt; clippy default + `--features mlas` + aarch64 `-D warnings`; `check_cross_compile.sh`; full ep-cpu suite both feature configs; mlas-sys; Miri task_runtime/strided/provider/dtype; default cdylib 0 MLAS symbols by `nm`/`nm -D`/`strings`/`ldd` with an 842-symbol positive control. Closes #1685 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…oncat it hid #1685 reported an intermittent 8-row hole of exact `0.0` in SDPA output under `--features mlas`, plus a SIGSEGV and `identity specialization diverged`. The defect is already fixed on main — incidentally, by PR #828's unrelated IR redesign, which added `wait_for_workers` to the MLAS work-stealing pool — and nothing guarded it. So this closes the coverage hole rather than the kernel. Root cause, now documented on `wait_for_workers`: `wait_for_completion` returns when the last *block* has run, but the worker that ran it is still inside `run_job` holding a by-value copy of the old `Job`. Pre-#828 the dispatcher released `dispatch_lock` there, so the next dispatch republished the loop bounds while that straggler was live; it then decremented the new job's `remaining` without running the new closure (a partition never executes, and a `beta = 0` SGEMM leaves rows of `C` unwritten) and invoked the *old* closure with an index from the *new* range (the SIGSEGV). The race needs two dispatches to overlap, so it is ~3% per process, not per call: 6000 `sdpa_f32` invocations and 3000 nested rayon x sgemm iterations reproduced nothing. `crates/mlas-sys/tests/concurrent_dispatch.rs` instead falsifies the fix — 4 dispatchers on an 8-thread pool. It passes in 0.10s on main and, with `wait_for_workers` removed, one test SIGSEGVs and the other reports the resulting deadlock through a watchdog. Coverage closed: - `AB_COVERED` gains `AttentionTranspose`, whose `PLAN` entry already claimed `Graduation::Partial` while the graduation rule reads `AB_COVERED`. - `sdpa_f32_native` / `sdpa_f32_mlas` name the two routes, since `sdpa_f32` short-circuits to MLAS and cannot be the native half of an A/B. - `native_vs_mlas_differential` gains 8 SDPA shapes x causal, a NaN-prefill fail-loud check that decodes a hole to (tile, row, column), a convex- combination oracle that still holds in a default MLAS-free build, and a concurrent-sessions test. - `identity_hook_specialization_matches_the_general_epilogue` prefills NaN. Falsified: dropping 8 rows/tile (6144 of 98304 elements, #1685's exact signature) leaves the old zero-prefill version *passing*. Two real defects the new coverage found: 1. `concat_cache` looped head-dim outside sequence, striding 512B per store through a contiguous `[b][h][s][d]` buffer and traversing it `dim` times — ~20ms of a ~28ms decode node. It is now two `copy_from_slice` calls per plane with the same fan-out threshold the sibling transforms use. `llama_decode_past1023` 27.2x -> 13.8x ORT, `chunk8` 24.8x -> 13.1x, `chunk32` 27.7x -> 18.5x, parity preserved. 2. No golden exercised the causal offset: every self-attention case has `past == 0` and the one past-KV case has `q_seq == 1`, so `causal = unidirectional && q_seq > 1` is false. Hard-coding `past_seq = 0` leaves all 12 pre-existing goldens green; the new `past_kv_chunked_prefill_causal` is the only one that fails (2.22 abs). `gen_mha.py` gains causal, decode and past-KV rows, and now *rejects* `unidirectional` with `q_seq != kv_seq` and no past-KV: a NumPy oracle swept over every offset in `0..=kv_seq` matches ORT at none of them, so that cell has no defined answer and its 24/384 parity failures were not a kernel bug. Ledger section 45 records the audit, the falsifier method, the 18-cell x 4-thread production-default matrix (432/432 parity PASS, A/A null control per cell), and the finding it exposed: `native_p50` is flat across 1..16 threads on every cell because `sdpa_f32_simd` — the route a default build actually takes — has no rayon fan-out at all, while the MLAS research route does. That is filed separately rather than folded in here. Gates: fmt; clippy default + `--features mlas` + aarch64 `-D warnings`; `check_cross_compile.sh`; full ep-cpu suite both feature configs; mlas-sys; Miri task_runtime/strided/provider/dtype; default cdylib 0 MLAS symbols by `nm`/`nm -D`/`strings`/`ldd` with an 842-symbol positive control. Closes #1685 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…oncat it hid (#1714) Closes #1685. ## The defect is already fixed on `main` — and nothing was guarding it #1685 reports an intermittent SDPA failure under `--features mlas`: ~3% of runs leave an 8-row hole of exact `0.0` in the output, or die with a SIGSEGV, tripping `identity specialization diverged`. The fix landed incidentally in **`0f40538b2` (PR #828, "Redesign inference metadata as a generic control-flow IR")**, which added 12 lines to `work_stealing_pool.rs` as a side effect. The reported-bad SHA `8aed77a17` has **0** occurrences of `wait_for_workers`/`observed`; current `main` has 7. **Mechanism.** `wait_for_completion` returns when `remaining` hits zero — when the last *block* has run. The worker that ran it is still inside `run_job`, holding a by-value copy of the old `Job` and looping in `claim_iterations` against the *shared* counters. Pre-fix the dispatcher released `dispatch_lock` at that moment; MLAS fans out under rayon in `sdpa_f32_fast`, so the next dispatch republished the loop bounds and bumped the epoch while the straggler was live. It then did two things at once: 1. decremented the **new** job's `remaining` without running the new closure → `wait_for_completion` returned early → a partition **never executed** → a `beta = 0` SGEMM left rows of `C` unwritten (the 8-row hole: `8 · dv` contiguous zeros at a tile boundary — MLAS had split the 128-row `probs·V` GEMM 16 ways); 2. invoked the **old** closure with an index from the **new** range, writing through raw pointers past the end of the previous GEMM's `C` → the SIGSEGV. ## Reproduction failed; falsification worked | attempt | result | |---|---| | 80 runs of the exact test | 0 failures | | 40 runs, `sdpa` filter, `--test-threads 8` | 0 | | pool widths 4 / 8 / 16 / 32, 40 runs each | 0 | | 3000 iterations, nested rayon × `sgemm`, NaN-prefilled `C` | 0 | | 6000 `sdpa_f32` invocations at the issue's shape | 0 | The race needs two dispatches to **overlap**, and `dispatch_lock` serialises them — so the window is only the gap between the last block finishing and the straggler leaving `run_job`. That is ~3% **per process**, not per invocation. So I stopped trying to trigger it and tried to **break the fix instead**. `crates/mlas-sys/tests/concurrent_dispatch.rs` runs 4 dispatcher threads against one 8-thread pool. On `main` it passes in **0.10 s**. With only `wait_for_workers` removed: - `overlapping_dispatches_never_return_with_work_unexecuted` → **SIGSEGV** - `parallel_for_returns_only_after_every_worker_has_left_the_closure` → the straggler drives `remaining` below zero, `usize` wraps, and the pool deadlocks; a watchdog reports that instead of hanging CI (A third test passed *under* the falsifier and was deleted rather than shipped as false assurance.) ## The audit — the actual deliverable | Instrument | State before this PR | |---|---| | `backend_ab.rs` `AB_COVERED` | `AttentionTranspose` **absent**, while its `PLAN` entry claims `Graduation::Partial` — and the graduation rule *reads* `AB_COVERED` | | `tests/native_vs_mlas_differential.rs` | no attention row at all | | `benches/native_vs_mlas.rs` | no attention row at all | | `identity_hook_specialization_matches_the_general_epilogue` | compared two **zero-prefilled** buffers | | `scripts/ort_ab/gen_mha.py` | 7 cells, all bidirectional encoder shapes; `unidirectional` supported but never set; no `q_seq = 1`; no past-KV | | `mha_parity/cases.rs` | 12 goldens; the only past-KV case has `q_seq = 1` | The family with a known reference-route defect was the family with no same-binary A/B. **Production-route coverage was already sound** — `mha_ort_parity.rs`, `msft_attention_ort_parity.rs` and `qwen35_ort_parity.rs` all run ORT goldens in the default MLAS-free build. The hole was entirely in the research/reference route and in the shape grid. ### Zero-prefill cannot see a dropped write Falsified by skipping 8 rows of the `probs·V` GEMM — 6144 of 98304 elements, #1685's exact signature: - old zero-prefill assertion → **passes** (an unwritten element is indistinguishable from a legitimate `0.0`, and a hole landing identically in both compared runs cancels out) - new NaN prefill → `left 6144 of 98304 output elements unwritten; first at index 7680 = (tile 0, row 120, column 0)` ## Two real defects the new coverage found ### 1. `concat_cache` walked a row-major buffer column-major ```rust for d in 0..dim { // head dim OUTSIDE for j in 0..past.seq { // sequence INSIDE data[((b * heads + h) * total + j) * dim + d] = past.at(b, h, j, d); ``` `Bnsh` is contiguous `[b][h][s][d]`, so consecutive stores were `dim` floats — **512 B** at Llama's head size — apart: every store touched a fresh cache line and the tensor was traversed `dim` times. ~4.2 M near-certain misses per tensor, **~20 ms of a ~28 ms decode node**. The operator spent most of a decode step copying its own cache. The two sibling transforms document this exact rule in their doc comments; this was the one place that broke it, and no benchmark row supplied a past-KV cache, so nothing measured it. Now two `copy_from_slice` calls per `(b, h)` plane, fanned out on the same `MIN_PARALLEL_TRANSPOSE_ELEMENTS` threshold: | cell (t=8) | before | after | |---|---|---| | `llama_decode_past1023` | 27.2× | **13.8×** | | `llama_chunk8_past1016` | 24.8× | **13.1×** | | `llama_chunk32_past992` | 27.7× | **18.5×** | Parity preserved on all three. ### 2. Nothing checked the causal offset `past_seq` is load-bearing only when `q_seq > 1` **and** the cache is non-empty. Every self-attention golden has `past == 0`; the one past-KV golden has `q_seq == 1`, so `causal = unidirectional && q_seq > 1` is **false**. Hard-coding `past_seq = 0` leaves **all 12 pre-existing goldens passing**. The new `past_kv_chunked_prefill_causal` is the only case that fails (`max abs diff 2.22`). Our convention was already correct; it is now pinned. Regenerating with ORT 1.26.0 was byte-stable for the existing 12 cases (+19 lines only). ## An undefined benchmark cell is not a slow one Adding causal/decode rows produced 24/384 parity failures on `q_seq=8, kv_seq=1024, unidirectional=1`, no past-KV. That looked like a kernel bug and was not one: | reference | vs ORT | |---|---| | `unidirectional = 0`, no mask | **1.2e-7** ✅ | | causal, offset swept over **every** value in `0..=kv_seq` | matches at **none** | | causal at offset `past_seq`, with a real past-KV cache | **2.4e-7** ✅ | ORT's `unidirectional` is simply undefined when `q_seq != kv_seq` with no past input — neither runtime computes a defined answer, so no timing from that cell is meaningful. It is replaced by past-KV cells (`llama_decode_past1023`, `llama_chunk8_past1016`, `llama_chunk32_past992`), which is how the runtime actually emits chunked prefill, and `build_mha` now **raises** on the invalid combination. ## Complete results — production default build, 18 cells × 4 thread counts Default MLAS-free `bench_generic`: `nm`, `nm -D`, `strings`, `ldd` → **0** MLAS symbols, no `libstdc++`. **Positive control:** the same probes on a `--features mlas` build report **842** symbols / 105 strings / `libstdc++` linked, so the probe demonstrably sees MLAS when present. `MultiHeadAttention` executes natively at 99.97% of node time, 1 call, no ORT fallback. **432/432 trials parity PASS.** `native/ort` p50, lower is better; `(…)` = same-invocation **A/A null control**: | cell | t=1 | t=4 | t=8 | t=16 | native ms t=1 → t=16 | |---|---|---|---|---|---| | `bert_base_b8_s128` | 5.36 (5.31) | 13.01 (8.97) | 15.50 (14.46) | 19.93 (20.97) | 38.2 → 42.3 | | `bert_base_decode_kv1024` | 1.51 (1.54) | 0.78 (0.76) | **0.55** (0.46) | 0.62 (0.58) | 1.2 → 1.3 | | `bert_base_s128` | 5.71 (5.73) | 9.99 (8.56) | 11.63 (10.64) | 11.65 (12.29) | 4.8 → 7.8 | | `bert_base_s384` | 7.04 (7.14) | 22.89 (15.69) | 23.36 (23.78) | 28.18 (32.36) | 52.7 → 59.7 | | `bert_large_s128` | 5.65 (5.67) | 11.29 (8.89) | 12.75 (14.01) | 15.75 (15.92) | 6.3 → 11.0 | | `clip_l14_s257` | 5.79 (5.74) | 14.00 (13.25) | 24.49 (25.95) | 24.83 (24.52) | 25.0 → 38.1 | | `llama_chunk32_past992` | 4.11 (4.18) | 11.69 (12.31) | 17.93 (17.64) | 24.29 (23.75) | 42.4 → 42.9 | | `llama_chunk8_past1016` | 3.41 (3.59) | 8.10 (7.59) | 11.77 (11.99) | 17.16 (17.15) | 17.6 → 21.9 | | `llama_decode_b8_kv1024` | 1.50 (1.50) | 1.44 (1.45) | 1.55 (1.53) | 1.53 (1.55) | 79.9 → 51.4 | | `llama_decode_kv1024` | 1.93 (1.88) | 1.14 (1.46) | 1.33 (1.22) | 1.24 (1.68) | 13.1 → 8.8 | | `llama_decode_kv128` | 1.50 (1.49) | 0.69 (0.75) | **0.63** (0.65) | 0.72 (0.63) | 0.6 → 0.8 | | `llama_decode_kv4096` | 1.59 (1.59) | 1.32 (1.35) | 1.40 (1.39) | 1.40 (1.68) | 42.2 → 31.8 | | `llama_decode_past1023` | 3.48 (3.49) | 8.21 (8.00) | 11.00 (9.65) | 15.03 (13.60) | 10.5 → 13.8 | | `llama_prefill_s128_causal` | 3.31 (3.35) | 5.67 (6.86) | 6.74 (7.67) | 9.96 (8.04) | 15.5 → 22.5 | | `llama_prefill_s512_causal` | 3.75 (3.73) | 12.34 (12.49) | 19.34 (19.30) | 21.86 (21.60) | 237.8 → 232.5 | | `phi35_prefill_s256_causal` | 4.02 (3.99) | 11.11 (10.67) | 15.20 (15.04) | 18.18 (17.25) | 51.8 → 52.8 | | `vit_b16_s197` | 5.64 (5.66) | 11.60 (11.33) | 11.37 (11.54) | 12.09 (16.70) | 10.9 → 11.3 | | `whisper_cross_s1500` | 6.16 (6.16) | 20.42 (21.55) | 31.52 (31.68) | 40.29 (40.46) | 182.6 → 181.3 | Regressions are included, not filtered. The null arm tracks the native arm within a few percent on every cell, so these are the measurement and not the instrument. ### What the table actually says Read the last column. **`native ms` is flat from 1 → 16 threads on every cell.** The ratio degrades with thread count purely because ORT scales and we do not (`llama_decode_past1023`: ORT 3.05 ms → 1.06 ms across the same sweep). The code agrees: `sdpa_f32_simd` — the route a **default** build takes on x86 and aarch64 — is a plain `for b { for n { … } }` with **no rayon fan-out at all**. `sdpa_f32_fast`, the MLAS *research* route, does use `par_chunks_mut`. The shipped route is the serial one. At `t = 1` the grid is 1.5–7.0×, which is a per-core efficiency gap. Everything above that is unclaimed parallelism worth roughly the thread count. The cells that already win (`bert_base_decode_kv1024` 0.55×, `llama_decode_kv128` 0.63×) are the ones small enough that one core suffices. **This is filed separately rather than folded into a coverage PR** — it is a large change and it touches pool ownership, which is Sebastian's lane. Now tracked as **#1718**. ## Gates | gate | result | |---|---| | `cargo fmt --all --check` | clean across all 11 files in this diff (main carries 3 pre-existing unformatted files in `onnx-genai-engine` / `onnx-genai-metadata`, untouched here) | | `clippy --locked -p onnx-runtime-ep-cpu -p mlas-sys --all-targets` | clean | | `clippy --no-default-features --features mlas --all-targets` | clean | | `clippy --target aarch64-unknown-linux-gnu --all-targets -- -D warnings` | clean (caught a real `needless_return` in my new `sdpa_f32_native` cfg ladder) | | `scripts/check_cross_compile.sh` | ✓ full offline set | | `cargo test -p onnx-runtime-ep-cpu` (default) | 1604 passed, 0 failed | | `cargo test -p onnx-runtime-ep-cpu --features mlas` | 1552 passed, 0 failed | | `cargo test -p mlas-sys` | 41 passed, 0 failed | | Miri `task_runtime` / `strided` / `provider` / `dtype` | 29 / 8 / 16 / 12 passed, no UB | | `default_artifacts_are_mlas_free` | 9 passed | | default cdylib MLAS symbols | `nm` 0, `nm -D` 0, `strings` 0, `ldd` no `libstdc++` — positive control: the same probes on a `--features mlas` build see **842** symbols, so the probe demonstrably detects MLAS when present | Note: `onnx-runtime-ep-cuda` has a **pre-existing** `approximate value of PI` clippy error (a CUDA C++ kernel string in `kernels/window.rs`). This diff touches 0 files in that crate. ## Review Reviewed by Opus (`claude-opus-4.8`): **no blocking issues**, with hand-verified index algebra for the `concat_cache` rewrite and confirmation that the new tests are non-vacuous. Two non-blocking robustness notes were raised and both are now addressed in `9ae8f7b26`: - the `peak > 1` overlap guard could misfire on a single-vCPU runner where workers may serialise; it is now gated on `available_parallelism() > 1`, so it still fails loudly wherever overlap is possible; - `concurrent_sdpa_sessions_lose_no_work` had no watchdog, so a reintroduced pool deadlock would have hung the suite rather than failing it. It now runs on a 120 s watchdog thread with the same diagnosis message as `concurrent_dispatch`. Neither touches production code. All gates re-run green after both changes and after the rebase onto `d3688e7e0`. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…d a knob that no longer exists (#1822) Follow-up to #1173, correcting two defects I shipped in it and repairing the rule they undermined. Docs, one ledger string, one new test, one new script. No production kernel or routing change. ## 1. The ledger named a route gate that had already been deleted `PLAN[MatMulF32].shape_gate` said the native `SimdX86` route "gates M=1 on `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, #1116)". #1183 shipped that GEMV on by default and removed the probe. `git merge-base --is-ancestor 5417d04 bdb4599` confirms it landed **before** #1173 merged — so the ledger was wrong the day it landed. Today `sgemm_simd` calls `sgemm_simd_variant(a, b, c, m, k, n, true)` unconditionally and `use_m1_gemv` is a plain parameter that only the in-process A/B harness passes as `false`. No environment variable reaches that route. `docs/performance/CPU_MATMUL_ASSIGNMENT.md:559` already recorded the correct fact ("It is measured now, and the route is the default. There is no env probe on the dispatch any more"). Two files in the same directory disagreed and nothing compared them. **Now guarded.** `ledger_prose_only_names_environment_variables_that_still_exist` requires every `NXRT_*` / `ONNX_GENAI_*` token in the ledger's prose to still exist as a string literal in the crate's sources. It cannot check that the description is *right*, only that the knob is *real* — which is the half that goes stale silently. Mutation-verified, not just observed green: ``` matmul_f32: ledger prose names environment variable `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`, but no source file in this crate contains the literal "ONNX_GENAI_CPU_MM_SIMD_M1_GEMV". ``` ## 2. The doc published a toggle A/B that could not have been run #1173 carried a table captioned **"same binary, same session, toggle the only difference"**, reporting `decode 1×2048×2048` at 0.146 with `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` off against 0.337 with it on, and called turning it on "the obvious next slice". Nothing reads that variable. Setting it measures the same route twice; it cannot produce two different columns. The table is withdrawn and the retraction kept in the text rather than quietly deleted. This is the failure mode the document's own graduation rule warns about — **an arm that was not on the route it was labelled with** — committed by the document that wrote the rule. It survived review because a plausible number in a well-formed table is not self-evidently unmeasured. Readers are pointed at `bench_f32_gemm_ab`, which holds the route as a function parameter and carries the M≥2 rows as a built-in control. ## 3. The gap table is re-measured and the ≥5% rule is repaired The old table was one unguarded invocation per row at an unstated width, taken before the decode-placement corrections (#1729, #1794, #1811) — i.e. when the decode pool put 16 workers on 8 physical cores. New harness: `scripts/bench_native_vs_mlas_width.py`. Arms interleaved rep by rep so host drift lands on both equally; per-rep `os.wait4` CPU-efficiency guard adapted from #1809; six reps per arm; two widths. Raw verdicts, spreads and discards are all reported rather than summarised away. **Three findings, all about method rather than kernels.** | | narrow (6 cores, 1 L3) | wide (32 logical CPUs) | |---|---|---| | `matmul_f32 16×512×512` | 1.581, spread 41% | 0.866, spread 134% | | `matmul_f32 decode 1×2048×2048` | 1.117, spread 21% | 0.934, spread 13% | - **Two cases change verdict on width alone.** Same binary, same half-hour, only the CPU mask differs. `x86_sgemm` parallelises over column strips and MLAS declines to parallelise some shapes, so interleaving the two *routes* inside one process does not protect the ratio — it changes both at once. - **`16×512×512` disagrees with itself on both arms**, alternating `keep-mlas` / `native-graduates` from a byte-identical binary. **One more run of the old table could have graduated a route on this row.** - **The narrow arm is more trustworthy despite having fewer cores** — spreads 4–42% against 5–134%, and it lost no reps to the guard. Isolation beat parallelism. **Softmax now decomposes cleanly**, because no vendored MLAS kernel has changed since #1173 (the only `mlas-sys` edits are the additive straggler handshake in `work_stealing_pool.rs`, #828/#1714, which adds waiting). At matched width the MLAS control arm is stationary to within 4% while native improved **1.24–1.27×** — matching #1416's claim for the row kernel. The f32 GEMM rows get no such attribution and now say so explicitly: their control moved **2.0× the wrong way**, so only the current ratio at a stated width is defensible. **The rule gains what it lacked**: spread must be smaller than the claimed win; reps that did not get the CPU are discarded rather than averaged; a verdict is valid only at a stated width. Under it, `decode 1×2048×2048` — the first f32 GEMM case to show a real native win — **still does not graduate**: it costs more CPU (cpu_ratio 0.875), does not hold at 32 threads, and its 21% spread exceeds its 12% win. ## The width claim is verified, not asserted #1815 landed while this was in progress and observed the neighbouring `bench_generic` harness spawning its ORT arm *outside* the affinity confinement it applied to the native arm. That hazard applies to any `taskset` claim, including mine, so I checked it instead of trusting it — sampling `Cpus_allowed_list` from `/proc/<pid>/task/*/status` 40× across a live narrow-arm run: ``` '16,20,22,26,28,30': 478 observations native_vs_mlas- 273, mlas-sys-ws-0..4 39 each, nxrt-task-0..4 2 each '0-31': 1 (the taskset process itself, before exec) ``` Both routes confined identically; no thread escaped. The rule now requires this check. ## Validation - `dispatch_ledger` **17/17**, including the new falsifier, after merging latest `main`. - `default_artifacts_are_mlas_free` **9/9** — the no-MLAS-in-defaults invariant is untouched. - `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` clean; `cargo fmt --check` clean. - Normal merge of `origin/main` (`aee2b9d11`), no rebase, no conflicts. ## Limitations - The narrow arm is six cores on one L3 of one x86-64 host. Nothing here transfers to aarch64 or to a two-socket box. - The `activations erf 1 Mi` row shows native 13.5% slower at matched width. The nearest scatter figure is the wide arm's 8% spread, but that is a spread of *ratios* against a move in a *native time*, so the two are not strictly commensurable. Its MLAS control also moved 11%. **Flagged for pinned re-measurement, not reported as a regression.** - The wide arm was taken with ~4–5 cores of unrelated load present. That is stated in the doc rather than hidden, and it is why its spreads are wider; the guard reports which reps were discarded instead of pretending the host was quiet. - No production behaviour changes here, so there is no performance claim to make about the shipped artifact. Refs #1173, #1183, #1809, #1815, #1416. ## Independent review, and what it changed An independent adversarial review of the full diff returned **no blockers** — it confirmed the ancestry argument behind the retraction, the stationary-control premise for the softmax attribution, and that the headline case is correctly *refused* by the rule (21% spread against a 12% win). It also found seven real defects, all now fixed in `f0323f9ed`. The one that mattered most was in the new test. It only proved the variable name appeared *somewhere* in the crate, so a variable whose read site had been deleted but whose name survived in an `EnvVarGuard::set(...)` line would still have passed — which is the precise shape of the defect this PR exists to correct. The test now requires the matching line to be an `env::var(` / `env::var_os(` read or an `_ENV: &str =` binding. Verified by mutation in **both** directions: | mutation | before | after | |---|---|---| | reinsert retired `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` into ledger prose | fails ✅ | fails ✅ | | retire the two real `NXRT_CPU_GEMM_BACKEND` reads, leaving the literal only in test guards | **passes ❌** | fails ✅ | The remaining six were prose defects in the doc: a stated spread range that contradicted its own table's 82% row, "within 4%" against a table reading −4.2%, a narrow-arm ratio fused with a wide-arm attribution, a spread quoted as 7.5% that was 8% *and* compared against an incommensurable quantity, the CPU-efficiency guard oversold as "what makes this table measurable at all" (in-process interleaving is what protects the ratio; the guard catches only *differential* descheduling), and a one-directional provenance argument standing in for the direct control measurement that actually carries the softmax attribution. **Two further defects I found myself while checking the tables against each other**, neither raised by the review: - The `ratio` column is a median of per-rep ratios while the `ns/unit` columns are medians of times. Medians do not distribute over division, so every row looked internally inconsistent to anyone who tried to divide it out (`0.0684 / 0.0617 = 1.109` against a stated `1.117`). Now documented, along with why the per-rep form is the correct one to quote: it pairs each MLAS invocation with the native invocation it was interleaved against, which is the entire point of interleaving. The then→now figures are relabelled as quotients of medians. - "wider than nine of the twelve wide-arm rows" was eleven of twelve. ## Adopting #1814 `aee2b9d11` (#1814) landed on `main` while this was in review, and it closes the exact hole the review found in the guard this document recommends. A differential CPU-efficiency check cannot see contention that lands evenly on both arms; #1814's confined-set meter reads busy jiffies on the process's own `Cpus_allowed_list` and subtracts the process's own CPU, so foreign load shows up directly. The rule now points at it, and the tables here are explicitly marked as predating it and guarded by the weaker method. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the row's siblings give either a full path from the repository root or a bare file name. This one is neither, so a reader following it finds nothing and cannot tell whether the test was moved, renamed, or never existed. Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing it anyway because the citation checker this branch runs flags it, and claiming "every citation in both documents resolves" while knowingly leaving one that does not would make the claim worth nothing. The audit that found it was prompted by the same class of defect found in a PR body upstream of this branch: published prose that nothing type-checks. Docs only; no code, schema, or fixture changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the row's siblings give either a full path from the repository root or a bare file name. This one is neither, so a reader following it finds nothing and cannot tell whether the test was moved, renamed, or never existed. Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing it anyway because the citation checker this branch runs flags it, and claiming "every citation in both documents resolves" while knowingly leaving one that does not would make the claim worth nothing. The audit that found it was prompted by the same class of defect found in a PR body upstream of this branch: published prose that nothing type-checks. Docs only; no code, schema, or fixture changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the row's siblings give either a full path from the repository root or a bare file name. This one is neither, so a reader following it finds nothing and cannot tell whether the test was moved, renamed, or never existed. Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing it anyway because the citation checker this branch runs flags it, and claiming "every citation in both documents resolves" while knowingly leaving one that does not would make the claim worth nothing. The audit that found it was prompted by the same class of defect found in a PR body upstream of this branch: published prose that nothing type-checks. Docs only; no code, schema, or fixture changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the row's siblings give either a full path from the repository root or a bare file name. This one is neither, so a reader following it finds nothing and cannot tell whether the test was moved, renamed, or never existed. Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing it anyway because the citation checker this branch runs flags it, and claiming "every citation in both documents resolves" while knowingly leaving one that does not would make the claim worth nothing. The audit that found it was prompted by the same class of defect found in a PR body upstream of this branch: published prose that nothing type-checks. Docs only; no code, schema, or fixture changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the row's siblings give either a full path from the repository root or a bare file name. This one is neither, so a reader following it finds nothing and cannot tell whether the test was moved, renamed, or never existed. Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing it anyway because the citation checker this branch runs flags it, and claiming "every citation in both documents resolves" while knowingly leaving one that does not would make the claim worth nothing. The audit that found it was prompted by the same class of defect found in a PR body upstream of this branch: published prose that nothing type-checks. Docs only; no code, schema, or fixture changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the row's siblings give either a full path from the repository root or a bare file name. This one is neither, so a reader following it finds nothing and cannot tell whether the test was moved, renamed, or never existed. Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing it anyway because the citation checker this branch runs flags it, and claiming "every citation in both documents resolves" while knowingly leaving one that does not would make the claim worth nothing. The audit that found it was prompted by the same class of defect found in a PR body upstream of this branch: published prose that nothing type-checks. Docs only; no code, schema, or fixture changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Redesigns the portable inference metadata contract around one typed structural
workflow IR (
sequence/invoke/loop/branch/emit), with no phases,strategies, or model-family dispatch.
The governing rule is an ownership split. Metadata declares structural and
semantic facts a runtime cannot infer. The runtime owns deployment policy —
memory budget, placement, EP, KV allocation and paging, compaction algorithm,
cache keys, scheduling, and the request/sequence table.
docs/genai/INFERENCE_METADATA_DECISIONS.mdis now the normative specification(goals/non-goals, terminology, ownership layers, schema model, workflow
semantics, batching, LoRA, multimodal, cache dependencies, state, sessions,
speculative, sharding, task profiles, legacy import, migration, invariants, and
conformance).
Removed, fail-closed
slot_ids,request_epochs,emit.row_ids,RuntimeInputRole::{RowIds,RequestEpochs}TensorContract.batch_layout(shared/request_aligned{axis}/token_packed{offsets,owner}/runtime_sequence_state) plus a runtime-mintedRowSelectiongatherKvQuantPolicyinonnx-genai-kvshared_bufferallocator flagaliasing: permitted | required | forbidden— the real graph ABI constraint it stood in for; defaults toforbiddenKvCacheSpec,QuantizationAxis,ForkPrecisionPolicyRetired fields are rejected by
deny_unknown_fields, so a stale document failsrather than being silently reinterpreted.
Added
RowScopedState::compact(permutation)/release(row). LoRA, native grammar, and vision state survive batchcompaction. An invariant, not a negotiated capability.
cache_dependencies()walks workflow SSA,state writes, and component dataflow, and always includes adapter artifacts,
externally-suppliable encoder results, and generation-affecting profiles. A
dependency cannot be omitted by silence.
bitwise | distribution_preserving | semanticdecidewhether the runtime may substitute an equivalent implementation unasked. An
absent contract counts as
semantic.speculation_safetyare independent axes. Thevalidator reads only
speculation_safety, and holds every component in theenclosing loop body to the bound — not just the proposer and target.
Runtime-owned state may not be published without one, and publication is
detected on the emitted value, so an emit cannot export private state under
an alias.
semantic_identity()) binding disposableplans and checkpoints to metadata semantics. Identity, not integrity or trust.
Normalization is conservative: it never merges two documents it cannot prove
equivalent from syntax alone.
a request-sourced typed workflow input with declared constraints. Anything else
fails loudly.
language dialect and version, task profiles with required-vs-ignorable
semantics, and legal TP/PP/EP sharding facts.
import_genai_config,--allow-lossy).No reverse synthesizer, because one would silently approximate.
Latest integration
maininto the feature branch (no rebase), preserving theworkflow-derived decoder ABI and current native/MTP changes.
the same canonical
pipeline.workflowIR; the runtime never dispatches onComfyUI
class_type; the crate remains in all offline build/clippy lanes and the crates.io publish order.update: { kind: replace }. Linear-attentionaccumulators and causal-convolution history share
kind: recurrentbut remainseparate groups because their shapes, ports, rollback, and checkpoint
boundaries differ. Replacement state is invariant and cannot declare a
sequence_axis; append/indexed-scatter state must declare one.*.onnx.textproto. Path-based loading preserves external-data descriptors, soQMoE mmap/offload coverage remains real without a binary graph. A regression
test rejects any tracked
*.onnxgraph.covering Gemma 4, Cosmos3 Edge/world rollout, Qwen3.5 VLM and hybrid speculative decoder, Whisper, Wav2Vec2,
PersonaPlex, Stable Diffusion, Qwen Image Edit, CogVideoX, LoRA, speculative
decoding, ESM-2, ProtBert, WeatherNext, local/sliding attention, linear
attention, causal convolution, static cache, and operator ABI differences.
These YAMLs are explicitly C/design-level examples, not fabricated E2E proof.
Qwen Image Edit reference/runtime PNGs, Whisper audio-to-text parity,
PersonaPlex reference/runtime WAVs and latency, CogVideoX animated output and
runtime parity/performance, Stable Diffusion generated output, and ESM-2 /
ProtBert embeddings, parity, batching, and throughput.
Verification
textproto migration.
all 35 non-ignored QMoE tests.
cargo clippy -D warningspasses for metadata, ComfyUI, engine, and loadertouched crates.
git ls-files '*.onnx'is empty.The evidence index labels unmeasured performance and reference-only artifacts
as gaps rather than estimating or overstating them.
Final manifest simplification
schema_versionis now the sole workflow syntax version;pipeline.workflow.manifest.ir_versionis rejected.manifest.onnx_opsetsis rejected.ports.rolesand state aliases remain authoritative, including numeric layer pairing and independently shaped K/V tensors.