Skip to content

Redesign inference metadata as a generic control-flow IR - #828

Merged
justinchuby merged 158 commits into
mainfrom
justinchuby/simplify-composite-metadata
Aug 21, 2026
Merged

justinchuby merged 158 commits into
mainfrom
justinchuby/simplify-composite-metadata

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 12, 2026 •

Copy link
Copy Markdown
Owner

Redesigns the portable inference metadata contract around one typed structural
workflow IR (sequence / invoke / loop / branch / emit), with no phases,
strategies, or model-family dispatch.

The governing rule is an ownership split. Metadata declares structural and
semantic facts a runtime cannot infer. The runtime owns deployment policy
—
memory budget, placement, EP, KV allocation and paging, compaction algorithm,
cache keys, scheduling, and the request/sequence table.

docs/genai/INFERENCE_METADATA_DECISIONS.md is now the normative specification
(goals/non-goals, terminology, ownership layers, schema model, workflow
semantics, batching, LoRA, multimodal, cache dependencies, state, sessions,
speculative, sharding, task profiles, legacy import, migration, invariants, and
conformance).

Removed, fail-closed

Removed Why Replacement
slot_ids, request_epochs, emit.row_ids, RuntimeInputRole::{RowIds,RequestEpochs} serialized the scheduler's private row identity into a portable package TensorContract.batch_layout (shared / request_aligned{axis} / token_packed{offsets,owner} / runtime_sequence_state) plus a runtime-minted RowSelection gather
package-level KV quant mode and tolerance selected a runtime-private storage representation from the package graph-visible representation inferred from dtype/Q-DQ/scale ports/Attention attrs; runtime-private KvQuantPolicy in onnx-genai-kv
shared_buffer allocator flag allocator policy, not contract aliasing: permitted | required | forbidden — the real graph ABI constraint it stood in for; defaults to forbidden
KvCacheSpec, QuantizationAxis, ForkPrecisionPolicy runtime-owned policy state group contracts; runtime-owned

Retired fields are rejected by deny_unknown_fields, so a stale document fails
rather than being silently reinterpreted.

Added

  • Mandatory row-scoped ABI. RowScopedState::compact(permutation) /
    release(row). LoRA, native grammar, and vision state survive batch
    compaction. An invariant, not a negotiated capability.
  • Derived cache dependencies. cache_dependencies() walks workflow SSA,
    state writes, and component dataflow, and always includes adapter artifacts,
    externally-suppliable encoder results, and generation-affecting profiles. A
    dependency cannot be omitted by silence.
  • Equivalence classes. bitwise | distribution_preserving | semantic decide
    whether the runtime may substitute an equivalent implementation unasked. An
    absent contract counts as semantic.
  • Effects. Retry class and speculation_safety are independent axes. The
    validator reads only speculation_safety, and holds every component in the
    enclosing loop body to the bound — not just the proposer and target.
  • Versioned checkpoint adapter, the only portable cross-build state path.
    Runtime-owned state may not be published without one, and publication is
    detected on the emitted value, so an emit cannot export private state under
    an alias.
  • Canonical semantic identity (semantic_identity()) binding disposable
    plans and checkpoints to metadata semantics. Identity, not integrity or trust.
    Normalization is conservative: it never merges two documents it cannot prove
    equivalent from syntax alone.
  • Structural generation contract. A caller may override only fields backed by
    a request-sourced typed workflow input with declared constraints. Anything else
    fails loudly.
  • Package facts — exact tokenizer/vocab/special-token bytes, constraint
    language dialect and version, task profiles with required-vs-ignorable
    semantics, and legal TP/PP/EP sharding facts.
  • One-way fail-closed legacy import (import_genai_config, --allow-lossy).
    No reverse synthesizer, because one would silently approximate.

Latest integration

  • Merged current main into the feature branch (no rebase), preserving the
    workflow-derived decoder ABI and current native/MTP changes.
  • Restored a one-way, fail-closed ComfyUI API-workflow importer that lowers into
    the same canonical pipeline.workflow IR; the runtime never dispatches on
    ComfyUI class_type; the crate remains in all offline build/clippy lanes and the crates.io publish order.
  • Added fixed recurrent state with update: { kind: replace }. Linear-attention
    accumulators and causal-convolution history share kind: recurrent but remain
    separate groups because their shapes, ports, rollback, and checkpoint
    boundaries differ. Replacement state is invariant and cannot declare a
    sequence_axis; append/indexed-scatter state must declare one.
  • Migrated every checked-in ONNX graph fixture to reviewable
    *.onnx.textproto. Path-based loading preserves external-data descriptors, so
    QMoE mmap/offload coverage remains real without a binary graph. A regression
    test rejects any tracked *.onnx graph.
  • Added a validated 20-example model catalogue
    covering Gemma 4, Cosmos3 Edge/world rollout, Qwen3.5 VLM and hybrid speculative decoder, Whisper, Wav2Vec2,
    PersonaPlex, Stable Diffusion, Qwen Image Edit, CogVideoX, LoRA, speculative
    decoding, ESM-2, ProtBert, WeatherNext, local/sliding attention, linear
    attention, causal convolution, static cache, and operator ABI differences.
    These YAMLs are explicitly C/design-level examples, not fabricated E2E proof.
  • Added durable real-output evidence:
    Qwen Image Edit reference/runtime PNGs, Whisper audio-to-text parity,
    PersonaPlex reference/runtime WAVs and latency, CogVideoX animated output and
    runtime parity/performance, Stable Diffusion generated output, and ESM-2 /
    ProtBert embeddings, parity, batching, and throughput.

Verification

  • Metadata schema sync and the complete metadata test suite pass.
  • Catalogue validation passes for all 20 YAMLs.
  • ComfyUI: 42 unit tests, 5 golden tests, and 5 engine E2E tests pass.
  • Workflow/static-cache/adapter/performance conformance E2Es pass after the
    textproto migration.
  • ORT textproto loading and native external-weight QMoE tests pass, including
    all 35 non-ignored QMoE tests.
  • cargo clippy -D warnings passes for metadata, ComfyUI, engine, and loader
    touched crates.
  • git ls-files '*.onnx' is empty.

The evidence index labels unmeasured performance and reference-only artifacts
as gaps rather than estimating or overstating them.

Final manifest simplification

  • schema_version is now the sole workflow syntax version; pipeline.workflow.manifest.ir_version is rejected.
  • ONNX opset imports are read from each component artifact; duplicated global manifest.onnx_opsets is rejected.
  • Semantic ports.roles and state aliases remain authoritative, including numeric layer pairing and independently shaped K/V tensors.

@codecov

codecov Bot commented Aug 12, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 81.57635% with 374 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.54%. Comparing base (798babb) to head (92d1217).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
crates/onnx-genai-comfyui-config/src/recognize.rs 62.31% 163 Missing and 62 partials ⚠️
...enai-comfyui-config/src/bin/comfyui_to_metadata.rs 0.00% 72 Missing ⚠️
crates/onnx-genai-comfyui-config/src/graph.rs 73.33% 35 Missing and 9 partials ⚠️
crates/onnx-genai-comfyui-config/src/lower.rs 98.76% 9 Missing and 3 partials ⚠️
crates/onnx-genai-comfyui-config/src/plan.rs 85.93% 9 Missing ⚠️
crates/onnx-genai-comfyui-config/src/lib.rs 93.04% 1 Missing and 7 partials ⚠️
crates/onnx-genai-cli/src/interactive.rs 60.00% 3 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #828      +/-   ##
==========================================
- Coverage   81.54%   80.54%   -1.00%     
==========================================
  Files         384      394      +10     
  Lines      180801   184785    +3984     
  Branches   180801   184785    +3984     
==========================================
+ Hits       147427   148839    +1412     
- Misses      28431    30820    +2389     
- Partials     4943     5126     +183     
Flag Coverage Δ
cli-ort-linux 72.47% <63.63%> (-10.14%) ⬇️
cli-ort-windows 72.06% <63.63%> (-10.14%) ⬇️
mlas 85.33% <100.00%> (+0.11%) ⬆️
offline 80.70% <81.58%> (-0.74%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/mlas-sys/src/work_stealing_pool.rs 92.28% <100.00%> (+0.96%) ⬆️
crates/onnx-genai-cli/src/generate.rs 94.51% <100.00%> (+3.31%) ⬆️
crates/onnx-genai-cli/src/lib.rs 92.41% <ø> (-1.15%) ⬇️
crates/onnx-genai-cli/src/model_inspection.rs 0.00% <ø> (-55.43%) ⬇️
crates/onnx-genai-comfyui-config/src/layout.rs 100.00% <100.00%> (ø)
...-genai-genai-config/src/bin/import_genai_config.rs 0.00% <ø> (ø)
...rates/onnx-genai-genai-config/src/compatibility.rs 61.63% <ø> (-19.57%) ⬇️
crates/onnx-genai-genai-config/src/graph_io.rs 62.00% <ø> (-7.63%) ⬇️
crates/onnx-genai-genai-config/src/import.rs 91.95% <ø> (ø)
...rates/onnx-genai-genai-config/src/json_builders.rs 60.00% <ø> (-26.65%) ⬇️
... and 31 more

... and 18 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@justinchuby
justinchuby marked this pull request as draft August 12, 2026 17:41
@justinchuby justinchuby changed the title Simplify composite inference metadata Redesign inference metadata as a generic control-flow IR Aug 12, 2026
@justinchuby

Copy link
Copy Markdown
Owner Author

Branch join blocker fixed in ad8e5e4f1a766eb95787c97266a6ed554e1b28ec.

Producer contract:

  • WorkflowNode::Branch.outputs: output/phi name -> { cases: {case: ValueRef}, default?: ValueRef }
  • WorkflowNode::Branch.effects: effect name -> { incoming, cases: {case: local_successor}, default?: local_successor, produces: joined_successor }
  • every case/default mapping is required, phi contracts must unify, joined effect tokens are distinct and globally linear, and only declared phi outputs plus explicit emits escape case scope

Updated paths: crates/onnx-genai-metadata/src/schema/ir.rs, crates/onnx-genai-metadata/src/validation.rs, schema/inference_metadata.schema.json, docs/WORKFLOW_POLICY_COMPONENTS.md, and engine/runtime tests.

Validation: metadata 20+49+2, engine 362 passed/1 ignored, workflow policy E2E 6, ORT admission 16. PR remains draft for Mobius generated-package validation.

@justinchuby

Copy link
Copy Markdown
Owner Author

Pushed preprocessing-adapter SSA support in f72782c64fbe04ac6e12593a69ecff47781f023b.

Key contract:

  • adapter ABI onnx-genai.image-preprocess@1, pinned in pipeline.workflow.manifest.adapter_abis
  • ordinary invoke input port encoded: uint8[encoded_bytes]
  • every preprocessing.image.outputs[] now requires source; workflow documents additionally require contract, a non-optional output, and name equal to an invoke-produced SSA value
  • runtime registry fails unsupported ABI versions at load and executes the adapter through generic invoke
  • ORT workflow admission no longer interprets preprocessing names as legacy component.input endpoints

Paths: crates/onnx-genai-metadata/src/schema/pipeline.rs, crates/onnx-genai-metadata/src/validation.rs, crates/onnx-genai-engine/src/pipeline/iterative.rs, docs/WORKFLOW_POLICY_COMPONENTS.md, generated schema, and E2E in crates/onnx-genai-engine/tests/workflow_policy_e2e.rs.

Validation: metadata 20 unit + 50 fixtures + 2 schema; preprocess 54; ORT admission 16; engine lib 362 passed/1 ignored; workflow policy E2E 7; formatting/schema/diff checks pass. Read-only review found no significant issues. PR remains draft pending cross-repo generated package validation.

@justinchuby

Copy link
Copy Markdown
Owner Author

Additional state soundness fix pushed in 21f6994dd34bd93a55e5853462cb8aeef4d2cd0c: every carried state initializer must exist after setup, satisfy the declared contract at metadata validation, and match the concrete runtime value before loop entry. Metadata fixtures (50) and workflow E2E (7) pass. This exposed a remaining Mobius decoder producer mismatch ([B,T] prompt declared as initializer for [B,1] token state), reported on #478.

@justinchuby

Copy link
Copy Markdown
Owner Author

Pushed generic loop induction SSA in e877270123d38d49012e70ecbf7254e720cd3431.

Contract:

kind: loop
iteration:
  value: loop.i
  contract: { dtype: int64, rank: 1, shape: [batch] }

The value is zero-based and materialized before each body execution; rank-0 scalar and explicit rank-1 broadcast contracts are supported. It is visible in the body/condition, does not escape the loop, and nested loops must use distinct lexical names. Reverse/remaining indices stay ordinary ONNX-derived values.

Validation rejects wrong dtype/rank/shape, shadowing, and post-loop references. Runtime E2E now binds diffusion.step directly into the solver and verifies [0,1,2]; nested loops verify outer [0,1] and inner [0,1,2,0,1,2]. Generated schema and docs/WORKFLOW_POLICY_COMPONENTS.md are updated.

Tests: metadata 20 unit + 51 fixtures + 2 schema; workflow E2E 8; engine lib 362 passed/1 ignored; formatting/schema/diff checks pass. Read-only review found no significant issues.

@justinchuby

Copy link
Copy Markdown
Owner Author

Cross-repo Mobius workflow execution update (2ad2e0d202902828b062e8f3a04199e6df3ee06f):

  • Executed packages generated from Mobius PR feat(pipeline): drive the decoder via a stateful PipelineDecoderComponent trait (native multi-component inc2a) #478 commit 774448b79098d92416c1aaca94c28d2972f3454b with its actual policy artifacts.
  • Masked-diffusion and codec packages load and execute end-to-end.
  • Decoder request binding, ONNX invokes, loop/effects, KV growth, and append emits execute end-to-end with a fixed-last-logit synthetic decoder.
  • Fixed generic runtime defects found by the producer packages:
    • component shape symbols are scoped per invocation, so prompt sequence=T does not conflict with decode sequence=1;
    • package dynamic-state symbol names no longer weaken component-local port unification;
    • scalar defaults can materialize singleton symbolic vectors (for example one EOS id);
    • adapters retain package symbols for allocation while validating all ports in an invocation-local symbol environment.
  • Added regular regression E2Es plus ignored cross-repo conformance tests in crates/onnx-genai-engine/tests/mobius_workflow_conformance.rs.

Validation:

  • workflow policy E2E: 11 passed
  • Mobius generated conformance: 3 passed
  • engine library: 362 passed, 1 ignored
  • metadata: 20 unit + 51 fixture + 2 schema passed
  • ORT pipeline admission: 16 passed
  • formatting and git diff --check: passed
  • final read-only code review: no significant issues

One producer-side blocker remains for normal decoders: full setup logits [B,T,V] are still declared as invariant carried state and updated with body logits [B,1,V]. Reported with the exact lowering fix at onnxruntime/mobius#478 (comment). Keeping this PR draft until that generated metadata is corrected and re-run.

@github-actions

github-actions Bot commented Aug 12, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/small_generic_f32_threads=8/1x256x256 32.22 µs 147.93 µs +359.1%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 40.40 µs 132.13 µs +227.1%
🔴 matmul/small_generic_f16_threads=8/1x256x256 28.35 µs 89.13 µs +214.3%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 29.88 µs 81.09 µs +171.4%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 40.90 µs 103.47 µs +153.0%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 516.59 µs 1.26 ms +144.0%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 361.58 µs 714.49 µs +97.6%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 28.00 µs 49.88 µs +78.1%
🔴 gather/large_f32_threads=1-internal/131072 27.36 µs 46.46 µs +69.8%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 3.72 ms 6.18 ms +66.1%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 349.42 µs 540.97 µs +54.8%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.26 ms 1.89 ms +50.0%
🔴 matmul/medium_generic_f32_threads=1/32x512x512 2.17 ms 3.16 ms +45.3%
🔴 gather/medium_bf16_threads=1-internal/32768 2.25 µs 3.04 µs +34.8%
🔴 gather/large_bf16_threads=1-internal/131072 11.50 µs 15.37 µs +33.7%
🔴 matmul/small_generic_bf16_threads=1/1x256x256 28.70 µs 37.36 µs +30.2%
⚠️ tokenization/encode_tokens_per_second 352.73 µs 442.51 µs +25.5%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 73.28 µs 87.64 µs +19.6%
⚠️ matmul/small_generic_f16_threads=1/1x256x256 27.89 µs 33.31 µs +19.4%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 77.97 µs 92.67 µs +18.9%
⚠️ tokenization/decode_tokens_per_second 5.76 ms 6.77 ms +17.6%
⚠️ reduce_mean/small_f32_threads=1-internal/4096 13.77 µs 16.03 µs +16.4%
✅ gather/small_f16_threads=1-internal/4096 440.2 ns 502.8 ns +14.2%
✅ gather/small_bf16_threads=1-internal/4096 440.3 ns 500.0 ns +13.6%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 8.68 ms 9.69 ms +11.6%
✅ add/small_bf16_threads=1-internal/1024 12.04 µs 13.41 µs +11.4%
✅ gather/medium_f32_threads=1-internal/32768 3.49 µs 3.85 µs +10.3%
✅ add/large_bf16_threads=1-internal/4194304 39.02 ms 42.97 ms +10.1%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 489.81 µs 538.96 µs +10.0%
✅ add/medium_f32_threads=1-internal/262144 2.33 ms 2.56 ms +9.6%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.83 ms 2.00 ms +9.3%
✅ add/large_f16_threads=1-internal/4194304 37.91 ms 41.34 ms +9.0%
✅ add/medium_f16_threads=1-internal/262144 2.37 ms 2.58 ms +8.8%
✅ grammar_masking/llguidance_compute_mask/32 75.34 µs 81.97 µs +8.8%
✅ gather/small_f32_threads=1-internal/4096 635.3 ns 690.7 ns +8.7%
✅ add/medium_bf16_threads=1-internal/262144 2.45 ms 2.65 ms +8.3%
✅ reduce_mean/medium_f32_threads=1-internal/65536 227.51 µs 246.20 µs +8.2%
✅ reduce_mean/large_f32_threads=1-internal/262144 914.98 µs 988.22 µs +8.0%
✅ sampling_latency/greedy_per_token 3.23 µs 3.47 µs +7.6%
✅ add/large_f32_threads=1-internal/4194304 38.29 ms 41.08 ms +7.3%
✅ matmul/small_generic_f32_threads=1/1x256x256 33.72 µs 36.13 µs +7.1%
✅ gather/medium_f16_threads=1-internal/32768 2.30 µs 2.47 µs +7.0%
✅ add/small_f16_threads=1-internal/1024 11.77 µs 12.56 µs +6.8%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 74.09 µs 78.57 µs +6.0%
✅ gather/large_f16_threads=1-internal/131072 13.20 µs 13.98 µs +6.0%
✅ kv_cache/alloc_dealloc_pages 38.12 µs 40.35 µs +5.8%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.82 ms 6.06 ms +4.0%
✅ matmul/medium_generic_f32_threads=8/32x512x512 951.04 µs 987.61 µs +3.8%
✅ sampling_latency/top_p_per_token 372.78 µs 386.62 µs +3.7%
✅ matmul/medium_generic_f16_threads=1/32x512x512 28.08 µs 28.73 µs +2.3%
✅ logit_processing/seven_processor_chain_per_step 320.14 µs 327.08 µs +2.2%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 532.60 µs 541.33 µs +1.6%
✅ qwen3_sampling_processors/top_k_top_p_fast 663.55 µs 674.25 µs +1.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.17 ms 2.16 ms -0.6%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.60 ms 3.53 ms -1.8%
✅ sampling_latency/min_p_per_token 218.72 µs 213.93 µs -2.2%
✅ add/small_f32_threads=1-internal/1024 197.0 ns 191.2 ns -2.9%
✅ qwen3_sampling_processors/top_k_partial_selection 146.46 µs 139.79 µs -4.5%
✅ sampling_latency/top_k_per_token 64.64 µs 60.38 µs -6.6%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.04 3.97 4.87 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

Legacy-removal milestone pushed in 7c4c4fa8a944adaf79a080fcf1789a880fbc48bd.

  • pipeline.workflow is now the only composite execution contract; legacy strategy/phases/schedulers and family-specific pipeline executors are deleted.
  • Composite compatibility synthesis, host image/audio family APIs, ComfyUI translation, legacy fixtures/examples, and stale routes/tests are removed.
  • Generic workflow execution, serving/resource governance, and KV mechanisms remain.
  • Review found branch-local session state could be dropped without a phi; validation now requires an escaping branch output and runtime commits the joined value.

Validation: metadata 24/24; workflow policy E2E 11/11; engine 308 passed (1 ignored); server 146 passed (2 ignored); ORT 116/116; CLI 105/105; all affected targets checked. Full workspace check remains blocked by the pre-existing missing crates/onnx-runtime-cpuinfo/vendor/cpuinfo/CMakeLists.txt. Final read-only review reported no significant issues.

@justinchuby

Copy link
Copy Markdown
Owner Author

Pushed producer/runtime contract unblocker in a9d6c5e0a0deafb18775b28a25b9c9ba8450ab4b.

  • emit.valid_length accepts a typed integer scalar/rank-1 SSA value, slices the emitted tensor's final-axis prefix, and works with replace/append/event while retaining final package-output validation.
  • Added generic bounded state recurrence for ONNX-computed truncation/selection, with branch-phi + subsequent-invoke conformance coverage. No host KV rollback fallback was added.
  • Updated the producer contract at docs/WORKFLOW_POLICY_COMPONENTS.md and regenerated schema/inference_metadata.schema.json.

Validation:

  • metadata: 20 unit + 3 fixture + 2 schema tests
  • workflow policy E2E: 12 passed
  • engine: 308 passed, 1 ignored (plus integration targets passed/ignored as configured)
  • ORT: 116 unit tests plus integration targets passed
  • genai-config: passed
  • formatting, schema sync, and affected crate checks: passed
  • read-only architecture review: no significant findings

Known unrelated baseline failures remain in removed legacy fixture tests: CLI image REPL tests reference missing tests/fixtures/tiny-vlm-image-input; two server HTTP JSON-format tests exceed the tiny fixture's context limit. The directly affected library/runtime suites pass. PR remains draft for Mobius cross-repo coordination.

@justinchuby

Copy link
Copy Markdown
Owner Author

Pushed the approved grammar/adaptive-K contracts in 1b1b42acce8679b78b7959b20ab6d828cc7a28dc.

Producer targets for #478:

  • Grammar ABI: onnx-genai.grammar-guidance@1, adapter.role: grammar_guidance, actions clone|lookahead|commit.
  • Grammar ports: state int64[B], tokens int64[B,T], valid_length int64[B], transition_table int64[S,V]; outputs next_state, consumed_length, logits_mask bool[B,V], forced_tokens int64[B,1], forced_length int64[B].
  • Adaptive ONNX policy: role: adaptive_proposal_budget with current_k, accepted, evaluated, committed_tokens, filled_proposal_budget, draft_ms, target_ms, estimates; outputs next_k, next_estimates.
  • State cells now declare class: semantic|advisory (semantic default). Advisory state is invocation-only/resettable and not session-serialized.
  • Optional timing ABI: onnx-genai.telemetry@1, actions start|elapsed.

Runtime now executes grammar clone/lookahead/commit and telemetry generically. The conformance workflow combines speculative verification, grammar-valid-prefix truncation, committed semantic grammar state, forced-token ONNX sampling, and loop-carried advisory adaptive K.

Exact producer documentation: docs/WORKFLOW_POLICY_COMPONENTS.md
Generated schema: schema/inference_metadata.schema.json

Validation: metadata 20+4+2, workflow policy E2E 14, engine 308/1 ignored, ORT 116, server lib 146/2 ignored, config, formatting, schema sync, and affected checks passed. Read-only review found no significant issues. PR remains draft for cross-repo generated-package validation.

@justinchuby

Copy link
Copy Markdown
Owner Author

World-model v1 acceptance fixture is pushed in 005acf49574959fd2e645518720e5c2929d05276.

Coverage:

  • executable synthetic ONNX observation encoder
  • semantic latent state with exclusive session ownership
  • ONNX action policy
  • action-driven branch across two ONNX environment-step components
  • loop carry and linear state/effect tokens
  • streaming action events and final latent emit
  • checkpoint_session / transactional restore_session_checkpoint
  • state advancement, restore, and deterministic replay from the checkpoint
  • advisory state explicitly excluded from checkpoint serialization

The checkpoint type records state contracts and rejects incompatible restores. Review identified a possible partial mutation if cloning failed during restore; restoration now stages every clone before touching live session state.

Validation: workflow policy E2E 15 passed; engine 308 passed/1 ignored; server library 146 passed/2 ignored; formatting and affected checks passed.

The PR remains draft only for the separate acceptance requirement that every existing supported family be demonstrated with its real producer graph/E2E; the synthetic world-model requirement itself is complete.

@justinchuby

Copy link
Copy Markdown
Owner Author

Execution-island milestone pushed in e0bc7f3.

  • Links adjacent pure same-device ONNX invokes into optimizer-visible sessions.
  • Uses stable CPU/CUDA IoBinding with per-shape CUDA capture/replay diagnostics and eager fallback.
  • Keeps stateful adapters, effects, overrides, control flow, and device changes as island boundaries.
  • Loads linked external initializers through ORT memory-file injection instead of embedding weights in generated protobufs.
  • Adds request-bound min-p -> sampler -> termination E2E plus external-weight and boundary-name collision regressions.

Validation: metadata suite 28 passed; workflow E2E 19 passed; engine lib 309 passed/1 ignored; ORT lib passed; config passed; CUDA feature check passed; formatting/schema sync passed. Read-only review finding on sanitized SSA-name collisions was fixed and covered by E2E.

@justinchuby
justinchuby force-pushed the justinchuby/simplify-composite-metadata branch from e0bc7f3 to b2157a2 Compare August 13, 2026 00:30
@justinchuby

Copy link
Copy Markdown
Owner Author

Trailer-only amend: execution-island milestone head is now b2157a26308a2d882c28611858b641e86e8815e5 (supersedes e0bc7f3b).

@justinchuby

Copy link
Copy Markdown
Owner Author

Mobius producer update 3c17705 now emits request-parameterized sampling_min_p and documents planner-derived execution islands/capture against schema b2157a2.

One exact runtime blocker remains for the required decoder policy island: crates/onnx-genai-engine/src/pipeline/islands.rs::pure_onnx_device rejects every component with application_overridable = true. Mobius samplers are intentionally overrideable, so even when no override is selected, the sampler always splits decoder/logits-processing/sampling/termination islands.

Required generic lowering: resolve application overrides before island partitioning, then test the resolved implementation's device, ONNX purity, ports, and effects. The package default and a selected pure same-device ONNX replacement should remain fusible; a stateful/host replacement should delimit the island. This preserves the versioned override ABI without sacrificing optimizer visibility or CUDA Graph capture. Please add conformance coverage for both the no-override and pure-ONNX-override cases.

@justinchuby
justinchuby force-pushed the justinchuby/simplify-composite-metadata branch from eadb31f to 8bacf8c Compare August 13, 2026 00:58
@justinchuby

Copy link
Copy Markdown
Owner Author

Performance conformance update (8bacf8c5):

  • Added workflow/island telemetry for TTFT, throughput, session boundaries, capture/replay, transfers, syncs, stable/external/device memory, and fallback reasons.
  • Removed boundary Identity nodes; logical component boundaries now add zero executable ONNX nodes.
  • Added paired 5-sample native-composite benchmarks and speculative capture conformance.
  • CPU medians (100 iterations): decoder policy ratio 1.032, min-p ratio 0.962.
  • H200 / ORT 1.28 CUDA medians: decoder ratio 0.903, min-p ratio 0.957; both captured once and replayed 503 times. Warm TTFT was 3.03 ms vs 17.93 ms and 4.36 ms vs 3.64 ms.
  • Cold workflow startup remains higher (467 ms vs 49 ms decoder; 231 ms vs 18 ms min-p) because the first run discovers output extents and creates stable bindings. This is documented as a prewarm/planner-shape optimization, not hidden.
  • Full engine: 309 passed, 1 ignored. Metadata/config/ORT suites passed. Speculative verifier/policy islands captured and replayed on CUDA.

The PR remains draft. Real producer-package/KV/per-row serving benchmarks and release ORT kernel/provider profiles remain readiness gates.

@justinchuby

Copy link
Copy Markdown
Owner Author

Producer review against current 8bacf8c found four exact schema/tooling blockers for the next #478 migration:

  1. Per-sequence speculative emit: WorkflowStep::Emit.valid_length still documents/validates scalar or rank-one with exactly one element, and runtime reads it through scalar extraction. Mobius must emit accepted_len int64[B] without ReduceMin. The surface/lowered schema and executor need per-row valid lengths aligned to value's batch axis, producing ragged/per-row appended output rather than one synchronized prefix.
  2. Per-row speculative state: WorkflowStateCell/carry has only one tensor recurrence and no per-row active/logical-length contract. Generic KV service needs loop-carried active_mask[B] and logical cache lengths [B], with rollback/truncation driven per row while physical dense/paged storage remains runtime-owned. Please define the state/emit admission rules; Mobius will not encode KV semantics in host metadata.
  3. Concise loop continuation: public WorkflowStep::Loop still requires condition; there is no continue_when/until polarity field. Mobius therefore still needs a pure boolean-Not artifact for done -> continue. Lexical iteration is now consumed directly at Mobius ada6303; the increment artifact is removed.
  4. Shipped semantic validator: onnx-genai-metadata exposes library validation but no validator binary/example. Mobius CI cannot invoke the shipped semantic validator on generated package directories without maintaining Rust glue. Please ship a stable CLI accepting metadata/package paths and artifact-root validation, so decoder/VLM/diffusion/TTS/speculative generated fixtures can be unconditional cross-repo tests instead of JSON-schema-only checks.

Also, workflow components/state currently have no explicit generic KV grouping/past-present alias/sequence-axis semantic fields beyond tensor names and shape recurrence. Please identify the intended existing contract or add those fields; Mobius will keep ordinary ONNX ports artifact-inferred and emit only this non-inferable KV service metadata.

@justinchuby

Copy link
Copy Markdown
Owner Author

Producer blocker contract is finalized and pushed at 8215649100a0a27be15b04045fddde777c8248fc (supersedes 5a2bcfb3). Mobius can resume migration against this commit.

Exact surfaces:

  • schema: schema/inference_metadata.schema.json
  • public/lowered IR: crates/onnx-genai-metadata/src/schema/ir.rs, crates/onnx-genai-metadata/src/lowering.rs
  • semantic/package validator: crates/onnx-genai-metadata/src/validation.rs, crates/onnx-genai-metadata/src/parser.rs
  • CLI: crates/onnx-genai-metadata/src/bin/validate_metadata.rs
  • migration guide: docs/WORKFLOW_POLICY_COMPONENTS.md (v1 producer migration checklist)

Migration rules:

  1. condition -> pre-test continue_when; initialize before entry. False initially is zero-trip. Carry the next active value when body policy changes it.
  2. Emit valid_length: int64[B] directly. Runtime produces output.row.<row>; event mode adds .<event>. Do not ReduceMin. Later appends remain in the same row stream.
  3. Serving values are active: bool[B], done: bool[B], accepted_len: int64[B], slot_ids: int64[B]. Inactive rows preserve logical carries.
  4. Cache state sets service_group. Each KV group declares sequence_axis, layout, semantic logical_lengths state (int64[B]), storage, and component/cell {input, output} aliases. shared_buffer explicitly permits past/present aliasing.
  5. Request source is only source: { kind: request }; the versioned runtime role is the sole request-field identity.
  6. Validate every package directory with cargo run -p onnx-genai-metadata --bin validate_metadata -- <package-dir>. It now checks semantic invariants, required artifacts, and package-root containment.

Conformance now includes B=2 speculative per-row acceptance/ragged emit, B=1 stable ragged namespace, mixed ragged+dense appends, active-row carry, KV logical-length validation, zero-trip, and valid/missing package artifacts.

Validation: engine 313 passed / 1 ignored; workflow E2E 19 passed; metadata/schema/CLI/config suites all passed; formatting and diff checks passed.

@justinchuby

Copy link
Copy Markdown
Owner Author

Pushed 5bd4783b786292c2140b5eb2de856ab941917870 for the next North Star runtime milestone.

  • keeps stable CUDA island outputs device-resident and materializes only at host control/adapter/ragged emit/public output boundaries
  • separates package outputs from internal SSA values
  • tracks tensor device IDs for correct multi-device materialization
  • fuses default application-overridable ONNX components and preserves a validated unfused replacement path
  • records host/device copies and synchronizing host transfers

Validation: engine suite passed (313 passed, 1 ignored); workflow policy E2E 19 passed; ORT suite passed (116 unit plus integration tests); metadata/config/schema/validator passed; CUDA feature compilation passed; formatting/diff checks passed. CPU native-matched median ratios: decoder 0.973, min-p 0.978. H200 benchmark could not run because the linked downloaded ORT exposes CPU only and lacks libonnxruntime_providers_cuda.so; the CUDA path compiled successfully but runtime capture/copy metrics still require a CUDA-enabled ORT run. PR remains draft.

@justinchuby
justinchuby force-pushed the justinchuby/simplify-composite-metadata branch from 5bd4783 to 76a37d8 Compare August 13, 2026 01:52
@justinchuby

Copy link
Copy Markdown
Owner Author

Trailer-order-only amend: the milestone commit is now 76a37d81710bac1c85bdf888fabc1406848b9f78 (supersedes 5bd4783b; tree contents unchanged).

@justinchuby

Copy link
Copy Markdown
Owner Author

North Star follow-up pushed:

  • 24f75804: fixes session checkpoint identity when multiple logical cells share one initializer; the executable world-model fixture now proves semantic replay while advisory session state continues independently.
  • 040c372e: removes top-level strategy, structured_output, generation, and tokens, closes the v1 root schema, regenerates JSON Schema (-576 lines), and rejects stale fields rather than silently accepting them. EOS defaults now come from tokenizer/request runtime inputs.

Validation: metadata 32/32 passed; engine 312 passed, 1 ignored plus all engine integration targets (workflow policy 19/19). Workspace-wide check remains blocked by the pre-existing uninitialized crates/onnx-runtime-cpuinfo/vendor/cpuinfo submodule.

H200 native-matched benchmark (ORT 1.27 CUDA 13, identical linked composite/session I/O/capture): decoder workflow/native 362.85/371.06 step/s = 0.978; min-p 205.57/207.01 = 0.993. Both captured once and replayed 1003 times with no fallback. Decoder workflow TTFT remains 20.60 ms vs 2.78 ms native due planner/session startup; steady-state graph execution meets the preliminary bar. Request upload/final materialization are included in aggregate transfer diagnostics; linked internal component boundaries do not round-trip through host.

@justinchuby

Copy link
Copy Markdown
Owner Author

Additional atomic cleanup pushed as c9bddd6e: deleted the now-unreferenced Rust API types and extensible vocabularies for legacy strategy, draft/verify strategy nodes, structured output, and special tokens (198 more lines removed). Metadata tests and engine compilation pass.

@justinchuby

Copy link
Copy Markdown
Owner Author

Explicit conformance status: implemented and executed, not schema-only.

  • 398b2985: the runtime grammar adapter executes clone (zero consumption), speculative lookahead over proposals, and accepted-prefix commit; it produces logits masks/forced tokens and carries grammar as semantic state. The adaptive-K artifact is an actual ONNX policy consuming accepted/evaluated/committed counts, filled-budget, advisory estimates, and draft/target latency. Tests vary K and timing: budget outputs change (3/5/4) while emitted [1,2,3] and final grammar state remain invariant. The generic telemetry adapter start/elapsed also has an executable runtime E2E.
  • 000f09b2: removed top-level speculative/speculator_config, implicit metadata-driven native MTP/shared-KV loading, stale precedence tests/fixtures, and 448 more schema lines. Speculation must now be workflow-executable or explicit transitional runtime config.

Validation: metadata all passed; ORT 116 unit plus integration targets passed; engine 310 passed/1 ignored plus all integration targets; workflow policy 19/19; native-backend compilation passed. PR remains draft for checked-in Mobius fixtures and the remaining explicit native speculative path/KV top-level migration.

@justinchuby

Copy link
Copy Markdown
Owner Author

Further legacy migration pushed as 3f9ab738: model-directory loading and CLI inspection no longer auto-discover HuggingFace speculator_config; stale speculator fixtures/tests are deleted. CLI sampling conformance now passes temperature/top-k explicitly as request policy inputs rather than relying on removed metadata defaults. Targeted request-sampling REPL tests, ORT 116-test suite, CLI/server compilation, and prior full engine suite pass. Remaining migration is the explicit transitional EngineConfig::speculative_mode implementation and top-level legacy KV declaration; generic workflow KV services remain untouched.

@justinchuby

justinchuby commented Aug 13, 2026 •

Copy link
Copy Markdown
Owner Author

Mobius checked-in producer fixtures are now available at onnxruntime/mobius@3b0445aec64868e7f8ecce51fd7db3aba57455de under tests/fixtures/onnx_genai_workflows/ (decoder, VLM, diffusion, masked diffusion, real tiny Qwen3-TTS, speculative, codec).

They are reproducibly generated by tests/generate_onnx_genai_validation_packages.py. Mobius CI is pinned to onnx-genai c9bddd6e, diffs regenerated packages against the checked-in tree, and runs validate_metadata on every package. All seven pass semantic validation. The c9 engine mobius_workflow_conformance runtime test also passes 3/3 for decoder, masked diffusion, and codec using this exact directory.

Please point unconditional ONNX GenAI fixture CI at the commit/path above.

@justinchuby

Copy link
Copy Markdown
Owner Author

KV legacy-path migration pushed in 46f4fc41:

  • removed kv_cache, model.runtime_configurable.kv_cache, and model.io.kv_update from public metadata and regenerated schema;
  • stopped engine/ORT defaults from selecting static/shared KV from model family, dtype, inference metadata, or artifact custom metadata;
  • preserved explicit low-level ORT static/shared APIs and workflow kv_service paging/shared contracts;
  • deleted stale engine static/shared batching and prefix fixtures;
  • added rejection and artifact-metadata non-selection regressions.

Validation: metadata all-targets, genai-config all-targets, engine all-targets (294 passed, 1 ignored plus integration targets; workflow policy 19/19), ORT all-targets (including 116 unit tests), CLI/server checks, native-backend check, formatting, and diff check. Independent serving/KV review found no high-confidence defects. PR remains draft only for checked-in Mobius fixture ingestion.

justinchuby and others added 10 commits August 21, 2026 17:42
Replace binary graph fixtures with reviewable ONNX TextFormat across workflow, package, ORT, and native-runtime tests. Keep external QMoE weights mmap-capable, update all artifact references, and add a guard against checked-in binary ONNX graphs.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Add durable audio, video, image, and protein outputs with upstream/runtime parity and measured performance where available. Preserve honest gaps for runs that did not retain runtime timing or a separate runtime render.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Add nineteen validated pipeline.workflow examples covering text, multimodal, speech, diffusion, scientific, recurrent-state, adapter, speculative, cache, and operator ABI cases. Document the config-only proof level and graph-visible versus runtime-private attention distinctions, and validate every catalogue YAML without requiring graph artifacts.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Reject sequence_axis on replace-updated state and remove stale binary-ONNX fixture exemptions now that all checked-in graphs use TextFormat, including external-weight QMoE fixtures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Add a validated Qwen3.5 decoder configuration that combines attention KV with linear-attention and causal-convolution replacement state. Document atomic rollback, snapshot-or-replay handling for partially accepted proposals, and persistent draft-model requirements.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Put onnx-genai-comfyui-config back into every offline build/clippy lane and the dependency-ordered crates.io release list so restoring the importer also restores its supported distribution surface.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Use schema_version as the sole workflow syntax version and read ONNX opset imports from each component artifact. Keep only semantic port roles and state aliases in metadata, and demonstrate independently shaped K/V and per-layer KV geometry in the Gemma 4 catalogue example.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Document that state role and layer provide semantic pairing and ordering while each ONNX port remains authoritative for its own geometry, including different per-layer and K-versus-V head counts.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Apply the repository Rust formatter after merging main so required quality and fast checks evaluate the combined tree cleanly.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
@justinchuby
justinchuby enabled auto-merge (squash) August 21, 2026 20:06
@justinchuby
justinchuby merged commit 0f40538 into main Aug 21, 2026
15 of 19 checks passed
@justinchuby
justinchuby deleted the justinchuby/simplify-composite-metadata branch August 21, 2026 20:22
justinchuby added a commit to onnxruntime/mobius that referenced this pull request Aug 21, 2026
## Summary

Produce the canonical ONNX GenAI `pipeline.workflow` ABI across Mobius
exports and validate it against
[onnx-genai#828](justinchuby/onnx-genai#828).

- remove serialized `model.io`, legacy phase/strategy duplication, and
runtime-private scheduling/allocation policy;
- emit typed SSA workflows for decoder generation, static cache,
speculative decoding, adapters, VLM, diffusion/image edit, video,
speech, codec/TTS, masked diffusion, ESM-2, and ProtBert;
- describe graph-visible state semantics including append,
indexed-scatter, and fixed-size recurrent replacement;
- emit canonical encoder embedding profiles and application-sourced
opaque auxiliary inputs;
- assign deterministic symbols to anonymous dynamic ONNX dimensions
instead of invalid YAML `null`;
- stop emitting duplicated workflow `ir_version` and `onnx_opsets`,
leaving schema and ONNX artifacts authoritative;
- pin cross-repository CI to ONNX GenAI commit
`509cd4e9c4471f4cbc59fe44b47168f0ae128fe3`;
- merge current `main` into the branch without rebasing.

## Validation

- 164 metadata producer tests passed, 4 skipped;
- all 13 generated package families match committed fixtures and pass
the authoritative `validate_metadata` parser/semantic validator;
- 11 ONNX GenAI engine conformance tests passed using real generated
Mobius packages;
- changed files pass lintrunner formatting.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Signed-off-by: justinchuby <justinchuby@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
Signed-off-by: Justin Chu <justinchu@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Justin Chu <justinchu@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 21, 2026
…n main) (#1686)

`Rust quality` is red on `main` and has been since #828: `cargo fmt
--all -- --check` reports five hunks in
`crates/onnx-runtime-session/src/executor/geometry.rs` and
`.../tests.rs`.

- CI job `96899749480` on `0f40538b2` fails with `Diff in
.../executor/geometry.rs:582`.
- Reproduced locally on `4c15a64b3`; both files are byte-identical to
`origin/main`, so this is not branch-local drift.

This PR is **`cargo fmt --all` output and nothing else**. Verified
locally on `4c15a64b3`:

| gate | result |
|---|---|
| `cargo fmt --all -- --check` | clean |
| `cargo check --locked -p onnx-runtime-session` | ok |
| `cargo test --locked -p onnx-runtime-session` | ok, 0 failed |

Split out of my MatMul work rather than folded into it, so the unblock
is reviewable on its own and lands for everyone.

Refs #1600.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 22, 2026
…oncat it hid

#1685 reported an intermittent 8-row hole of exact `0.0` in SDPA output under
`--features mlas`, plus a SIGSEGV and `identity specialization diverged`. The
defect is already fixed on main — incidentally, by PR #828's unrelated IR
redesign, which added `wait_for_workers` to the MLAS work-stealing pool — and
nothing guarded it. So this closes the coverage hole rather than the kernel.

Root cause, now documented on `wait_for_workers`: `wait_for_completion` returns
when the last *block* has run, but the worker that ran it is still inside
`run_job` holding a by-value copy of the old `Job`. Pre-#828 the dispatcher
released `dispatch_lock` there, so the next dispatch republished the loop
bounds while that straggler was live; it then decremented the new job's
`remaining` without running the new closure (a partition never executes, and a
`beta = 0` SGEMM leaves rows of `C` unwritten) and invoked the *old* closure
with an index from the *new* range (the SIGSEGV).

The race needs two dispatches to overlap, so it is ~3% per process, not per
call: 6000 `sdpa_f32` invocations and 3000 nested rayon x sgemm iterations
reproduced nothing. `crates/mlas-sys/tests/concurrent_dispatch.rs` instead
falsifies the fix — 4 dispatchers on an 8-thread pool. It passes in 0.10s on
main and, with `wait_for_workers` removed, one test SIGSEGVs and the other
reports the resulting deadlock through a watchdog.

Coverage closed:
- `AB_COVERED` gains `AttentionTranspose`, whose `PLAN` entry already claimed
  `Graduation::Partial` while the graduation rule reads `AB_COVERED`.
- `sdpa_f32_native` / `sdpa_f32_mlas` name the two routes, since `sdpa_f32`
  short-circuits to MLAS and cannot be the native half of an A/B.
- `native_vs_mlas_differential` gains 8 SDPA shapes x causal, a NaN-prefill
  fail-loud check that decodes a hole to (tile, row, column), a convex-
  combination oracle that still holds in a default MLAS-free build, and a
  concurrent-sessions test.
- `identity_hook_specialization_matches_the_general_epilogue` prefills NaN.
  Falsified: dropping 8 rows/tile (6144 of 98304 elements, #1685's exact
  signature) leaves the old zero-prefill version *passing*.

Two real defects the new coverage found:

1. `concat_cache` looped head-dim outside sequence, striding 512B per store
   through a contiguous `[b][h][s][d]` buffer and traversing it `dim` times —
   ~20ms of a ~28ms decode node. It is now two `copy_from_slice` calls per
   plane with the same fan-out threshold the sibling transforms use.
   `llama_decode_past1023` 27.2x -> 13.8x ORT, `chunk8` 24.8x -> 13.1x,
   `chunk32` 27.7x -> 18.5x, parity preserved.

2. No golden exercised the causal offset: every self-attention case has
   `past == 0` and the one past-KV case has `q_seq == 1`, so
   `causal = unidirectional && q_seq > 1` is false. Hard-coding
   `past_seq = 0` leaves all 12 pre-existing goldens green; the new
   `past_kv_chunked_prefill_causal` is the only one that fails (2.22 abs).

`gen_mha.py` gains causal, decode and past-KV rows, and now *rejects*
`unidirectional` with `q_seq != kv_seq` and no past-KV: a NumPy oracle swept
over every offset in `0..=kv_seq` matches ORT at none of them, so that cell has
no defined answer and its 24/384 parity failures were not a kernel bug.

Ledger section 45 records the audit, the falsifier method, the 18-cell x 4-thread
production-default matrix (432/432 parity PASS, A/A null control per cell), and
the finding it exposed: `native_p50` is flat across 1..16 threads on every cell
because `sdpa_f32_simd` — the route a default build actually takes — has no
rayon fan-out at all, while the MLAS research route does. That is filed
separately rather than folded in here.

Gates: fmt; clippy default + `--features mlas` + aarch64 `-D warnings`;
`check_cross_compile.sh`; full ep-cpu suite both feature configs; mlas-sys;
Miri task_runtime/strided/provider/dtype; default cdylib 0 MLAS symbols by
`nm`/`nm -D`/`strings`/`ldd` with an 842-symbol positive control.

Closes #1685

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 22, 2026
…oncat it hid

#1685 reported an intermittent 8-row hole of exact `0.0` in SDPA output under
`--features mlas`, plus a SIGSEGV and `identity specialization diverged`. The
defect is already fixed on main — incidentally, by PR #828's unrelated IR
redesign, which added `wait_for_workers` to the MLAS work-stealing pool — and
nothing guarded it. So this closes the coverage hole rather than the kernel.

Root cause, now documented on `wait_for_workers`: `wait_for_completion` returns
when the last *block* has run, but the worker that ran it is still inside
`run_job` holding a by-value copy of the old `Job`. Pre-#828 the dispatcher
released `dispatch_lock` there, so the next dispatch republished the loop
bounds while that straggler was live; it then decremented the new job's
`remaining` without running the new closure (a partition never executes, and a
`beta = 0` SGEMM leaves rows of `C` unwritten) and invoked the *old* closure
with an index from the *new* range (the SIGSEGV).

The race needs two dispatches to overlap, so it is ~3% per process, not per
call: 6000 `sdpa_f32` invocations and 3000 nested rayon x sgemm iterations
reproduced nothing. `crates/mlas-sys/tests/concurrent_dispatch.rs` instead
falsifies the fix — 4 dispatchers on an 8-thread pool. It passes in 0.10s on
main and, with `wait_for_workers` removed, one test SIGSEGVs and the other
reports the resulting deadlock through a watchdog.

Coverage closed:
- `AB_COVERED` gains `AttentionTranspose`, whose `PLAN` entry already claimed
  `Graduation::Partial` while the graduation rule reads `AB_COVERED`.
- `sdpa_f32_native` / `sdpa_f32_mlas` name the two routes, since `sdpa_f32`
  short-circuits to MLAS and cannot be the native half of an A/B.
- `native_vs_mlas_differential` gains 8 SDPA shapes x causal, a NaN-prefill
  fail-loud check that decodes a hole to (tile, row, column), a convex-
  combination oracle that still holds in a default MLAS-free build, and a
  concurrent-sessions test.
- `identity_hook_specialization_matches_the_general_epilogue` prefills NaN.
  Falsified: dropping 8 rows/tile (6144 of 98304 elements, #1685's exact
  signature) leaves the old zero-prefill version *passing*.

Two real defects the new coverage found:

1. `concat_cache` looped head-dim outside sequence, striding 512B per store
   through a contiguous `[b][h][s][d]` buffer and traversing it `dim` times —
   ~20ms of a ~28ms decode node. It is now two `copy_from_slice` calls per
   plane with the same fan-out threshold the sibling transforms use.
   `llama_decode_past1023` 27.2x -> 13.8x ORT, `chunk8` 24.8x -> 13.1x,
   `chunk32` 27.7x -> 18.5x, parity preserved.

2. No golden exercised the causal offset: every self-attention case has
   `past == 0` and the one past-KV case has `q_seq == 1`, so
   `causal = unidirectional && q_seq > 1` is false. Hard-coding
   `past_seq = 0` leaves all 12 pre-existing goldens green; the new
   `past_kv_chunked_prefill_causal` is the only one that fails (2.22 abs).

`gen_mha.py` gains causal, decode and past-KV rows, and now *rejects*
`unidirectional` with `q_seq != kv_seq` and no past-KV: a NumPy oracle swept
over every offset in `0..=kv_seq` matches ORT at none of them, so that cell has
no defined answer and its 24/384 parity failures were not a kernel bug.

Ledger section 45 records the audit, the falsifier method, the 18-cell x 4-thread
production-default matrix (432/432 parity PASS, A/A null control per cell), and
the finding it exposed: `native_p50` is flat across 1..16 threads on every cell
because `sdpa_f32_simd` — the route a default build actually takes — has no
rayon fan-out at all, while the MLAS research route does. That is filed
separately rather than folded in here.

Gates: fmt; clippy default + `--features mlas` + aarch64 `-D warnings`;
`check_cross_compile.sh`; full ep-cpu suite both feature configs; mlas-sys;
Miri task_runtime/strided/provider/dtype; default cdylib 0 MLAS symbols by
`nm`/`nm -D`/`strings`/`ldd` with an 842-symbol positive control.

Closes #1685

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 22, 2026
…oncat it hid (#1714)

Closes #1685.

## The defect is already fixed on `main` — and nothing was guarding it

#1685 reports an intermittent SDPA failure under `--features mlas`: ~3%
of runs leave an 8-row hole of exact `0.0` in the output, or die with a
SIGSEGV, tripping `identity specialization diverged`.

The fix landed incidentally in **`0f40538b2` (PR #828, "Redesign
inference metadata as a generic control-flow IR")**, which added 12
lines to `work_stealing_pool.rs` as a side effect. The reported-bad SHA
`8aed77a17` has **0** occurrences of `wait_for_workers`/`observed`;
current `main` has 7.

**Mechanism.** `wait_for_completion` returns when `remaining` hits zero
— when the last *block* has run. The worker that ran it is still inside
`run_job`, holding a by-value copy of the old `Job` and looping in
`claim_iterations` against the *shared* counters. Pre-fix the dispatcher
released `dispatch_lock` at that moment; MLAS fans out under rayon in
`sdpa_f32_fast`, so the next dispatch republished the loop bounds and
bumped the epoch while the straggler was live. It then did two things at
once:

1. decremented the **new** job's `remaining` without running the new
closure → `wait_for_completion` returned early → a partition **never
executed** → a `beta = 0` SGEMM left rows of `C` unwritten (the 8-row
hole: `8 · dv` contiguous zeros at a tile boundary — MLAS had split the
128-row `probs·V` GEMM 16 ways);
2. invoked the **old** closure with an index from the **new** range,
writing through raw pointers past the end of the previous GEMM's `C` →
the SIGSEGV.

## Reproduction failed; falsification worked

| attempt | result |
|---|---|
| 80 runs of the exact test | 0 failures |
| 40 runs, `sdpa` filter, `--test-threads 8` | 0 |
| pool widths 4 / 8 / 16 / 32, 40 runs each | 0 |
| 3000 iterations, nested rayon × `sgemm`, NaN-prefilled `C` | 0 |
| 6000 `sdpa_f32` invocations at the issue's shape | 0 |

The race needs two dispatches to **overlap**, and `dispatch_lock`
serialises them — so the window is only the gap between the last block
finishing and the straggler leaving `run_job`. That is ~3% **per
process**, not per invocation.

So I stopped trying to trigger it and tried to **break the fix
instead**. `crates/mlas-sys/tests/concurrent_dispatch.rs` runs 4
dispatcher threads against one 8-thread pool. On `main` it passes in
**0.10 s**. With only `wait_for_workers` removed:

- `overlapping_dispatches_never_return_with_work_unexecuted` →
**SIGSEGV**
- `parallel_for_returns_only_after_every_worker_has_left_the_closure` →
the straggler drives `remaining` below zero, `usize` wraps, and the pool
deadlocks; a watchdog reports that instead of hanging CI

(A third test passed *under* the falsifier and was deleted rather than
shipped as false assurance.)

## The audit — the actual deliverable

| Instrument | State before this PR |
|---|---|
| `backend_ab.rs` `AB_COVERED` | `AttentionTranspose` **absent**, while
its `PLAN` entry claims `Graduation::Partial` — and the graduation rule
*reads* `AB_COVERED` |
| `tests/native_vs_mlas_differential.rs` | no attention row at all |
| `benches/native_vs_mlas.rs` | no attention row at all |
| `identity_hook_specialization_matches_the_general_epilogue` | compared
two **zero-prefilled** buffers |
| `scripts/ort_ab/gen_mha.py` | 7 cells, all bidirectional encoder
shapes; `unidirectional` supported but never set; no `q_seq = 1`; no
past-KV |
| `mha_parity/cases.rs` | 12 goldens; the only past-KV case has `q_seq =
1` |

The family with a known reference-route defect was the family with no
same-binary A/B.

**Production-route coverage was already sound** — `mha_ort_parity.rs`,
`msft_attention_ort_parity.rs` and `qwen35_ort_parity.rs` all run ORT
goldens in the default MLAS-free build. The hole was entirely in the
research/reference route and in the shape grid.

### Zero-prefill cannot see a dropped write

Falsified by skipping 8 rows of the `probs·V` GEMM — 6144 of 98304
elements, #1685's exact signature:

- old zero-prefill assertion → **passes** (an unwritten element is
indistinguishable from a legitimate `0.0`, and a hole landing
identically in both compared runs cancels out)
- new NaN prefill → `left 6144 of 98304 output elements unwritten; first
at index 7680 = (tile 0, row 120, column 0)`

## Two real defects the new coverage found

### 1. `concat_cache` walked a row-major buffer column-major

```rust
for d in 0..dim {            // head dim OUTSIDE
    for j in 0..past.seq {   // sequence INSIDE
        data[((b * heads + h) * total + j) * dim + d] = past.at(b, h, j, d);
```

`Bnsh` is contiguous `[b][h][s][d]`, so consecutive stores were `dim`
floats — **512 B** at Llama's head size — apart: every store touched a
fresh cache line and the tensor was traversed `dim` times. ~4.2 M
near-certain misses per tensor, **~20 ms of a ~28 ms decode node**. The
operator spent most of a decode step copying its own cache. The two
sibling transforms document this exact rule in their doc comments; this
was the one place that broke it, and no benchmark row supplied a past-KV
cache, so nothing measured it.

Now two `copy_from_slice` calls per `(b, h)` plane, fanned out on the
same `MIN_PARALLEL_TRANSPOSE_ELEMENTS` threshold:

| cell (t=8) | before | after |
|---|---|---|
| `llama_decode_past1023` | 27.2× | **13.8×** |
| `llama_chunk8_past1016` | 24.8× | **13.1×** |
| `llama_chunk32_past992` | 27.7× | **18.5×** |

Parity preserved on all three.

### 2. Nothing checked the causal offset

`past_seq` is load-bearing only when `q_seq > 1` **and** the cache is
non-empty. Every self-attention golden has `past == 0`; the one past-KV
golden has `q_seq == 1`, so `causal = unidirectional && q_seq > 1` is
**false**.

Hard-coding `past_seq = 0` leaves **all 12 pre-existing goldens
passing**. The new `past_kv_chunked_prefill_causal` is the only case
that fails (`max abs diff 2.22`). Our convention was already correct; it
is now pinned. Regenerating with ORT 1.26.0 was byte-stable for the
existing 12 cases (+19 lines only).

## An undefined benchmark cell is not a slow one

Adding causal/decode rows produced 24/384 parity failures on `q_seq=8,
kv_seq=1024, unidirectional=1`, no past-KV. That looked like a kernel
bug and was not one:

| reference | vs ORT |
|---|---|
| `unidirectional = 0`, no mask | **1.2e-7** ✅ |
| causal, offset swept over **every** value in `0..=kv_seq` | matches at
**none** |
| causal at offset `past_seq`, with a real past-KV cache | **2.4e-7** ✅
|

ORT's `unidirectional` is simply undefined when `q_seq != kv_seq` with
no past input — neither runtime computes a defined answer, so no timing
from that cell is meaningful. It is replaced by past-KV cells
(`llama_decode_past1023`, `llama_chunk8_past1016`,
`llama_chunk32_past992`), which is how the runtime actually emits
chunked prefill, and `build_mha` now **raises** on the invalid
combination.

## Complete results — production default build, 18 cells × 4 thread
counts

Default MLAS-free `bench_generic`: `nm`, `nm -D`, `strings`, `ldd` →
**0** MLAS symbols, no `libstdc++`. **Positive control:** the same
probes on a `--features mlas` build report **842** symbols / 105 strings
/ `libstdc++` linked, so the probe demonstrably sees MLAS when present.
`MultiHeadAttention` executes natively at 99.97% of node time, 1 call,
no ORT fallback. **432/432 trials parity PASS.**

`native/ort` p50, lower is better; `(…)` = same-invocation **A/A null
control**:

| cell | t=1 | t=4 | t=8 | t=16 | native ms t=1 → t=16 |
|---|---|---|---|---|---|
| `bert_base_b8_s128` | 5.36 (5.31) | 13.01 (8.97) | 15.50 (14.46) |
19.93 (20.97) | 38.2 → 42.3 |
| `bert_base_decode_kv1024` | 1.51 (1.54) | 0.78 (0.76) | **0.55**
(0.46) | 0.62 (0.58) | 1.2 → 1.3 |
| `bert_base_s128` | 5.71 (5.73) | 9.99 (8.56) | 11.63 (10.64) | 11.65
(12.29) | 4.8 → 7.8 |
| `bert_base_s384` | 7.04 (7.14) | 22.89 (15.69) | 23.36 (23.78) | 28.18
(32.36) | 52.7 → 59.7 |
| `bert_large_s128` | 5.65 (5.67) | 11.29 (8.89) | 12.75 (14.01) | 15.75
(15.92) | 6.3 → 11.0 |
| `clip_l14_s257` | 5.79 (5.74) | 14.00 (13.25) | 24.49 (25.95) | 24.83
(24.52) | 25.0 → 38.1 |
| `llama_chunk32_past992` | 4.11 (4.18) | 11.69 (12.31) | 17.93 (17.64)
| 24.29 (23.75) | 42.4 → 42.9 |
| `llama_chunk8_past1016` | 3.41 (3.59) | 8.10 (7.59) | 11.77 (11.99) |
17.16 (17.15) | 17.6 → 21.9 |
| `llama_decode_b8_kv1024` | 1.50 (1.50) | 1.44 (1.45) | 1.55 (1.53) |
1.53 (1.55) | 79.9 → 51.4 |
| `llama_decode_kv1024` | 1.93 (1.88) | 1.14 (1.46) | 1.33 (1.22) | 1.24
(1.68) | 13.1 → 8.8 |
| `llama_decode_kv128` | 1.50 (1.49) | 0.69 (0.75) | **0.63** (0.65) |
0.72 (0.63) | 0.6 → 0.8 |
| `llama_decode_kv4096` | 1.59 (1.59) | 1.32 (1.35) | 1.40 (1.39) | 1.40
(1.68) | 42.2 → 31.8 |
| `llama_decode_past1023` | 3.48 (3.49) | 8.21 (8.00) | 11.00 (9.65) |
15.03 (13.60) | 10.5 → 13.8 |
| `llama_prefill_s128_causal` | 3.31 (3.35) | 5.67 (6.86) | 6.74 (7.67)
| 9.96 (8.04) | 15.5 → 22.5 |
| `llama_prefill_s512_causal` | 3.75 (3.73) | 12.34 (12.49) | 19.34
(19.30) | 21.86 (21.60) | 237.8 → 232.5 |
| `phi35_prefill_s256_causal` | 4.02 (3.99) | 11.11 (10.67) | 15.20
(15.04) | 18.18 (17.25) | 51.8 → 52.8 |
| `vit_b16_s197` | 5.64 (5.66) | 11.60 (11.33) | 11.37 (11.54) | 12.09
(16.70) | 10.9 → 11.3 |
| `whisper_cross_s1500` | 6.16 (6.16) | 20.42 (21.55) | 31.52 (31.68) |
40.29 (40.46) | 182.6 → 181.3 |

Regressions are included, not filtered. The null arm tracks the native
arm within a few percent on every cell, so these are the measurement and
not the instrument.

### What the table actually says

Read the last column. **`native ms` is flat from 1 → 16 threads on every
cell.** The ratio degrades with thread count purely because ORT scales
and we do not (`llama_decode_past1023`: ORT 3.05 ms → 1.06 ms across the
same sweep).

The code agrees: `sdpa_f32_simd` — the route a **default** build takes
on x86 and aarch64 — is a plain `for b { for n { … } }` with **no rayon
fan-out at all**. `sdpa_f32_fast`, the MLAS *research* route, does use
`par_chunks_mut`. The shipped route is the serial one.

At `t = 1` the grid is 1.5–7.0×, which is a per-core efficiency gap.
Everything above that is unclaimed parallelism worth roughly the thread
count. The cells that already win (`bert_base_decode_kv1024` 0.55×,
`llama_decode_kv128` 0.63×) are the ones small enough that one core
suffices.

**This is filed separately rather than folded into a coverage PR** — it
is a large change and it touches pool ownership, which is Sebastian's
lane. Now tracked as **#1718**.

## Gates

| gate | result |
|---|---|
| `cargo fmt --all --check` | clean across all 11 files in this diff
(main carries 3 pre-existing unformatted files in `onnx-genai-engine` /
`onnx-genai-metadata`, untouched here) |
| `clippy --locked -p onnx-runtime-ep-cpu -p mlas-sys --all-targets` |
clean |
| `clippy --no-default-features --features mlas --all-targets` | clean |
| `clippy --target aarch64-unknown-linux-gnu --all-targets -- -D
warnings` | clean (caught a real `needless_return` in my new
`sdpa_f32_native` cfg ladder) |
| `scripts/check_cross_compile.sh` | ✓ full offline set |
| `cargo test -p onnx-runtime-ep-cpu` (default) | 1604 passed, 0 failed
|
| `cargo test -p onnx-runtime-ep-cpu --features mlas` | 1552 passed, 0
failed |
| `cargo test -p mlas-sys` | 41 passed, 0 failed |
| Miri `task_runtime` / `strided` / `provider` / `dtype` | 29 / 8 / 16 /
12 passed, no UB |
| `default_artifacts_are_mlas_free` | 9 passed |
| default cdylib MLAS symbols | `nm` 0, `nm -D` 0, `strings` 0, `ldd` no
`libstdc++` — positive control: the same probes on a `--features mlas`
build see **842** symbols, so the probe demonstrably detects MLAS when
present |

Note: `onnx-runtime-ep-cuda` has a **pre-existing** `approximate value
of PI` clippy error (a CUDA C++ kernel string in `kernels/window.rs`).
This diff touches 0 files in that crate.


## Review

Reviewed by Opus (`claude-opus-4.8`): **no blocking issues**, with
hand-verified index algebra for the `concat_cache` rewrite and
confirmation that the new tests are non-vacuous.

Two non-blocking robustness notes were raised and both are now addressed
in `9ae8f7b26`:

- the `peak > 1` overlap guard could misfire on a single-vCPU runner
where workers may serialise; it is now gated on `available_parallelism()
> 1`, so it still fails loudly wherever overlap is possible;
- `concurrent_sdpa_sessions_lose_no_work` had no watchdog, so a
reintroduced pool deadlock would have hung the suite rather than failing
it. It now runs on a 120 s watchdog thread with the same diagnosis
message as `concurrent_dispatch`.

Neither touches production code. All gates re-run green after both
changes and after the rebase onto `d3688e7e0`.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…d a knob that no longer exists (#1822)

Follow-up to #1173, correcting two defects I shipped in it and repairing
the rule they undermined. Docs, one ledger string, one new test, one new
script. No production kernel or routing change.

## 1. The ledger named a route gate that had already been deleted

`PLAN[MatMulF32].shape_gate` said the native `SimdX86` route "gates M=1
on `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, #1116)".

#1183 shipped that GEMV on by default and removed the probe. `git
merge-base --is-ancestor 5417d04 bdb4599` confirms it landed
**before** #1173 merged — so the ledger was wrong the day it landed.
Today `sgemm_simd` calls `sgemm_simd_variant(a, b, c, m, k, n, true)`
unconditionally and `use_m1_gemv` is a plain parameter that only the
in-process A/B harness passes as `false`. No environment variable
reaches that route.

`docs/performance/CPU_MATMUL_ASSIGNMENT.md:559` already recorded the
correct fact ("It is measured now, and the route is the default. There
is no env probe on the dispatch any more"). Two files in the same
directory disagreed and nothing compared them.

**Now guarded.**
`ledger_prose_only_names_environment_variables_that_still_exist`
requires every `NXRT_*` / `ONNX_GENAI_*` token in the ledger's prose to
still exist as a string literal in the crate's sources. It cannot check
that the description is *right*, only that the knob is *real* — which is
the half that goes stale silently.

Mutation-verified, not just observed green:

```
matmul_f32: ledger prose names environment variable `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`,
but no source file in this crate contains the literal "ONNX_GENAI_CPU_MM_SIMD_M1_GEMV".
```

## 2. The doc published a toggle A/B that could not have been run

#1173 carried a table captioned **"same binary, same session, toggle the
only difference"**, reporting `decode 1×2048×2048` at 0.146 with
`ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` off against 0.337 with it on, and
called turning it on "the obvious next slice".

Nothing reads that variable. Setting it measures the same route twice;
it cannot produce two different columns. The table is withdrawn and the
retraction kept in the text rather than quietly deleted.

This is the failure mode the document's own graduation rule warns about
— **an arm that was not on the route it was labelled with** — committed
by the document that wrote the rule. It survived review because a
plausible number in a well-formed table is not self-evidently
unmeasured. Readers are pointed at `bench_f32_gemm_ab`, which holds the
route as a function parameter and carries the M≥2 rows as a built-in
control.

## 3. The gap table is re-measured and the ≥5% rule is repaired

The old table was one unguarded invocation per row at an unstated width,
taken before the decode-placement corrections (#1729, #1794, #1811) —
i.e. when the decode pool put 16 workers on 8 physical cores.

New harness: `scripts/bench_native_vs_mlas_width.py`. Arms interleaved
rep by rep so host drift lands on both equally; per-rep `os.wait4`
CPU-efficiency guard adapted from #1809; six reps per arm; two widths.
Raw verdicts, spreads and discards are all reported rather than
summarised away.

**Three findings, all about method rather than kernels.**

| | narrow (6 cores, 1 L3) | wide (32 logical CPUs) |
|---|---|---|
| `matmul_f32 16×512×512` | 1.581, spread 41% | 0.866, spread 134% |
| `matmul_f32 decode 1×2048×2048` | 1.117, spread 21% | 0.934, spread
13% |

- **Two cases change verdict on width alone.** Same binary, same
half-hour, only the CPU mask differs. `x86_sgemm` parallelises over
column strips and MLAS declines to parallelise some shapes, so
interleaving the two *routes* inside one process does not protect the
ratio — it changes both at once.
- **`16×512×512` disagrees with itself on both arms**, alternating
`keep-mlas` / `native-graduates` from a byte-identical binary. **One
more run of the old table could have graduated a route on this row.**
- **The narrow arm is more trustworthy despite having fewer cores** —
spreads 4–42% against 5–134%, and it lost no reps to the guard.
Isolation beat parallelism.

**Softmax now decomposes cleanly**, because no vendored MLAS kernel has
changed since #1173 (the only `mlas-sys` edits are the additive
straggler handshake in `work_stealing_pool.rs`, #828/#1714, which adds
waiting). At matched width the MLAS control arm is stationary to within
4% while native improved **1.24–1.27×** — matching #1416's claim for the
row kernel. The f32 GEMM rows get no such attribution and now say so
explicitly: their control moved **2.0× the wrong way**, so only the
current ratio at a stated width is defensible.

**The rule gains what it lacked**: spread must be smaller than the
claimed win; reps that did not get the CPU are discarded rather than
averaged; a verdict is valid only at a stated width. Under it, `decode
1×2048×2048` — the first f32 GEMM case to show a real native win —
**still does not graduate**: it costs more CPU (cpu_ratio 0.875), does
not hold at 32 threads, and its 21% spread exceeds its 12% win.

## The width claim is verified, not asserted

#1815 landed while this was in progress and observed the neighbouring
`bench_generic` harness spawning its ORT arm *outside* the affinity
confinement it applied to the native arm. That hazard applies to any
`taskset` claim, including mine, so I checked it instead of trusting it
— sampling `Cpus_allowed_list` from `/proc/<pid>/task/*/status` 40×
across a live narrow-arm run:

```
'16,20,22,26,28,30': 478 observations
  native_vs_mlas- 273, mlas-sys-ws-0..4 39 each, nxrt-task-0..4 2 each
'0-31': 1  (the taskset process itself, before exec)
```

Both routes confined identically; no thread escaped. The rule now
requires this check.

## Validation

- `dispatch_ledger` **17/17**, including the new falsifier, after
merging latest `main`.
- `default_artifacts_are_mlas_free` **9/9** — the no-MLAS-in-defaults
invariant is untouched.
- `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` clean;
`cargo fmt --check` clean.
- Normal merge of `origin/main` (`aee2b9d11`), no rebase, no conflicts.

## Limitations

- The narrow arm is six cores on one L3 of one x86-64 host. Nothing here
transfers to aarch64 or to a two-socket box.
- The `activations erf 1 Mi` row shows native 13.5% slower at matched
width. The nearest scatter figure is the wide arm's 8% spread, but that
is a spread of *ratios* against a move in a *native time*, so the two
are not strictly commensurable. Its MLAS control also moved 11%.
**Flagged for pinned re-measurement, not reported as a regression.**
- The wide arm was taken with ~4–5 cores of unrelated load present. That
is stated in the doc rather than hidden, and it is why its spreads are
wider; the guard reports which reps were discarded instead of pretending
the host was quiet.
- No production behaviour changes here, so there is no performance claim
to make about the shipped artifact.

Refs #1173, #1183, #1809, #1815, #1416.



## Independent review, and what it changed

An independent adversarial review of the full diff returned **no
blockers** — it confirmed the ancestry argument behind the retraction,
the stationary-control premise for the softmax attribution, and that the
headline case is correctly *refused* by the rule (21% spread against a
12% win). It also found seven real defects, all now fixed in
`f0323f9ed`.

The one that mattered most was in the new test. It only proved the
variable name appeared *somewhere* in the crate, so a variable whose
read site had been deleted but whose name survived in an
`EnvVarGuard::set(...)` line would still have passed — which is the
precise shape of the defect this PR exists to correct. The test now
requires the matching line to be an `env::var(` / `env::var_os(` read or
an `_ENV: &str =` binding.

Verified by mutation in **both** directions:

| mutation | before | after |
|---|---|---|
| reinsert retired `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` into ledger prose |
fails ✅ | fails ✅ |
| retire the two real `NXRT_CPU_GEMM_BACKEND` reads, leaving the literal
only in test guards | **passes ❌** | fails ✅ |

The remaining six were prose defects in the doc: a stated spread range
that contradicted its own table's 82% row, "within 4%" against a table
reading −4.2%, a narrow-arm ratio fused with a wide-arm attribution, a
spread quoted as 7.5% that was 8% *and* compared against an
incommensurable quantity, the CPU-efficiency guard oversold as "what
makes this table measurable at all" (in-process interleaving is what
protects the ratio; the guard catches only *differential* descheduling),
and a one-directional provenance argument standing in for the direct
control measurement that actually carries the softmax attribution.

**Two further defects I found myself while checking the tables against
each other**, neither raised by the review:

- The `ratio` column is a median of per-rep ratios while the `ns/unit`
columns are medians of times. Medians do not distribute over division,
so every row looked internally inconsistent to anyone who tried to
divide it out (`0.0684 / 0.0617 = 1.109` against a stated `1.117`). Now
documented, along with why the per-rep form is the correct one to quote:
it pairs each MLAS invocation with the native invocation it was
interleaved against, which is the entire point of interleaving. The
then→now figures are relabelled as quotients of medians.
- "wider than nine of the twelve wide-arm rows" was eleven of twelve.

## Adopting #1814

`aee2b9d11` (#1814) landed on `main` while this was in review, and it
closes the exact hole the review found in the guard this document
recommends. A differential CPU-efficiency check cannot see contention
that lands evenly on both arms; #1814's confined-set meter reads busy
jiffies on the process's own `Cpus_allowed_list` and subtracts the
process's own CPU, so foreign load shows up directly. The rule now
points at it, and the tables here are explicitly marked as predating it
and guarded by the weaker method.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the
file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the
row's siblings give either a full path from the repository root or a bare file
name. This one is neither, so a reader following it finds nothing and cannot
tell whether the test was moved, renamed, or never existed.

Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing
it anyway because the citation checker this branch runs flags it, and claiming
"every citation in both documents resolves" while knowingly leaving one that
does not would make the claim worth nothing. The audit that found it was
prompted by the same class of defect found in a PR body upstream of this
branch: published prose that nothing type-checks.

Docs only; no code, schema, or fixture changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the
file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the
row's siblings give either a full path from the repository root or a bare file
name. This one is neither, so a reader following it finds nothing and cannot
tell whether the test was moved, renamed, or never existed.

Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing
it anyway because the citation checker this branch runs flags it, and claiming
"every citation in both documents resolves" while knowingly leaving one that
does not would make the claim worth nothing. The audit that found it was
prompted by the same class of defect found in a PR body upstream of this
branch: published prose that nothing type-checks.

Docs only; no code, schema, or fixture changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the
file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the
row's siblings give either a full path from the repository root or a bare file
name. This one is neither, so a reader following it finds nothing and cannot
tell whether the test was moved, renamed, or never existed.

Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing
it anyway because the citation checker this branch runs flags it, and claiming
"every citation in both documents resolves" while knowingly leaving one that
does not would make the claim worth nothing. The audit that found it was
prompted by the same class of defect found in a PR body upstream of this
branch: published prose that nothing type-checks.

Docs only; no code, schema, or fixture changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the
file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the
row's siblings give either a full path from the repository root or a bare file
name. This one is neither, so a reader following it finds nothing and cannot
tell whether the test was moved, renamed, or never existed.

Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing
it anyway because the citation checker this branch runs flags it, and claiming
"every citation in both documents resolves" while knowingly leaving one that
does not would make the claim worth nothing. The audit that found it was
prompted by the same class of defect found in a PR body upstream of this
branch: published prose that nothing type-checks.

Docs only; no code, schema, or fixture changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the
file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the
row's siblings give either a full path from the repository root or a bare file
name. This one is neither, so a reader following it finds nothing and cannot
tell whether the test was moved, renamed, or never existed.

Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing
it anyway because the citation checker this branch runs flags it, and claiming
"every citation in both documents resolves" while knowingly leaving one that
does not would make the claim worth nothing. The audit that found it was
prompted by the same class of defect found in a PR body upstream of this
branch: published prose that nothing type-checks.

Docs only; no code, schema, or fixture changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 25, 2026
`tests/adapter_artifact_compat.rs` resolves from no directory in the tree: the
file is `crates/onnx-genai-metadata/tests/adapter_artifact_compat.rs`, and the
row's siblings give either a full path from the repository root or a bare file
name. This one is neither, so a reader following it finds nothing and cannot
tell whether the test was moved, renamed, or never existed.

Pre-existing, from the redesign in #828, and outside this PR's subject. Fixing
it anyway because the citation checker this branch runs flags it, and claiming
"every citation in both documents resolves" while knowingly leaving one that
does not would make the claim worth nothing. The audit that found it was
prompted by the same class of defect found in a PR body upstream of this
branch: published prose that nothing type-checks.

Docs only; no code, schema, or fixture changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

human Assigned to a human

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants