Skip to content

feat(engine): one interpreter executes every package's declared workflow - #1723

Merged
justinchuby merged 64 commits into
mainfrom
feat/native-workflow-backend
Aug 23, 2026
Merged

justinchuby merged 64 commits into
mainfrom
feat/native-workflow-backend

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 22, 2026 •

Copy link
Copy Markdown
Owner

One runtime, one interpreter, one drive

Every package — a bare decoder, a composite pipeline, a speculative block —
serializes pipeline.workflow, and every generated token now comes out of
run_workflow_node walking that document. The iteration bound, the liveness
predicate, the carried state, the emit that publishes the token stream and the
stop are read from the package rather than supplied by a Rust loop written
beside it.

What varies between packages is which executor implements a declared node, and
that is answered by the contract the node names.
A component declaring
onnx-genai.autoregressive-decode is run by the fused decode session — paged KV,
the device sampling fast paths, the captured CUDA graph — when this process
holds one, and invoked from its own artifact when it does not.
holds_decode_core() asks whether that session exists; nothing asks what kind
of package was loaded.

Measured, not asserted

authored_body_selects_executor.rs drives three packages declaring three
different bodies through one entry point and reads a per-contract execution
counter recorded inside the interpreter's node dispatch:

body contracts executed components invoked
tiny-llm (runtime decode + policy) 5 × decode, 5 × policy for 5 tokens —
decoder fixture (in-graph sampler) none model, plus a fused policy island
speculative fixture none proposer/verifier island, grammar adapters

The zero matters more than the counts: a package that declares a graph for every
step must never have the runtime supply one, or the declared document is
decorative.

Deleted

pipeline/canonical_decode.rs (771 lines), Engine::lowered_workflow,
canonical_workflow() and its nine guard call sites, WorkflowShape /
declared_workflow_shape, is_workflow(), workflow_shape(),
WorkflowShapeReport, from_pipeline_dir*,
generate_pipeline_with_callback(s), workflow_generate,
forget_canonical_workflow_for_test, the server's start_pipeline /
run_pipeline_driver / run_pipeline_generation / GeneratePipeline, and the
CLI's Backend::{Text,Pipeline} enum.

The shape-dispatch gate's allowance went from 23 files to zero (the gate
itself aside). Those symbols do not exist anywhere in the tree.

Three things the convergence surfaced

  • The decoder component declared no contract. The emitter had gained
    AUTOREGRESSIVE_DECODE_CONTRACT, but the fourteen converted fixtures predated
    it — nothing in the tree declared the step it was executing. Re-emitted through
    a new migrate_model_io --reemit, which reads the ABI back out of the workflow
    the package already declares.
  • The sequence cell's recurrence was wrong. It said invariant while
    carrying the prompt on the first iteration and one token thereafter. Nothing
    could honour that, and nothing had to, because no interpreter was walking the
    loop. It now declares bounded.
  • The liveness read was skipped for host-written predicates. The
    max_iterations_only optimization is sound for an in-graph EOS predicate a
    caller disabled; a predicate a runtime executor writes carries the context
    bound and stop strings too. Without this the context-limit stop was ignored and
    generation ran into a model-window overflow.

Real finish reasons, and a load-bearing emit

The composite drive no longer fabricates MaxTokens, prefix_cache_hit_len: 0
and logprobs: None. It reports the core's stop when a core reached one, the
interpreter's own record of whether the predicate ended the loop otherwise, and
the prefix/logprob facts of whoever scored the tokens — None for a package
whose sampler is an ONNX component, which is a fact about the package rather
than a missing feature. The emitted stream and the committed one are compared,
and a mismatch refuses.

Sessions

A session is the conversation a client is having; where the runtime keeps it is
not something a client can act on. /v1/completions and /v1/sessions no longer
reject workflow packages, and create_session / session_token_count /
close_session mean the same thing for every package.

Real-model evidence

.goldens/REAL_MODEL_EVIDENCE.md — one probe, two builds (clean
origin/main@cb81745b0 worktree versus this branch), same machine, pinned
package revisions.

  • qwen3-0.6b-onnx-genai@a89c02d3 decodes byte-identically on both, through
    completely different execution paths: main ran it on the decode core from its
    genai_config.json, head runs the eleven-component workflow it declares.
  • gemma4-e2b@b1189a9b and gemma4-e2b-speculative@3e5d3b6a cannot run on
    main at all
    (one fails before generating, one fails to load). Both load on
    head.
  • Two behaviour changes are reported, not summarised away: seeded sampling
    differs (head uses the sampler the package declares), and multi-turn
    continuation is lost for qwen3_0_6b
    because its declared workflow has no
    session-scoped state. That is called out as a regression with both fixes named.
  • A gap is stated: both Gemma4 packages require serving inputs a bare prompt
    cannot seed, so composite speculative widths 1/2/4 against those packages
    have no measurement. The hermetic gemma4_chained CUDA gate is not substituted
    for it.

Two defects the pinned packages found are fixed here: generate_in_session
failed on a package with no decode core because only one wrapper had been routed
to the interpreted session path, and a text prompt against a package declaring a
prompt_tokens input was refused even when the package shipped tokenizer.json.

CUDA residency — partial, and measured

WorkflowRuntime counts every device→host materialization, and the H200 chained
parity gate reads it across a propose/verify/reject/rollback decode.

Landed: borrowed proposer bindings are narrowed in place (a pointer view via
the new Value::alias_with_offset) instead of being copied down in full and
narrowed after; the per-draft-token argmax reads four bytes off the device
instead of a ≈300 KiB logits row; the embedding table is read once rather than
per proposal. tensor_ops_for is fail-closed — a device value with no device
implementation is an error naming both remedies, never a quiet copy.

Not landed, and the assertion says so: the count is a ratchet at 16, not a
zero. The folded carry seed and the folded carry itself still cross the bus,
because the fused proposer input is assembled on the host. Removing them needs a
device-side scatter this seam does not have, and faking it with a staging copy
would satisfy the counter and change nothing.

Validation

465 engine lib tests, every engine integration test, 241 server tests, metadata /
CLI / C API suites, cargo fmt, clippy -D warnings on every crate touched, and
all 12 native_workflow_parity cases on an H200 including both chained
speculative gates. origin/main was merged (never rebased); the zero-allowlist
IR11/opset24 fixture guard passes.

Known pre-existing failure, verified on origin/main@cb81745b0 in a clean
worktree and unrelated to this branch:
onnx-genai-ort::loader::tests::selected_non_dense_candidate_fails_explicitly.

justinchuby and others added 2 commits August 22, 2026 04:12
Make `EngineDecodeBackend::Native` a first-class backend for the one
universal `pipeline.workflow` interpreter instead of another parallel
native workflow. The interpreter (SSA, loops, branches, emits,
loop-carried and session state, adapters, checkpoints) is reused
unchanged; only *component execution* is now backend-neutral.

What changed
- Add `PipelineEngine::invoke_onnx_component`, a backend-neutral seam the
  node interpreter calls for every ONNX component. The ORT branch is
  behavior-preserving (it keeps the `Session::run` and the CUDA-graph /
  `IoBinding` stable-island fast path); the native branch runs a pure-Rust
  `InferenceSession`.
- Add `pipeline::native_component` (feature `native-backend`): one native
  session per declared component, loaded once and reused across recurring
  loop edges, bridging the interpreter's `onnx_genai_ort::Value` currency
  to/from native `Tensor` with a faithful, fail-closed dtype mapping. No
  tensor is routed through the host-only `ComponentTensor` seam.
- Flip `validate_pipeline_backend_request` to accept Native when the
  native backend is compiled in, and fail closed (no silent ORT fallback)
  otherwise. Execution islands (an ORT CUDA-graph optimization) are not
  planned under Native, so the interpreter drives individual component
  nodes correctly but unfused.
- Remove the redundant host-only `ComponentSession`/`ComponentTensor`
  (onnx-genai-metadata) and `OrtComponentSession` (onnx-genai-ort). They
  had no callers and duplicated the component-invocation concept this seam
  now owns device-resident (Rule 3, Rule 10).

Why
- `Value` is already a device-capable handle, so keeping it as the neutral
  currency lets native tensors and loop-carried/shared state stay
  backend/device-resident without a host round-trip, and reuses the whole
  interpreter rather than cloning workflow logic.

Autoregressive decode runs through the one interpreter loop (the decoder
Loop over policy components), reusing the single sampling/stopping/commit
policy — no third token loop is added.

Tests (feature `native-backend`)
- `native_workflow_parity`: same package run under ORT and Native with
  identical outputs for single-pass, nested loop+branch, loop-carried
  state, read-only shared session state, diffusion loop, and autoregressive
  decode; a native-used counter proves native sessions ran (not an ORT
  fallback); a routing test proves recurring loop edges reuse resident
  native sessions; a fail-closed test proves an unsupported op errors with
  an actionable diagnostic instead of silently falling back to ORT.

Design note and staged follow-up boundaries (CUDA zero-copy residency,
specialized AR executor, direct-Engine facade, Gemma4 target+assistant)
are in docs/architecture/NATIVE_WORKFLOW_BACKEND.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
@codecov

codecov Bot commented Aug 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 48.97959% with 100 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.62%. Comparing base (b9217fa) to head (14da4e3).
⚠️ Report is 4 commits behind head on main.

Files with missing lines Patch % Lines
...rates/onnx-genai-genai-config/src/compatibility.rs 51.56% 28 Missing and 3 partials ⚠️
...s/onnx-genai-metadata/src/bin/validate_metadata.rs 31.70% 24 Missing and 4 partials ⚠️
crates/onnx-genai-cli/src/interactive.rs 47.05% 24 Missing and 3 partials ⚠️
crates/onnx-genai-cli/src/model_inspection.rs 0.00% 10 Missing ⚠️
crates/onnx-genai-cli/src/transcribe.rs 0.00% 4 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1723      +/-   ##
==========================================
+ Coverage   80.21%   80.62%   +0.41%     
==========================================
  Files         413      413              
  Lines      202063   203061     +998     
  Branches   202063   203061     +998     
==========================================
+ Hits       162091   163728    +1637     
+ Misses      34435    33789     -646     
- Partials     5537     5544       +7     
Flag Coverage Δ
cli-ort-linux 72.51% <36.92%> (+0.03%) ⬆️
cli-ort-windows 72.01% <36.92%> (+0.03%) ⬆️
mlas 85.33% <ø> (+0.13%) ⬆️
offline 80.77% <54.96%> (+0.42%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-genai-genai-config/src/import.rs 91.97% <100.00%> (+0.02%) ⬆️
crates/onnx-genai-metadata/src/decoder_abi.rs 91.15% <100.00%> (+8.37%) ⬆️
crates/onnx-genai-metadata/src/decoder_workflow.rs 96.71% <ø> (ø)
crates/onnx-genai-metadata/src/lib.rs 95.23% <ø> (ø)
crates/onnx-genai-metadata/src/parser.rs 58.60% <ø> (-10.35%) ⬇️
...ates/onnx-genai-metadata/src/schema/decoder_abi.rs 96.29% <ø> (ø)
...rates/onnx-genai-metadata/src/schema/generation.rs 46.66% <ø> (-20.00%) ⬇️
crates/onnx-genai-metadata/src/schema/ir.rs 97.56% <ø> (ø)
crates/onnx-genai-metadata/src/schema/mod.rs 81.39% <ø> (-8.74%) ⬇️
crates/onnx-genai-metadata/src/validation.rs 69.87% <ø> (-0.05%) ⬇️
... and 6 more

... and 20 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

justinchuby and others added 27 commits August 22, 2026 04:43
… plan

Phase 0 of collapsing the two execution engines into the one canonical
`pipeline.workflow` runtime (see the new
docs/architecture/WORKFLOW_RUNTIME_UNIFICATION.md).

The `Engine` session/decode model is already unified across ORT and native:
ORT vs native is a `DecodeLoopBackend` choice beneath the single
`run_decode_loop` policy, and the `*_native_*` session methods were already
deprecated shims that delegate to the backend-dispatched
`create_session`/`close_session`/`generate_in_session`/`rewind_session_by`.

Per Rule 3 (no compatibility shims for pre-release APIs), delete them:
- `Engine::create_native_session` -> `create_session`
- `Engine::close_native_session` -> `close_session`
- `Engine::generate_native_in_session[_with_callback]` -> `generate_in_session[_with_callback]`
- `Engine::rewind_native_session` -> `rewind_session_by`

The internal `generate_native_in_session_with_callbacks` (the real native
executor beneath the unified session path) is retained. The only external
caller, `onnx-genai-bench`'s `multiturn`, is migrated to the unified API.

Proven behavior-preserving: `generate_in_session_with_priority_and_callback`
already dispatches to the native executor when `decode_backend == Native`;
the native session tests (`native_engine`, `multi_session`) already use the
unified API and pass, and the native workflow parity/smoke tests still pass.

Also documents the full sole-runtime plan (one interpreter, one state/session
model, one decode policy, ORT/native executors beneath) with a phased,
independently-testable deletion order, and normalizes fmt drift inherited from
the merged main under the pinned toolchain.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Speculative decode expressed as a workflow (draft → verify → accept/reject →
correction) produces identical output on ORT and native. The accept/reject/
rollback semantics live in the backend-agnostic interpreter; the native
executor only runs component forward passes and never re-derives speculative or
state-transition semantics. This is the single-parity-case pattern the Gemma4
target+assistant speculative workflow (onnx-genai #1716/#1696) will rely on.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…ase 1

Record the concrete Gemma4-driven requirement from onnx-genai #1696/#1716: the
chained speculative proposer (folded_carry_output or recurrent) must be lifted
from the direct-Engine Rust propose() (speculative/mod.rs:1273-1356) into a
backend-agnostic interpreter construct keyed by SpeculativeProposalExecution::
Chained, so examples 22 (recurrent) and 24 (folded, cacheless drafter) become
single ORT/native parity cases with only the per-step forward pass per-backend.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…oint

Sibling (#1716) confirmed REPLACE, single implementation: no dependency on the
old SpeculatorConfig / SharedKvProposer::propose path for examples 22/24. Record
the two cross-PR Phase-1 preconditions: parent ratifies "replace + delete the
old proposer path" and assigns single interpreter-seam ownership; and the
pre-existing gates gemma4_assistant_full / chained_proposer_real are re-pointed
at the interpreter Chained construct so the deletion drops no coverage.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…dependency gate

Parent ratified REPLACE + DELETE (no facade/no delegator) and assigned the full
consolidation to this PR. Record the measured execution gate: Phases 1-4 are
blocked on #1716 landing on main (a ~50k-insertion, still-evolving branch that
edits the interpreter seam this PR owns) and on an executable chained-proposal
fixture (none exists yet; gemma4-real-packages' package not built). Note that
the AR/chained interpreter node is a real Engine-decode-core integration
(SessionDecodeLoopBackend is Engine-tied), so it must land after #1716 to avoid
redo against its pipeline/* seam edits.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
… + fail-closed bridge

Rubber-duck of #1723 found two substantive, unblocked blockers; fix both now.

Blocker 1 — no ORT sessions under Native. PipelineEngine::build built an ORT
Session for every component even under EngineDecodeBackend::Native, pulling in
the ORT runtime, double-loading weights, misreporting the EP, and letting ORT
reject a native-only graph. Now build components with
PipelineModels::load_with_ort_session_filter(.., |_| false) under Native: zero
ORT sessions, the package's I/O contract stays available as backend-neutral
graph_io_metadata, and execution_provider_status() reports the real native
device. Tests: native_backend_builds_no_ort_sessions (no ORT sessions + native
run + native-cpu EP), ort_backend_builds_ort_sessions (contrast).

Blocker 2 — explicit device/provider + fail-closed residency. native_component
no longer calls InferenceSession::load (auto-detects a CPU EP, ignores the
requested device). It resolves the device via resolve_native_decode_device(
config.native_device, ..) and builds each session with the explicit EP
(CpuExecutionProvider / CudaExecutionProvider), mirroring native_decode/load.rs;
CUDA without the feature and plugin EPs fail closed. The Value<->Tensor bridge
gates on residency: host values/tensors round-trip through bytes (no device
round-trip on CPU); a device-resident value/tensor fails closed with an
actionable diagnostic instead of reading device memory as host bytes
(Tensor::as_bytes is host-only) or silently host-copying. Root-caused: the
native runtime has the device-tensor APIs, so the zero-copy device bridge is a
wiring task (boundary A), not a fundamental CPU-only limit.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…H200-verified)

Adds a native-cuda-gated test proving, on real CUDA hardware: the Native backend
resolves and reports its CUDA device (Blocker 2 device/provider propagation, not
an auto-detected CPU EP), still builds zero ORT sessions (Blocker 1), and at a
device-resident boundary either bridges or fails closed with an actionable
native diagnostic — never panics or silently host-copies. Verified passing on an
H200 (cargo test --features native-cuda). The CUDA code path
(CudaExecutionProvider::initialized + DevicePreference::Gpu) also compiles clean
under --features native-cuda.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Rubber-duck blocker 2 required a real device-resident bridge, not a
fail-closed intermediate. Implement it end-to-end behind `native-cuda`
so a multi-component native CUDA workflow keeps intermediate and
recurring/loop-carried/state tensors on the device with no host
round-trip on recurring edges.

onnx-genai-ort:
- Add guard-carrying `Value::from_external_memory_with_owner` so a
  `Value` OWNS the native `Tensor`/`DeviceIoBinding` behind the device
  pointer it wraps. `TensorBacking::External` gains an
  `Option<Arc<dyn Any + Send + Sync>>` owner, freed by `Value::Drop`
  AFTER the `OrtValue` is released (no leak, no use-after-free).
- `try_alias_clone` now aliases an owned `External` value by sharing the
  owner `Arc`, so a device value flows through the interpreter's
  `clone_value` fast path instead of failing a host read.

engine (pipeline/native_component.rs):
- CUDA path binds device-resident inputs zero-copy via
  `ExternalMemorySpec` + `device_binding_from_external_memory`, binds a
  device output buffer (`allocate_device_output_binding`) for every
  output whose concrete shape the interpreter resolved from bound input
  symbols, runs via `run_with_device_bindings`, and republishes each
  device output as an owning `Value`. Dynamic-shape outputs fall back to
  host materialization. Cross-session stream correctness via
  `Tensor::sync()` / allocator sync. Device/EP stay explicit; genuinely
  unsupported device/provider still fails closed.
- Track `(device_input_bindings, device_outputs)` counters exposed via
  `native_device_residency_counts()` to prove no host round-trip.

engine (pipeline/workflow.rs):
- `resolve_component_output_shapes` resolves output contract shapes from
  the invocation's bound symbols; threaded into the native seam. ORT
  path ignores it.

tests:
- `native_cuda_device_resident_multicomponent` (H200): ORT/native
  bit-for-bit parity on a two-step static-cache decode AND non-zero
  device-residency counters proving the recurring KV stayed on-device.
- Pin the CPU native device in the backend-agnostic parity helpers so
  they are deterministic under `native-cuda`; add `native_cuda_engine`.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Make the smoke run deterministic regardless of build features or a GPU
being present (device-residency is covered in native_workflow_parity),
consistent with the CPU-pinned parity helpers.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…he gate

gemma4-real-packages committed the tiny executable gemma4_chained
target+assistant fixture on the #1716 schema (branch
justinchuby/gemma4-chained-fixture, commit c9c62bc), reusing the proven
tiny-gemma4-assistant graphs. Static contract review (this PR): accepted,
no regen — ABI/state_service/speculative match the agreed chained
contract.

So the second execution-gate condition (an executable chained fixture)
is now satisfied; only "#1716 on main" remains. Record the fixture, the
Phase-1 wire-in plan (required ORT/native + native-CUDA device-resident
parity case in native_workflow_parity.rs), and the one open contract
item flagged to gemma4-metadata-audit: the source-vs-destination meaning
of port_bindings.target_hidden_context for the folded-carry case (carry_0
is the target hidden_states.0 [H], not inputs_embeds [2H]).

No code change; consolidation remains gated on #1716 landing on main.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
gemma4-real-packages confirmed from the fixture graph ground truth:
carry_0 = target.hidden_states.0 (H=16); each proposer step
inputs_embeds[2H] = concat(embed(last_token)[leading H], carry[trailing
H]) (order pinned by the assistant Slice starts=[16] ends=[32]);
input_embedding.f32 = [vocab=32, hidden=16]. steps = single-pass body,
speculative: = loop driver.

Record the coverage boundary: this tiny drafter reads only the carry
half (embed[0:16] is present for the 2H ABI but unused), so the fixture
isolates folded-carry threading + borrowed-KV + rollback and does NOT
validate embed-gather correctness (Phase 1 covers that separately).

Narrow the open items to the two gemma4-metadata-audit field rulings:
(1) direction of port_bindings.target_hidden_context; (2) where the
interpreter sources the embedding table for real models (shared_weights
omitted). Both remain the field owner's call; consolidation stays gated
on #1716 landing on main.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…fields resolved

gemma4-real-packages updated the fixture to the final #1716 schema
(commit 8a66e2c, merging #1716 tip 9f47bb6) and made the folded carry
explicit. Both previously-open contract items are now RESOLVED by
ratified, mandatory proposal_execution fields (verified in
crates/onnx-genai-metadata/src/validation.rs, required whenever
folded_carry_output is set):
- folded_carry_seed {component, output} names carry_0's source
  (fixture: target.hidden_states.0) — supersedes the target_hidden_context
  ambiguity.
- token_embedding {component, table} names the embed-gather source
  (fixture: target.hidden_table) — generalizes to real models, no
  model-name gate.

The Phase-1 chained driver reads both fields directly (no convention
inference). Bump the recorded fixture SHA and point Phase 1 at the
resolved fields. Consolidation stays gated on #1716 landing on main.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…ined contract

gemma4-real-packages published the faithful E2B fix (parent kept the
schema strict: folded_carry_seed names a real target output, no
request-input escape hatch). mobius gemma4 now emits the post-final-norm
hidden as hidden_states.{idx} (PR #546 @ 710d4927, backward-compatible),
and the real E2B packages carry the SAME folded_carry_seed/token_embedding
contract at scale with real ports (target hidden_states.34,
model.embed_tokens.weight [262144,1536] fp16).

This empirically confirms the Phase-1 field-reading chained driver
generalizes from the tiny fixture to real models with no model-name gate.
Record the real-package coordinates as an optional post-parity scale case;
the tiny gemma4_chained @ 8a66e2c stays the required hermetic parity
fixture (unchanged, revalidated). No seam change; still gated on #1716.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…backend

Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>

# Conflicts:
#	crates/onnx-genai-engine/src/pipeline/mod.rs
Tiny target+assistant Gemma4 fixture for native_workflow_parity.rs, driven by
the onnx-genai #1716 `proposal_execution: chained` contract (metadata-driven),
not the legacy SharedKvProposerConfig.

- Reuses the proven tiny-gemma4-assistant graphs (hidden 16, 2H=32, vocab 32,
  kv_heads 2, head_dim 8; target layer 0 = sliding owner, layer 1 = full owner).
- Assistant is a borrowed-KV drafter: inputs_embeds [b,q,2H] + read_only
  shared_kv.{full,sliding}_attention.{key,value} -> logits + projected_state
  (folded carry); no own KV, no present.*.
- Combined inference_metadata.yaml: components {target, assistant}; assistant
  read_only aliases keyed on the owner state cell (no output/no layer);
  proposal_execution {kind: chained, token_embedding_input: inputs_embeds,
  logits_output: logits, folded_carry_output: projected_state};
  port_bindings.target_hidden_context: inputs_embeds; shared_state
  [full_attention, sliding_attention]; vocabulary identical;
  distribution_preserving: true; rollback_state = 4 target KV cells.
- Greedy / logits-emit (external argmax); standard ops only (ORT and native CPU
  EP); FLOAT graphs. No RNG uint64 Cast.

Rust-valid: `cargo run -p onnx-genai-metadata --bin validate_metadata` passes.
Both graphs execute on ORT CPU with the declared contract shapes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add the explicit folded-carry fields introduced by #1716 @ 9f47bb6 to the
chained proposal_execution, using actual graph names (no placeholders):
- folded_carry_seed: {component: target, output: hidden_states.0}
  (the tiny target exposes hidden_states.0 = Gather(hidden_table, input_ids))
- token_embedding: {component: target, table: hidden_table}
target_hidden_context/folded_carry_output retained.

Rust validate_metadata (built from the merged #1716 tip) + python jsonschema
both pass; both graphs still execute on ORT CPU with the declared shapes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…terpreter

A `chained` proposer emits one distribution per invocation, so materializing a
proposal block means driving the same typed component repeatedly while threading
its recurrence forward. That loop lived in a direct-`Engine` Rust `propose()`
bound to ORT `Session`s, which made it a second execution engine the
backend-neutral component seam could not drive.

It now lives in `pipeline::speculative`, on the interpreter's own seam:

* `PipelineEngine::propose_chained` drives the chain through
  `invoke_component_values` -> `invoke_onnx_component`, so ORT and native run the
  identical loop with no second "invoke a component" implementation.
* Everything the loop needs is read from `speculative.proposal_execution`:
  `token_embedding_input`, `logits_output`, `recurrent[]`, `folded_carry_output`,
  `folded_carry_seed` ({component, output}), and `token_embedding`
  ({component, table}). No model name, port spelling, or shape heuristic.
* `accept_chained_proposal` and `rollback_speculative_state` complete the round:
  the accepted prefix plus the target's free correction, and truncation of every
  declared `rollback_state` cell along its serving group's own `sequence_axis`.
  The folded carry owns no state cell and is deliberately never restored.
* The embedding half of the fused input is gathered from the initializer the
  contract names, read out of the target artifact.

Two interpreter facts had to become explicit for this to work at all:

* `run_pipeline_retained` keeps the pass's whole SSA map. A proposal binds the
  proposer's borrowed read-only shared KV and masks from the values the workflow
  itself bound, rather than a caller's reconstruction of them.
* Island fusion elides values no later step reads, which swallowed exactly the
  tensors a proposal needs. `plan_execution_islands` now takes the externally
  used set, computed from the declared contract, so a speculative package keeps
  its folded-carry seed and proposer bindings live (and a package without a
  chained proposer contributes nothing).

Ports sharing the fused input's position symbol narrow to the step's single
position, derived from the declared port contracts, so a borrowed-KV drafter's
`kv_sequence`-keyed mask is left alone while its `sequence`-keyed position ids
are not.

Proven on the hermetic `gemma4_chained` package (imported here from
gemma4-real-packages @ 8a66e2c): contract reading, chain driving, embedding
gather against the package's own second copy of the table, acceptance,
rejection + rollback, width bounding, and a full propose/verify/accept/rollback
decode that reproduces plain greedy decoding token for token.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
… its gates

The interpreter now owns chained proposal driving, so the direct-`Engine`
shared-KV proposer is a second implementation of the same contract. Rule 3: it
is deleted, not kept behind a facade.

Removed: `SpeculativeMode::SharedKv`, `SharedKvProposerConfig`,
`SharedKvBinding`, `validate_shared_kv_proposer_config`, `SharedKvProposer` and
its `propose()`, `SharedKvProposerModel`, `load_shared_kv_proposer`, the paged-KV
slicer `shared_kv_slices_from_materialized`, and the native
`NativeSharedKvProposerModel` / `NativeSpeculationKind::SharedKv` dispatch —
which was already unreachable, since `native_shared_kv_proposer` was hard-wired
to `None` at load and its arm could only ever return "requested without a loaded
proposer session".

Coverage moves rather than disappearing. Both legacy gates asserted the same
property — speculative decoding is token-identical to plain greedy — through the
deleted KV slicer:

* `gemma4_assistant_full` -> `gemma4_chained_workflow::speculative_decode_equals_greedy_decode`
  (plus the ORT/native/CUDA parity case), on the same tiny graphs driven through
  the declared `proposal_execution: chained` contract.
* `gemma4_assistant_mixed` -> `gemma4_chained_workflow::mixed_head_dim_speculative_decode_equals_greedy_decode`,
  on a new hermetic `gemma4_chained_mixed` package built from the same
  heterogeneous graphs (layer 0 sliding head_dim 8, layer 1 full head_dim 16).
  The property is stronger where it now lives: there is no global KV geometry to
  get wrong, because each shared-KV group's ports declare their own head width
  and the interpreter binds them straight through.

Also fixes two merge-fallout parity cases against the #1716 tip: the authored
nested-loop package must declare the `linear_effects` capability it uses, and
the `speculative` fixture now has an asymmetric `verifier.past_key_values.0.value`.

Verified: `gemma4_chained_workflow` (8), `native_workflow_parity` (10 on
native-backend, 12 on native-cuda incl. the H200 device-resident chained parity
case).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
With the interpreter owning chained proposal driving, nothing reaches the
supporting legacy surface, so it goes rather than lingering as a facade:

* `onnx-genai-ort::shared_kv_proposer` — `SharedKvProposerSession`, its
  signature detection, and its step contract. Its only caller was the deleted
  `SharedKvProposer`.
* `onnx-genai-ort::gemma4_assistant` — the model-name-gated Gemma4
  `Gemma4SharedKvSpec`/`Gemma4AssistantSignature` path. It was an orphan file,
  never declared as a module and therefore never compiled; the chained contract
  supersedes it with no model-name conditional at all.
* `native_decode::proposer` (`NativeProposerSession`) and
  `NativeDecodeSession::shared_kv_inputs`, plus the two unit tests pinning them.
* `SpeculativeProposerContext::shared_kv_slices` / `CandidateProposalInputs`, a
  field every remaining caller now passes as `None`.
* `SharedKvProposerSpec` and `resolve_shared_kv` in the metadata parser.

A legacy HuggingFace `proposal_type: shared_kv` block still parses — configs in
the wild carry it — but now degrades to `Unknown` with a diagnostic naming the
contract that replaced it (`speculative.proposal_execution: {kind: chained}`),
so a package author is pointed at the right block instead of at a resolution
path nothing can run. The parser's six shared-KV resolution tests collapse into
one that pins exactly that.

Verified: workspace builds clean (the `onnx-genai-bench` compare warnings are
pre-existing on main); `onnx-genai-engine` lib 568, `onnx-genai-metadata` lib 19,
`native_workflow_parity` 10, `gemma4_chained_workflow` 8, `native_workflow_smoke`
2, `onnx_genai_workflow_conformance` 14, `workflow_policy_e2e` 20,
`static_cache_workflow` 7, `native_speculative_driver` 3, `native_engine` 14.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`items_after_test_module` fires on `debug_shapes_enabled`, which #1716 left
below `captured_step_retry_tests`. It blocks
`cargo clippy --all-targets -p onnx-genai-ort -- -D warnings`, so move the
function above the module rather than leave the gate red for the next change.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…mains

The plan document was written as a gate; both gate conditions are now met and
Phase 1 is code, so it records what executes rather than what is proposed:

* a table mapping each `speculative.proposal_execution` field to the driver's
  use of it, so the "no model-name gate" claim is checkable field by field;
* the coverage boundary the tiny drafter imposes (it slices only the carry half,
  so the parity case cannot prove embedding-gather) and the separate test that
  does prove it;
* the exact deletions, and the gates each was re-pointed to;
* a new §6 listing the remaining work by symbol — the decode core to lift
  (`run_decode_loop` / `DecodeLoopBackend`, tied to `Engine` state), the missing
  plain-decoder `WorkflowSpec` synthesizer, the server's
  `if handle.pipeline` dispatch sites, the bench `--pipeline` mode, and the
  `SpeculativeProposer` implementations that still need to become contract-keyed
  interpreter constructs.

`NATIVE_WORKFLOW_BACKEND.md` boundary D is marked done, and the CLI capability
inventory no longer lists a `SpeculativeMode::SharedKv` that does not exist.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
… real-package gate

Three things the chained lift depends on that nothing was proving:

* **Liveness is scoped.** `externally_used_values` force-keeps values so island
  fusion cannot elide a proposal's inputs. If that ever stopped being conditional
  on a *chained* contract, every workflow in the repo would silently lose fusion
  with no test failing. Two unit cases pin it: no contract and a `block`
  proposer both contribute nothing, and a chained contract keeps exactly its own
  bindings plus the folded-carry seed — not the target logits the workflow emits.

* **The folded-carry contract is required, not inferred.** A folded carry with no
  `folded_carry_seed`, no `token_embedding`, or a chain with neither a
  recurrence nor a folded carry, must fail to resolve with a diagnostic naming
  the missing field — otherwise a runtime would be free to guess, which is the
  behavior this whole design removes.

* **Real packages, not just the fixture.** `chained_proposer_real` gains an
  opt-in `ONNX_GENAI_CHAINED_WORKFLOW_PACKAGE` case that resolves a real
  workflow package's chained contract through the interpreter and gathers its
  declared embedding table from the real target artifact — the check a per-model
  heuristic would pass vacuously and a contract read cannot. The published
  Gemma4-E2B speculative package carries the identical field shape, so the same
  driver serves both. Verified against the hermetic package
  (`REAL_CHAINED_WORKFLOW_EVIDENCE proposer=assistant target=target width<=4
  rollback_cells=4`); it skips cleanly when unset.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
The user requirement is one runtime and no caller-side dispatch. Two public
runtime types made that impossible by construction: every caller had to decide
which one it held, and each decided separately.

`PipelineEngine` is gone. The workflow interpreter is now internal state
(`pipeline::WorkflowRuntime`, `pub(crate)`) held by the one public `Engine`, and
the operations that used to exist twice exist once:

* `Engine::from_dir` resolves the package shape itself
  (`PipelineModelDirectory::load_if_declared`) and loads either half. Callers
  stop running that probe and picking a constructor — the server, the CLI, and
  the benchmarks each used to do it separately.
* `Engine::generate*` is one entry point: a package declaring
  `pipeline.workflow` is driven by the interpreter over its declared `tokens`
  output, a bare decoder by the decode core. The caller does not choose.
* `tokenize`, `execution_provider_status`, `effective_max_context`,
  `resource_snapshot`, `governor`, and the island/adapter diagnostics all resolve
  the same way, so a report reads the same regardless of package shape.
* Exactly one resource governor per runtime. `Engine::governor` is `Option`
  precisely so a workflow package does not get a second one that would
  double-count every reservation; the accessor returns whichever half owns it.
* `Engine::tokenizer` is `Option` because a workflow package may legitimately
  ship none (an image-generation pipeline). That is a load-time fact with an
  actionable error, not a panic in the decode loop.

Server, CLI, C ABI, benchmarks:

* `EngineBackend::{Single,Pipeline}` in the server driver collapses to one
  owned runtime, and `run_engine_driver` asks `engine.is_workflow()` instead of
  destructuring a variant.
* `ModelHandle::pipeline` is deleted. The completions route's four
  `if handle.pipeline` branches become one `submit_generation` call; the
  sessions and admin routes ask the runtime. `EngineDriver::warmup` no longer
  takes a caller-supplied `pipeline: bool` — it reads the flag it derived from
  the engine at startup, so a route and the driver cannot disagree.
* The CLI, `run_comfyui`, `profile_native`, the workflow example, and 40+ tests
  now name one type.

Dead API removed rather than left as a second surface: `WorkflowRuntime`'s
`from_dir`, `generate`, `detokenize`, `resource_snapshot`, `device_authority`,
`set_vram_limit`, and `clear_prefix_cache` are superseded by `Engine`'s and are
deleted.

New `one_runtime_e2e` (5 cases) pins the property against regression to two
runtimes: one constructor loads both shapes, autoregressive text generation and
session lifecycle are one API each, shape-specific diagnostics degrade uniformly
rather than erroring, and tokenization is served from whichever half owns the
tokenizer.

Verified: full `onnx-genai-engine` suite under `native-backend` (lib 572 + every
integration target, zero failures), server 228 / CLI / C ABI / bench suites,
`cargo fmt --all --check`, and
`cargo clippy --locked --all-targets -D warnings` on engine, server, cli, capi.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…y case

`Engine::models()` returns a `Result` now that one type serves both package
shapes, and this assertion sits behind `--features native-cuda` so the
non-CUDA build never type-checked it.

Verified on H200: `native_workflow_parity` 12/12 twice in a row, including
`chained_speculative_proposal_parity_native_cuda` and
`native_cuda_device_resident_multicomponent`. (Running three CUDA test binaries
concurrently on one device can make the residency case flake on contention; it
passes deterministically when the file runs on its own.)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
… blocker

Phase 3 is done and §6's caller items are struck through with what replaced
them. §7 is new and states the blocker precisely, with the three facts that make
it a decision rather than an implementation:

* `validate_metadata` rejects `model.io` beside `pipeline.workflow` by ratified
  rule, and every bare-decoder package declares `model.io`;
* `PipelineModelDirectory::load` errors without a `pipeline` section;
* `decoder_io_is_legacy()` — surfaced on `/v1/debug/config` — flips for every
  existing package the moment one is synthesized.

Three options are laid out with their costs, including the one that needs no
schema change (a `binding` component carrying
`contract.id: onnx-genai.autoregressive-decode`, dispatched the way adapters
already are). The second, independent gate is recorded too: the bare-decoder
generate path is covered by real-model tests that skip without weights, so the
rewire cannot be proven green here.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…anonical workflow

Parent decision A: keep authored bare-decoder metadata unchanged and valid, and
lower it at load into an internal canonical WorkflowSpec.

`onnx_genai_metadata::canonical` is that compiler. It is a pure function of the
declared ABI -- same `ModelIoSpec` in, byte-identical document out -- which is
what makes the lowered form *derived* rather than a second authored answer: it
cannot drift from `model.io`, because it is recomputed from it on every load.
Nothing writes it back, so `validate_model_io_against_workflow` never sees a
pair and no published package needs re-authoring.

It emits canonical YAML and parses it through the ordinary schema path. That is
deliberate: the lowered form is exactly what an author would have had to write,
it is diffable, and a field this module gets wrong is a loud error at lowering
time instead of a subtly malformed spec the interpreter trips over later. It
caught two on the first run -- a missing `increment` on `growing` recurrence and
a missing `contract` on the loop induction value.

No schema change. Both canonical components are `binding` components identified
by contract id -- `onnx-genai.autoregressive-decode` for one decoder forward
pass, `onnx-genai.token-policy` for next-token selection and stop detection --
dispatched the way workflow adapters already are. The published JSON schema is
untouched. KV is declared `management: runtime`, the schema's existing word for
"the runtime owns these buffers", so the paged / share-buffer / CUDA-graph
executors keep owning their KV and no per-step host round-trip is introduced.

`/v1/debug/config` gains `workflow_provenance` (authored | lowered | none).
`pipeline` keeps its meaning -- does the package *serialize* a workflow -- so a
lowered decoder reports `pipeline: false` with `workflow_provenance: lowered`,
and no report claims the file contains something it does not.

Tests: 7 unit cases (determinism, schema round-trip, KV becomes runtime-managed
state, an authored workflow is never lowered beside itself, the package is never
mutated and stays valid, unpaired KV refused, a missing port named) plus
`canonical_lowering_corpus` against the real corpus through the engine's own
loader -- 7 packages covered (phi35-mini int4 shared-buffer, qwen3-0.6b int4,
qwen2.5-0.5b CUDA, qwen2.5-coder-7b, phi4-mini CUDA, gpt-oss-20b MoE,
qwen3.5-2b) -- asserting deterministic lowering and that the lowered decoder's
ports mirror the resolved ABI exactly: no port invented, none dropped.

That corpus test immediately earned its keep. Routing `Engine::from_dir` through
`load_if_declared` had turned a shape guess into a hard load error for a package
recognized as multi-component that ships no native metadata. A resolution
failure now means "not a workflow package" and the decoder loader handles it, as
it always did.

Greedy goldens over 5 real models (phi35-mini, phi4-mini CUDA, qwen2.5-0.5b
CUDA, qwen3-0.6b, gpt-oss-20b) are byte-identical before and after. `.goldens/`
holds the harness and the c58eb5b baseline so later work diffs against the
previous runtime rather than only against itself.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
Delete temporary commit-message and local golden-evidence files that were unintentionally included with PR #1723. Runtime code, tests, and maintained documentation are unchanged.

Signed-off-by: justinchuby <justinchu@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

CI status, and what is and is not this branch's

CUDA compile (Linux x86_64) fails on main itself, at b9217fafc
(run),
along with CLI ORT on both platforms and Rust coverage (Windows).

The CUDA job's failing step is verify_cuda_test_honesty.py, which builds
onnx-runtime-ep-cuda's tests. Those do not compile against origin/main:

error[E0432]: unresolved import
  `onnx_runtime_ep_cuda::granule_transition::transition_granule_range_with_phase8_faults`
  --> crates/onnx-runtime-ep-cuda/tests/content_preserving_transition_gpu.rs:52:5
note: found an item that was configured out

Reproduced in a clean origin/main worktree with no changes from this branch.
This branch touches no file in onnx-runtime-ep-cuda.

What this branch's own CUDA surface does. The step that is about this work —
cargo clippy --locked -p onnx-genai-engine --features native-cuda --all-targets -- -D warnings — passes locally, and so do the runtime gates on an H200:

  • all 12 native_workflow_parity cases, including
    chained_speculative_proposal_parity_native_cuda with its device→host staging
    ratchet;
  • all 9 gemma4_chained_workflow cases.

Fixed here after CI caught it. The first CUDA compile failure on this PR
was this branch's: the ort-cuda-gated test module in pipeline/mod.rs had not
been updated for the governor move, and two divergence audits called a native
method whose signature changed. Both are fixed in cd3d9bd46, which is why the
job's clippy step now passes and only the main-side script failure remains.

Rust quality (fmt + clippy) and codecov/project pass. Coverage is
+0.32% against b9217fa.

One further pre-existing failure worth recording, unrelated to CI:
onnx-genai-ort::loader::tests::selected_non_dense_candidate_fails_explicitly
fails identically on origin/main@cb81745b0 in a clean worktree.

justinchuby added a commit that referenced this pull request Aug 23, 2026
Continuous batching and speculative decoding were the two loops #1723 left
outside the interpreter. Both took their bound and their stop from Rust, and
which one ran was decided beside the workflow rather than by it. They are now
declared steps: a body naming `onnx-genai.continuous-batch` is run by the
row-major batch executor, a body naming `onnx-genai.speculative-block` by the
propose-verify-accept executor, and the interpreter's `Loop` owns the iteration
in both cases exactly as it does for a single token.

## How a body is chosen

`decoder_workflow::iteration_variant` re-authors a package's own declared
generation loop into the body for a named `IterationPolicy`. The bound, the
liveness predicate, the carried cells, the termination and the emits stay
byte-identical; only the body's node changes. It sits in the module that authors
the single-token body, so the three cannot drift into three statements of one
loop.

`WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is
selected, and it selects by naming a policy. Nothing reads a package's decode
path, its file layout or its model identity to decide which iteration runs.
Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer,
native CUDA persistent batch) is still a question about a resource this process
holds -- the same family as `holds_decode_core()` -- and all three implement one
`BatchedDecodeSession` behind the one authored body.

## A batch publishes the batch's liveness, not each row's

The interpreter reads a loop's predicate *before* the body runs and preserves an
inactive row's carried state afterwards, so a slot that published `false` cannot
come back within the pass. A continuous batch recycles slots, so a per-slot bit
would permanently retire a slot the scheduler is about to backfill -- dropping
the backfilled request's tokens out of the declared stream and ending the pass
while that request was still running. The batch answers one question for every
slot: will this batch decode again? The per-row question the emit needs -- did
this slot contribute a token this iteration -- is `accepted_len`, which is zero
for a slot that sat the step out.

`a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and
fails with a per-slot bit.

## Measured

`authored_iteration_executors.rs` drives one runtime as a single-token
generation, a continuous batch and a speculative request, and reads the
per-contract counter recorded inside the interpreter's node dispatch. The three
produce **disjoint** contract sets. A batched generation that merely produced
the right tokens would keep passing if Rust went back to picking the row-major
loop from `ModelDecodePath::StaticCache`; a count against
`onnx-genai.continuous-batch` cannot be produced by a shape `match`.

`authored_iteration_e2e.rs` checks the executors still do the job: a batched row
stops on a declared end token that is **not** the first listed, rows with
different budgets finish independently, and a wider declared draft width
proposes more candidates per block while the rejected suffix is rolled back --
with the stream byte-identical to plain greedy.

The manager's scripted-decode unit tests now run on the authored body, so
multi-EOS stops, per-row device/host sampling parity, priority admission and
occupancy are exercised through the declared document.

## Deleted

`generate_batched_static` had its own row loop *and* its own copy of the per-row
token policy; a fixed batch is the continuous-batch algorithm with all-at-once
admission, so it now runs the same authored body and the copy is gone (with it,
the stale "past/present batching is deferred" refusal -- shared-buffer packages
batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one
`WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }`
is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is
`advance_block` -- both implementing the *same* contract, so the ORT and native
paths are two implementations of one declared step rather than two loops.

## The emit carries a length now

A batch row may sit out a step and a verified block may contribute several
tokens, so a re-authored emit publishes a per-row prefix bounded by the step's
declared `accepted_len`. That also makes the token output row-wise, giving each
row its own accumulated stream. Each row's stream is held against what the
executor committed for that row when the pass ends -- `emitted_token_rows` reads
every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on
the run-to-completion routes and in the server's driver, so the check runs where
batches are actually driven.

`tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with
`aliasing: permitted`: one declaration, and the package offers the shared buffer
a row-major batch decodes into. It exists so the authored batch has a hermetic
package to run on.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
Continuous batching and speculative decoding were the two loops #1723 left
outside the interpreter. Both took their bound and their stop from Rust, and
which one ran was decided beside the workflow rather than by it. They are now
declared steps: a body naming `onnx-genai.continuous-batch` is run by the
row-major batch executor, a body naming `onnx-genai.speculative-block` by the
propose-verify-accept executor, and the interpreter's `Loop` owns the iteration
in both cases exactly as it does for a single token.

## How a body is chosen

`decoder_workflow::iteration_variant` re-authors a package's own declared
generation loop into the body for a named `IterationPolicy`. The bound, the
liveness predicate, the carried cells, the termination and the emits stay
byte-identical; only the body's node changes. It sits in the module that authors
the single-token body, so the three cannot drift into three statements of one
loop.

`WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is
selected, and it selects by naming a policy. Nothing reads a package's decode
path, its file layout or its model identity to decide which iteration runs.
Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer,
native CUDA persistent batch) is still a question about a resource this process
holds -- the same family as `holds_decode_core()` -- and all three implement one
`BatchedDecodeSession` behind the one authored body.

## A batch publishes the batch's liveness, not each row's

The interpreter reads a loop's predicate *before* the body runs and preserves an
inactive row's carried state afterwards, so a slot that published `false` cannot
come back within the pass. A continuous batch recycles slots, so a per-slot bit
would permanently retire a slot the scheduler is about to backfill -- dropping
the backfilled request's tokens out of the declared stream and ending the pass
while that request was still running. The batch answers one question for every
slot: will this batch decode again? The per-row question the emit needs -- did
this slot contribute a token this iteration -- is `accepted_len`, which is zero
for a slot that sat the step out.

`a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and
fails with a per-slot bit.

## Measured

`authored_iteration_executors.rs` drives one runtime as a single-token
generation, a continuous batch and a speculative request, and reads the
per-contract counter recorded inside the interpreter's node dispatch. The three
produce **disjoint** contract sets. A batched generation that merely produced
the right tokens would keep passing if Rust went back to picking the row-major
loop from `ModelDecodePath::StaticCache`; a count against
`onnx-genai.continuous-batch` cannot be produced by a shape `match`.

`authored_iteration_e2e.rs` checks the executors still do the job: a batched row
stops on a declared end token that is **not** the first listed, rows with
different budgets finish independently, and a wider declared draft width
proposes more candidates per block while the rejected suffix is rolled back --
with the stream byte-identical to plain greedy.

The manager's scripted-decode unit tests now run on the authored body, so
multi-EOS stops, per-row device/host sampling parity, priority admission and
occupancy are exercised through the declared document.

## Deleted

`generate_batched_static` had its own row loop *and* its own copy of the per-row
token policy; a fixed batch is the continuous-batch algorithm with all-at-once
admission, so it now runs the same authored body and the copy is gone (with it,
the stale "past/present batching is deferred" refusal -- shared-buffer packages
batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one
`WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }`
is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is
`advance_block` -- both implementing the *same* contract, so the ORT and native
paths are two implementations of one declared step rather than two loops.

## The emit carries a length now

A batch row may sit out a step and a verified block may contribute several
tokens, so a re-authored emit publishes a per-row prefix bounded by the step's
declared `accepted_len`. That also makes the token output row-wise, giving each
row its own accumulated stream. Each row's stream is held against what the
executor committed for that row when the pass ends -- `emitted_token_rows` reads
every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on
the run-to-completion routes and in the server's driver, so the check runs where
batches are actually driven.

`tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with
`aliasing: permitted`: one declaration, and the package offers the shared buffer
a row-major batch decodes into. It exists so the authored batch has a hermetic
package to run on.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…l-closed same-device checks

Fix-forward revision of rejected PR #1854 (author Deckard locked out of this
artifact). Addresses all 9 Cycle-17 blocking findings from Roy's review:

1. Range-level (not value-level) rollback tracking: if range N within one
   ValueId commits and a LATER range in that same value hits Fatal, the
   earlier committed range is now correctly included in the rollback.
2. BoundaryApplicationOutcome.fatal_progress surfaces committed_count and
   poisoned_range straight from TransitionOutcome::Fatal (no more '..' discard).
3. Rollback-of-rollback failures are now explicit: a new RollbackFailure
   struct (value, range, detail, committed_count, poisoned_range,
   quarantined) replaces the previous silent all_ok=false.
4. New deterministic PLAN-ENTRY fault injection: a test-only
   apply_residency_plan_at_boundary_with_phase8_faults entry point accepts a
   per-ValueId DriverFaultPlan map, enabling 4 new GPU tests that exercise
   partial-commit-then-fatal, rollback-of-rollback, same-device rejection,
   and expert-group atomicity end to end (not just at the primitive level).
5. Same-device fail-closed: check_same_device validates allocator.device_key()
   and both pools' device_ordinal against the requested device_ordinal
   before any mutation.
6. Corrected the module doc's false 'one production consumer called at model
   load' claim; #1723 is now honestly described as technical current-main
   context only, with zero production callers today.
7. Cross-ValueId atomic expert-group tiering: new graph-derived
   ExpertWeightGroup type + expert_weight_groups(graph) (onnx-runtime-ep-api),
   purely structural from QMoe/BlockQuantizedMoe node inputs (no tensor-name
   heuristics). apply_residency_plan_at_boundary now takes an
   expert_groups parameter and enforces all-or-none group fallback via a
   validation-only precheck phase that completes for an entire group before
   any member is mutated.
8. Renamed misleading telemetry: device_bytes_before/after -> single
   truthful device_bytes_released; COARSE_RESIDENCY_PROFILE_ENV ->
   COARSE_RESIDENCY_ENABLE_ENV (deprecated alias kept for compatibility).

Also fixes a pre-existing (unrelated) Cargo feature-forwarding bug:
onnx-runtime-ep-cuda's gpu-tests feature did not forward to
onnx-runtime-cuda-memory/gpu-tests, hiding CudaVmmAllocator test-only items
needed to compile these tests.

Tests: 7/7 coarse_residency_plan_gpu (3 pre-existing + 4 new) pass on A100
(CUDA_VISIBLE_DEVICES=4), 4/4 coarse_residency_plan (non-GPU) pass, 3/3 +
91/91 unit tests (onnx-runtime-ep-cuda coarse_residency + onnx-runtime-ep-api)
pass, 530/530 onnx-runtime-ep-cuda lib tests pass (no regressions), and all
29 Slice2-4 GPU regressions (composable_vmm_production_gpu,
content_preserving_transition_gpu, expert_bank_remap_cost_gpu) pass
unchanged. fmt/clippy clean on all touched crates. Independent code review
(not Deckard) found no blocking issues.

PR remains in draft; Roy's explicit re-review is still required before merge.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
Continuous batching and speculative decoding were the two loops #1723 left
outside the interpreter. Both took their bound and their stop from Rust, and
which one ran was decided beside the workflow rather than by it. They are now
declared steps: a body naming `onnx-genai.continuous-batch` is run by the
row-major batch executor, a body naming `onnx-genai.speculative-block` by the
propose-verify-accept executor, and the interpreter's `Loop` owns the iteration
in both cases exactly as it does for a single token.

## How a body is chosen

`decoder_workflow::iteration_variant` re-authors a package's own declared
generation loop into the body for a named `IterationPolicy`. The bound, the
liveness predicate, the carried cells, the termination and the emits stay
byte-identical; only the body's node changes. It sits in the module that authors
the single-token body, so the three cannot drift into three statements of one
loop.

`WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is
selected, and it selects by naming a policy. Nothing reads a package's decode
path, its file layout or its model identity to decide which iteration runs.
Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer,
native CUDA persistent batch) is still a question about a resource this process
holds -- the same family as `holds_decode_core()` -- and all three implement one
`BatchedDecodeSession` behind the one authored body.

## A batch publishes the batch's liveness, not each row's

The interpreter reads a loop's predicate *before* the body runs and preserves an
inactive row's carried state afterwards, so a slot that published `false` cannot
come back within the pass. A continuous batch recycles slots, so a per-slot bit
would permanently retire a slot the scheduler is about to backfill -- dropping
the backfilled request's tokens out of the declared stream and ending the pass
while that request was still running. The batch answers one question for every
slot: will this batch decode again? The per-row question the emit needs -- did
this slot contribute a token this iteration -- is `accepted_len`, which is zero
for a slot that sat the step out.

`a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and
fails with a per-slot bit.

## Measured

`authored_iteration_executors.rs` drives one runtime as a single-token
generation, a continuous batch and a speculative request, and reads the
per-contract counter recorded inside the interpreter's node dispatch. The three
produce **disjoint** contract sets. A batched generation that merely produced
the right tokens would keep passing if Rust went back to picking the row-major
loop from `ModelDecodePath::StaticCache`; a count against
`onnx-genai.continuous-batch` cannot be produced by a shape `match`.

`authored_iteration_e2e.rs` checks the executors still do the job: a batched row
stops on a declared end token that is **not** the first listed, rows with
different budgets finish independently, and a wider declared draft width
proposes more candidates per block while the rejected suffix is rolled back --
with the stream byte-identical to plain greedy.

The manager's scripted-decode unit tests now run on the authored body, so
multi-EOS stops, per-row device/host sampling parity, priority admission and
occupancy are exercised through the declared document.

## Deleted

`generate_batched_static` had its own row loop *and* its own copy of the per-row
token policy; a fixed batch is the continuous-batch algorithm with all-at-once
admission, so it now runs the same authored body and the copy is gone (with it,
the stale "past/present batching is deferred" refusal -- shared-buffer packages
batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one
`WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }`
is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is
`advance_block` -- both implementing the *same* contract, so the ORT and native
paths are two implementations of one declared step rather than two loops.

## The emit carries a length now

A batch row may sit out a step and a verified block may contribute several
tokens, so a re-authored emit publishes a per-row prefix bounded by the step's
declared `accepted_len`. That also makes the token output row-wise, giving each
row its own accumulated stream. Each row's stream is held against what the
executor committed for that row when the pass ends -- `emitted_token_rows` reads
every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on
the run-to-completion routes and in the server's driver, so the check runs where
batches are actually driven.

`tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with
`aliasing: permitted`: one declaration, and the package offers the shared buffer
a row-major batch decodes into. It exists so the authored batch has a hermetic
package to run on.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…r the unified interpreter (#1883)

## Context

Follow-up from the merged S3 capacity-emission slice (#1860): the user
directed re-validating that work's integration against the newly-merged
unified-workflow runtime (#1723, "one runtime, one interpreter, one
drive").

## What broke

#1723 made package loading eagerly resolve **every** declared
`pipeline.workflow` component's on-disk artifact at load time
(`crates/onnx-genai-ort/src/loader.rs`). Two fixtures added earlier in
#1832 — `tests/fixtures/tiny-deepseek-v4-qmoe/` and
`tests/fixtures/tiny-glm52-full-attention/` — hand-authored a full
workflow document referencing ten auxiliary policy components
(`cache_length_update`, `token_sampler`, `termination`,
`token_state_update`, `last_token_logits`, `decoder_state_initializer`,
`decoder_step_update`, `termination_batch_initializer`, `token_to_slot`,
`generated_length_update`) whose `policies/*.onnx` artifact files were
never committed. Under the old canonical-dispatch loader this apparently
went unnoticed; under the new interpreter it hard-fails:

```
Failed to resolve workflow package: IO error: workflow ONNX component
file not found: .../tiny-deepseek-v4-qmoe/policies/cache_length_update.onnx
```

This broke 3/4 tests in `deepseek_v4_tiny_qmoe_e2e.rs` and 3/4 in
`glm_tiny_full_attention_e2e.rs` — both are correctness gates for
already-merged Sapper work (S3 KV capacity-emission, GLM-5.2
enablement).

## Fix

Re-emitted both fixtures' `inference_metadata.yaml` with the
repository's own `migrate_model_io --reemit` tool — the exact remedy
#1723 used for the 14 fixtures it found in the same stale state
(`tiny-llm`, `tiny-mtp-full`). Metadata-only: `model.onnx.textproto`,
`manifest.json`, and `tokenizer.json` are byte-identical/untouched, so
DeepSeek-V4's dense-CSA/QMoE graph and GLM-5.2's full-attention graph
are completely unaffected. No Rust source touched.

Confirmed the regenerated documents still:
- declare `onnx-genai.autoregressive-decode`
- use `bounded` (not `invariant`) sequence recurrence — the exact
latent-bug class #1723 called out
- correctly omit MTP (matching `num_nextn_predict_layers: 0` / no MTP
field)
- validate cleanly against the schema (`validate_metadata`)

## Verification (fresh `origin/main` @ `7a35dba77`, real A100 GPU)

| Test | Result |
|---|---|
| `deepseek_v4_tiny_qmoe_e2e` | 4/4 pass, `captures=1 replays=10
fallbacks=0`, anchor IDs identical across native CPU/CUDA/stock-ORT |
| `glm_tiny_full_attention_e2e` | 4/4 pass, anchor IDs identical across
native CPU/CUDA/stock-ORT |
| `glm_tiny_qmoe_native_cuda_e2e` | 4/4 pass (unaffected, sanity
re-check) |
| `deepseek_v2_tiny_qmoe_native_e2e` | 1/1 pass (unaffected, sanity
re-check) |
| `authored_body_selects_executor` | 4/4 pass |
| `one_runtime_e2e` | 6/6 pass |
| `native_workflow_parity` | 10/10 pass |
| `native_workflow_smoke` | 2/2 pass |
| `canonical_execution_parity` | 7/7 pass |

Also ran `cargo test -p onnx-genai-engine --features native-backend`
(full lib+test suite): 613/614 pass; the one failure
(`native_generate_rejects_over_kv_byte_budget_before_backend_run`) is
**pre-existing on main without this change** (confirmed via `git stash`
bisection) — unrelated to this fix, out of scope here.

Independent review: **APPROVE**, no blocking findings (contract/MTP/EOS
semantics verified unchanged, `model.onnx.textproto` confirmed
byte-identical).

## Scope

Metadata-only fixture fix. No workflow-runtime refactor, no
native-vs-ORT branching, no side decode loop, no changes to the S3
capacity-emission mechanism itself (`rewrite_kv_capacity_appends`) —
that remains verified compatible with the unified
`run_workflow_node`/contract-dispatch model as-is, since it operates
purely at ONNX-graph-build time below the workflow-interpreter layer.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Signed-off-by: Sapper <sapper@squad.local>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
Continuous batching and speculative decoding were the two loops #1723 left
outside the interpreter. Both took their bound and their stop from Rust, and
which one ran was decided beside the workflow rather than by it. They are now
declared steps: a body naming `onnx-genai.continuous-batch` is run by the
row-major batch executor, a body naming `onnx-genai.speculative-block` by the
propose-verify-accept executor, and the interpreter's `Loop` owns the iteration
in both cases exactly as it does for a single token.

## How a body is chosen

`decoder_workflow::iteration_variant` re-authors a package's own declared
generation loop into the body for a named `IterationPolicy`. The bound, the
liveness predicate, the carried cells, the termination and the emits stay
byte-identical; only the body's node changes. It sits in the module that authors
the single-token body, so the three cannot drift into three statements of one
loop.

`WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is
selected, and it selects by naming a policy. Nothing reads a package's decode
path, its file layout or its model identity to decide which iteration runs.
Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer,
native CUDA persistent batch) is still a question about a resource this process
holds -- the same family as `holds_decode_core()` -- and all three implement one
`BatchedDecodeSession` behind the one authored body.

## A batch publishes the batch's liveness, not each row's

The interpreter reads a loop's predicate *before* the body runs and preserves an
inactive row's carried state afterwards, so a slot that published `false` cannot
come back within the pass. A continuous batch recycles slots, so a per-slot bit
would permanently retire a slot the scheduler is about to backfill -- dropping
the backfilled request's tokens out of the declared stream and ending the pass
while that request was still running. The batch answers one question for every
slot: will this batch decode again? The per-row question the emit needs -- did
this slot contribute a token this iteration -- is `accepted_len`, which is zero
for a slot that sat the step out.

`a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and
fails with a per-slot bit.

## Measured

`authored_iteration_executors.rs` drives one runtime as a single-token
generation, a continuous batch and a speculative request, and reads the
per-contract counter recorded inside the interpreter's node dispatch. The three
produce **disjoint** contract sets. A batched generation that merely produced
the right tokens would keep passing if Rust went back to picking the row-major
loop from `ModelDecodePath::StaticCache`; a count against
`onnx-genai.continuous-batch` cannot be produced by a shape `match`.

`authored_iteration_e2e.rs` checks the executors still do the job: a batched row
stops on a declared end token that is **not** the first listed, rows with
different budgets finish independently, and a wider declared draft width
proposes more candidates per block while the rejected suffix is rolled back --
with the stream byte-identical to plain greedy.

The manager's scripted-decode unit tests now run on the authored body, so
multi-EOS stops, per-row device/host sampling parity, priority admission and
occupancy are exercised through the declared document.

## Deleted

`generate_batched_static` had its own row loop *and* its own copy of the per-row
token policy; a fixed batch is the continuous-batch algorithm with all-at-once
admission, so it now runs the same authored body and the copy is gone (with it,
the stale "past/present batching is deferred" refusal -- shared-buffer packages
batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one
`WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }`
is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is
`advance_block` -- both implementing the *same* contract, so the ORT and native
paths are two implementations of one declared step rather than two loops.

## The emit carries a length now

A batch row may sit out a step and a verified block may contribute several
tokens, so a re-authored emit publishes a per-row prefix bounded by the step's
declared `accepted_len`. That also makes the token output row-wise, giving each
row its own accumulated stream. Each row's stream is held against what the
executor committed for that row when the pass ends -- `emitted_token_rows` reads
every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on
the run-to-completion routes and in the server's driver, so the check runs where
batches are actually driven.

`tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with
`aliasing: permitted`: one declaration, and the package offers the shared buffer
a row-major batch decodes into. It exists so the authored batch has a hermetic
package to run on.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
… the matrix (main is red) (#1888)

`main` is red on a **required** check and has been since #1883, so
nothing in the repository can merge.

```
the_matrix_holds_for_every_maintained_workflow
  tests/fixtures/tiny-deepseek-v4-qmoe/inference_metadata.yaml: graph component count
    left: 1   right: 11
```

`crates/onnx-genai-metadata/tests/decoder_recognizer_agreement.rs` runs
in `Fast (Linux x86_64)`, which with `Rust quality` is one of only two
required status checks on `main`. Same class of repo-wide block as the
`op_rules` catalog pin (#1860 → #1872/#1882): a pin that fell behind a
deliberate change.

## Bisected

| commit | result |
|---|---|
| `182d1f776` (#1864, parent) | **ok**, 14/14 |
| `7adfae901` (#1883) | **FAILED**, 13/14 |
| `7a0cb6c39` (tip) | **FAILED**, 13/14 |

Both re-emitted fixtures are affected — DeepSeek-V4 fails first, and
GLM-5.2 has the identical 11 → 1 collapse behind it.

## Why the fixtures are right and the matrix is wrong

#1832 hand-authored both documents with ten auxiliary policy components
(`cache_length_update`, `token_sampler`, `termination`, …) whose
`policies/*.onnx` artifacts **were never committed**. #1723 made package
loading resolve every declared component eagerly, turning that latent
staleness into a hard load failure, and #1883 re-emitted both with
`migrate_model_io --reemit` — the same remedy #1723 applied to the
fourteen fixtures it found in the same state.

So the ten components were never real. The re-emission is the fix; the
matrix is what fell behind it.

## Delta accounted for, not made to match

Each field justified from the re-emitted document rather than read off
the failure:

| field | was | now | why |
|---|---|---|---|
| components | 11 | 1 | the document declares a single `decoder`; the
ten policies are gone because their artifacts never existed |
| cardinality | `Composite` | `SingleGraph` | direct consequence of the
above |
| decoder | `Some("model")` | `Some("decoder")` | the re-emit tool names
it `decoder`; the hand-written document said `model` |
| layer 1 | `false` | `true` | exactly one decoder component, no
competing roles |
| layer 2 | `None` | `Some("decoder")` | both layers now resolve |

## Checked for silently deleted coverage first

Updating a pin can quietly retire the only witness to a classification
path, so I checked before touching the rows rather than after:

- the `11` / `Composite` / `Some("model")` shape these two used to pin
is **still pinned** by
`tests/fixtures/onnx_genai_workflows/decoder/inference_metadata.yaml`;
- **44 of 64** rows still carry the `false` / `None` two-layer split
that motivates reporting the layers separately.

No path loses its only witness here. That reasoning is recorded in a
comment beside the rows so the next person does not have to redo it.

## Verification

`cargo test -p onnx-genai-metadata` — **14/14** on the agreement suite,
all 22 targets green. `cargo fmt --all -- --check` clean. `cargo clippy
--locked --all-targets -p onnx-genai-metadata -- -D warnings` clean.

Not my area — I found this because it blocked #1877. If the intent was
for these fixtures to keep an eleven-component workflow, the correct fix
is to commit the missing `policies/*.onnx` artifacts instead and this PR
should be closed. On the evidence in #1883 that is not the intent.

Auto-merge armed, waiting on required checks. No `--admin`, no bypass.

Refs #1832, #1723, #1883.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
#1893)

## Summary

While running a read-only tiny-MoE E2E baseline (3 fixtures) through the
unified workflow/runtime architecture from #1723, `cargo build --release
-p onnx-genai-bench --features native-cuda --bin profile_native` **did
not compile at all** on `origin/main`. `onnx-genai-bench` has zero CI
coverage (absent from `.github/workflows/ci.yml`), so this — plus two
separate, independent, pre-existing bugs — went undetected. This PR is
the tooling-only repair; the E2E baseline results themselves are
reported separately (see below).

**Scope: `crates/onnx-genai-bench` only. No `onnx-genai-engine`
(production) changes.**

## What broke, and why

1. **`Engine::from_dir_with_config` doesn't exist post-#1723** →
`Engine::from_dir`.
2. **`Pipeline::clear_prefix_cache()` doesn't exist post-#1723** (no
public replacement was ever wired into this bench crate) → replaced both
call sites with the public, per-request `GenerateOptions::cold_start`,
the documented #1723 successor. Scoped to one request instead of
mutating shared engine state — strictly better.
3. **`generate_with_callback(pipeline_request(..)?, ..)` type mismatch**
(needs a `GenerateRequest`, got a `PipelineGenerateRequest`) at two call
sites → `generate_with_pipeline_callbacks`. One of the two was hidden by
a rustc diagnostic-suppression quirk on the first error and only
surfaced after fixing it.
4. **`NativeDecodeSession::generate` and family became `pub(crate)`** →
added a local greedy-only `native_session_generate()` reimplementation
(with matching `loop.next_logits`/`loop.sampling` `prof_span!`
instrumentation so the `--synthetic` profiling regression test keeps
passing) and a manual `token_logprob()` log-softmax/top-k helper for
`--dump-logprobs`.
5. **Separate, pre-existing bug (unrelated to #1723):**
`NativeDecodeSession::from_session(..)` requires shape/dtype-unambiguous
ports. The bench crate's synthetic decoder's 3 `Int64 [-1,-1]` inputs
(`input_ids`/`attention_mask`/`position_ids`) are structurally
indistinguishable, so this deterministically failed ("N ports match") —
by design, confirmed via an existing engine-crate test. Fixed with the
documented `from_session_with_io(..)` escape hatch and a new
`synthetic_decoder_io()` `DecoderAbi` (adapted from
`fused_batch_prefill.rs`'s `synthetic_io()` helper — which is itself
independently bit-rotted, see "Found but not fixed" below).
6. **Separate packaging bug:** `onnx-genai-metadata` (needed for
`DecoderAbi`) was `[dev-dependencies]`-only, so no regular binary target
could ever call that public API surface. Promoted to `[dependencies]`.
7. Two pre-existing clippy lints inside `profile_native.rs`
(`manual_contains`, `needless_borrows_for_generic_args`) and one on
`tests/profile_native.rs`'s `assert_floor` helper
(`too_many_arguments`), fixed to keep a scoped clippy gate green.

## Independent review

An independent code-review pass on the initial repair found two
substantive issues, both fixed in this PR:

- **`token_logprob()` could panic** on an overflowed (`+inf`) logit:
`max` became infinite, `v - max` on that entry produced `NaN`, which
propagated through every ranked log-probability, and the old
`partial_cmp(..).expect("logits must not be NaN")` sort comparator
panicked on the resulting NaNs. Now mirrors
`decode_loop::logprob_for_token`'s NaN/Inf-safe approach (filters
non-finite logits, uses `total_cmp`). Added regression tests.
- **`native_session_generate`'s bail condition used
`!chain.is_empty()`**, an incidental proxy. Replaced with
`!chain.preserves_argmax()` — the actual eligibility test the engine's
own device-greedy fast path uses — and updated the
`--repetition-penalty` doc comment, which had drifted out of sync (it
read as though non-default values worked in raw/default mode, when in
fact — like `--top-p`/`--top-k`/`--min-p` — they require
`--steady`/`--pipeline`).

## Found but not fixed (out of scope, reported for follow-up)

- `crates/onnx-genai-bench/tests/fused_batch_prefill.rs` references a
nonexistent `kv_update` field on `DecoderAbi` — independently broken,
unrelated to #1723 or this repair.
- `bench_generic.rs`/`compare.rs` carry their own pre-existing clippy
failures under `--all-targets` (const-assertion lint, collapsible-if) —
left alone; this PR's clippy gate is deliberately scoped to `--bin
profile_native --test profile_native`, not `--all-targets`.
- Two of the three tiny-MoE fixtures (`tiny-deepseek-v4-qmoe`,
`tiny-glm52-full-attention`) currently fail to load at all
(`Engine::from_dir`) because their committed `inference_metadata.yaml`
declares 10 `policies/*.onnx` workflow-component artifacts that were
never committed to the fixture directories. Confirmed via the
*unmodified* authoritative engine-crate regression tests
(`deepseek_v4_tiny_qmoe_e2e.rs`, `glm_tiny_full_attention_e2e.rs`) —
this is a pre-existing fixture-generation gap, not caused by this PR or
#1723. `generate.py --mobius-root <path>` (a local Mobius checkout
exists) can presumably regenerate the missing artifacts; tracked as a
follow-up, not attempted here (out of scope: this PR does not touch
fixtures).

## Verification

- `cargo build --release -p onnx-genai-bench --features native-cuda
--bin profile_native` — clean, no warnings.
- `cargo test --release -p onnx-genai-bench --features native-cuda --bin
profile_native --test profile_native` — 8 unit + 4/5 non-ignored
integration tests pass (1 correctly `#[ignore]`d, needs real multi-node
hardware).
- `cargo fmt -p onnx-genai-bench -- --check` — clean.
- `cargo clippy --release -p onnx-genai-bench --features native-cuda
--bin profile_native --test profile_native -- -D warnings` — clean.
- Re-verified end-to-end correctness on an idle A100 (GPU 1) after all
fixes: `profile_native --model tests/fixtures/tiny-glm52-qmoe-indexshare
--ep cuda --steady --tokens 12 ...` still produces `generated_token_ids:
[33, 221, 186, 142, 162, 143, 148, 18, 18, 18, 18, 18]`, byte-identical
to the locked `ANCHOR_IDS` regression constant in
`crates/onnx-genai-engine/tests/glm_tiny_qmoe_native_cuda_e2e.rs`.

Closes: tooling blocker for issue #82's E2E baseline effort (see
separate report for the actual 3-fixture baseline measurements).

Working as Sebastian (Performance Engineer).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…ntly (#1899)

## Why

`profile_native` is the tool every decode benchmark under
`docs/benchmarks/` is produced with — and **nothing in CI compiles it**.

- The Fast job's `cargo build --locked -p …` list cannot reach it: it is
a **bin** that needs `--features bench-native`.
- `benchmark.yml` runs only `--bench no_model` and `--bench model`,
neither of which builds it.
- A root `cargo build` does not enable the feature either.

So #1723 broke it across five API changes and nobody noticed until I
tried to use the tool by hand. That was #1878 (fixed by #1893) — but the
*reason it landed* is this gap, and without closing it the next engine
API change breaks it again just as quietly.

## What

One step in the Fast (Linux) job that **builds, not runs**, the bin.
Cheap: it compiles crates the job already builds, plus the bin itself.

```yaml
- name: Build the native profiling bin
  run: >-
    cargo build --locked
    -p onnx-genai-bench --features bench-native --bin profile_native
```

## Verification

- The command itself is verified green on current `main` (`1be9f2cc2`):
`Finished dev profile in 19m 08s` — that is how I confirmed #1893
actually fixed #1878 before closing it.
- `ci.yml` parses as valid YAML and the step lands in the `fast-linux`
job (checked by loading the workflow and inspecting the parsed step
list, not by eyeballing the indentation).

## Note on cost

This adds a compile of `onnx-genai-bench` with `bench-native` to the
Fast job. That feature pulls in the native decode stack, which the job
already builds for other crates, so the marginal cost is the bench crate
itself rather than a new dependency tree. If it turns out to lengthen
the job more than is wanted, the honest alternative is a separate
scheduled job — but a silent breakage that reaches a benchmarking tool
is worse than a slower Fast job, and this is the cheapest form of the
check.

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…scheduler (#1900)

## Root cause

#1723 ("one runtime, one interpreter, one drive") changed
`generate_with_callbacks`'s top-level dispatch from checking the
`decode_backend` field to checking `holds_decode_core()`
(`native_session.is_some() || session.is_some()`).

Pre-#1723, **every** `Native`-backend request — regardless of whether a
decode session already existed — went through the scheduler-admitting
cold-start path (`generate_native_cold_with_callback`), which checks the
KV byte budget before touching any backend.

Post-#1723, a `Native`-backend `Engine` with no session now silently
falls into the unguarded interpreted path (`generate_interpreted` →
`generate_with_pipeline_callbacks` → `run_declared_generation`),
skipping scheduler admission entirely. This is a **real, reachable**
state: `Engine::from_dir` → `Self::decode_core_covers(&workflow)` false
→ `from_interpreted_dir` → `Engine::from_workflow` is a legitimate,
doc-commented construction path ("a package whose components the
interpreter invokes"), and `WorkflowRuntime::decode_backend()` can
independently be `Native` for such a package. Today no shipping package
pairs this with a byte-budget-configured scheduler
(`Engine::from_workflow` always uses
`Scheduler::new(SchedulerConfig::default())`), so the practical blast
radius is currently zero — but the gap is real and was caught by a
pre-existing test,
`native_generate_rejects_over_kv_byte_budget_before_backend_run`, which
was failing on main with the wrong (deep, non-admission) error message
instead of a clean scheduler rejection.

## Fix

- Added `admit_interpreted_generate_request` (`engine/runtime.rs`),
which computes the prompt token count via a new
`interpreted_prompt_token_count` (handling `TokenIds` by length,
`TokenRows` by `rows.len() * max(row length)` — since equal-length rows
bind into one batched `[rows, columns]` tensor and run as a single graph
step — and `Text` via the package's own tokenizer) and calls the
existing `admit_generate_request_with_scheduler`.
- Called this from `generate_with_pipeline_callbacks`'s
(`engine/workflow_api.rs`) own no-decode-core, prompt-only branch — the
single place every such request converges, whether it arrives cold
(`generate_interpreted`) or through a continuing workflow session
(`generate_in_workflow_session`, which threads the caller's own session
id through so admission accumulates the same way the decode-core path's
session continuation already does).
- Tensor-bound (multimodal) requests are deliberately left unchanged,
matching pre-existing, unrelated precedent (CLI
`interactive.rs`/`transcribe.rs`, server `driver.rs`) that skips
admission regardless of decode-core status.

## Before / after evidence

Before the fix,
`native_generate_rejects_over_kv_byte_budget_before_backend_run` failed
with:
```
workflow component 'decoder' input 'past.key' references unavailable value 'decoder_kv.0000'
```
(a deep node-execution failure, not a scheduler rejection). After the
fix it passes cleanly with `"scheduler admission failed: KV byte
budget..."`.

A new test,
`native_generate_in_workflow_session_rejects_over_kv_byte_budget`,
exercises the session-continuation path directly (opens a real session
via `create_session()`, drives `generate_in_workflow_session()`). I
confirmed by temporarily reverting `workflow_api.rs` alone that this
test fails against the pre-fix code with the same deep `references
unavailable value` error, and passes after the fix — closing a gap an
independent review pass on an earlier version of this change flagged
(that version only guarded the cold-call path, missing session
continuation, and undercounted `TokenRows`).

## Review

Independently reviewed twice: once on an initial version (which lived
inline in `generate_interpreted` and only covered the cold-call path —
REQUEST CHANGES, two findings), and once on this final restructured
version, which addresses both findings by moving the admission call site
to the shared no-decode-core branch of
`generate_with_pipeline_callbacks` and fixing the `TokenRows` formula.
Second pass: **APPROVE**, after tracing control flow, confirming
`scheduler.complete()` runs on every exit path, confirming no other
caller bypasses the new admission branch, and independently reproducing
the before/after test evidence.

## Tests (all local; no CI wait per current session directive)

- `cargo test -p onnx-genai-engine --features native-backend --lib`: 618
passed, 0 failed, 1 ignored
- `cargo test -p onnx-genai-engine --features native-cuda --lib`: 648
passed, 0 failed, 4 ignored
- `authored_body_selects_executor` (4), `one_runtime_e2e` (6),
`native_workflow_parity` (13), `native_workflow_smoke` (2),
`canonical_execution_parity` (7): 32/32 passed
- `cargo fmt --check -p onnx-genai-engine`: clean
- `cargo clippy -p onnx-genai-engine --features native-cuda
--all-targets -- -D warnings`: clean
- GPU e2e on idle A100 (`CUDA_VISIBLE_DEVICES=1`, `--test-threads=1`):
- `deepseek_v4_tiny_qmoe_e2e`: 4/4, `captures=1 replays=10 fallbacks=0`
  - `glm_tiny_qmoe_native_cuda_e2e`: 4/4
  - `glm_tiny_full_attention_e2e`: 4/4
  - `deepseek_v2_tiny_qmoe_native_e2e`: 2/2

## Scope

No workflow-unification refactors, no CUDA residency/paging/scheduler
changes, no DeepSeek/GLM export/model-name branching. Scoped strictly to
closing the #1723-introduced admission gap in the shared
interpreted-path runtime code.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…l-closed same-device checks

Fix-forward revision of rejected PR #1854 (author Deckard locked out of this
artifact). Addresses all 9 Cycle-17 blocking findings from Roy's review:

1. Range-level (not value-level) rollback tracking: if range N within one
   ValueId commits and a LATER range in that same value hits Fatal, the
   earlier committed range is now correctly included in the rollback.
2. BoundaryApplicationOutcome.fatal_progress surfaces committed_count and
   poisoned_range straight from TransitionOutcome::Fatal (no more '..' discard).
3. Rollback-of-rollback failures are now explicit: a new RollbackFailure
   struct (value, range, detail, committed_count, poisoned_range,
   quarantined) replaces the previous silent all_ok=false.
4. New deterministic PLAN-ENTRY fault injection: a test-only
   apply_residency_plan_at_boundary_with_phase8_faults entry point accepts a
   per-ValueId DriverFaultPlan map, enabling 4 new GPU tests that exercise
   partial-commit-then-fatal, rollback-of-rollback, same-device rejection,
   and expert-group atomicity end to end (not just at the primitive level).
5. Same-device fail-closed: check_same_device validates allocator.device_key()
   and both pools' device_ordinal against the requested device_ordinal
   before any mutation.
6. Corrected the module doc's false 'one production consumer called at model
   load' claim; #1723 is now honestly described as technical current-main
   context only, with zero production callers today.
7. Cross-ValueId atomic expert-group tiering: new graph-derived
   ExpertWeightGroup type + expert_weight_groups(graph) (onnx-runtime-ep-api),
   purely structural from QMoe/BlockQuantizedMoe node inputs (no tensor-name
   heuristics). apply_residency_plan_at_boundary now takes an
   expert_groups parameter and enforces all-or-none group fallback via a
   validation-only precheck phase that completes for an entire group before
   any member is mutated.
8. Renamed misleading telemetry: device_bytes_before/after -> single
   truthful device_bytes_released; COARSE_RESIDENCY_PROFILE_ENV ->
   COARSE_RESIDENCY_ENABLE_ENV (deprecated alias kept for compatibility).

Also fixes a pre-existing (unrelated) Cargo feature-forwarding bug:
onnx-runtime-ep-cuda's gpu-tests feature did not forward to
onnx-runtime-cuda-memory/gpu-tests, hiding CudaVmmAllocator test-only items
needed to compile these tests.

Tests: 7/7 coarse_residency_plan_gpu (3 pre-existing + 4 new) pass on A100
(CUDA_VISIBLE_DEVICES=4), 4/4 coarse_residency_plan (non-GPU) pass, 3/3 +
91/91 unit tests (onnx-runtime-ep-cuda coarse_residency + onnx-runtime-ep-api)
pass, 530/530 onnx-runtime-ep-cuda lib tests pass (no regressions), and all
29 Slice2-4 GPU regressions (composable_vmm_production_gpu,
content_preserving_transition_gpu, expert_bank_remap_cost_gpu) pass
unchanged. fmt/clippy clean on all touched crates. Independent code review
(not Deckard) found no blocking issues.

PR remains in draft; Roy's explicit re-review is still required before merge.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
…1892)

## Turn 3 was turn 3's prompt

`justinchuby/qwen3-0.6b-onnx-genai@a89c02d3` declares all 65 of its
state cells
`scope: invocation`, and its 56 cache cells `release_boundary:
invocation`. The
interpreter honours that. Measured on the published package, before this
branch:

| | tokens |
|---|---|
| turn 3 of a session | `[576, 3491, 374, 311, 1477, 279, 1372, 315]` |
| the same prompt sent cold, no session | `[576, 3491, 374, 311, 1477,
279, 1372, 315]` |
| one request carrying the whole conversation | `[358, 1079, 264, 5458,
315, 279, 12103, 315]` |

Turn 3 was byte-identical to the cold request. This is the regression
#1723
called out and named two possible fixes for. The other one — *prefer the
decode
core when a package's workflow is decoder-shaped despite shipping
several
components* — is the runtime supplying a step the package declares a
graph for,
which is the thing #1723 removed. This is the first one: **the package
declares
its conversation, and the runtime carries it.**

## Root cause: the document, not the interpreter

`scope: session` says *keep this*. It never said **how the next
invocation
reaches what was kept**, and there are only two answers:

* **the graph reads it** — a loop carries the cell, or a step consumes
the value
its `initializer` names. `moshiko-full-duplex` does this with 72 cells,
and the
  runtime already carried them correctly.
* **the request binding rejoins it** — and nothing could say so.

A decoder like this one has no reader of the second kind and cannot grow
one.
Its `decoder_state_initializer` takes `(prompt_tokens, prompt_lengths,
max_iterations)` and emits empty `past_key_values.*`; its setup then
prefills the
turn's prompt from that empty cache. There is no port for a prior
session length,
so a runtime handing it turn 1's cache would produce a mask and a cache
that
disagree. Declaring `scope: session` on the existing cells would not
have fixed
multi-turn — it would have produced *garbage*, which is worse. That is
why this
is not a one-line metadata patch.

## The declaration

```yaml
conversation:
  contract: {dtype: int64, rank: 2, shape: [batch_size, conversation_length],
             batch_layout: {kind: request_aligned, axis: 0}}
  class: semantic
  scope: session
  initializer: request.input_ids
  recurrence: {kind: bounded, axis: 1, max: package.max_context}
  management: runtime
  release_boundary: session
  session:
    policy: exclusive
    continuation:
      kind: prompt_prefix
      prompt_input: request.input_ids     # must carry role prompt_tokens
      tokens_output: tokens               # must carry role tokens
```

A turn's prompt is the cell's value followed by the caller's tokens; the
cell
then absorbs both that prompt and what the turn published. It is bound
**before
the request is bound**, so every input derived from the prompt — the
token
tensor, `prompt_lengths`, the mask the initializer builds — is derived
from the
whole conversation, exactly as for a caller who sent it in one request.
Rewriting
the bound tensors afterwards would leave the derived ones describing
only the
turn.

Nothing model-specific, no `model.io`, no branch on package kind: the
only
questions asked are *does this workflow declare a continuation* and
*does this
request carry a session*.

## Three layers, each changed where it was wrong

**Metadata** — `SessionLeaseContract.continuation`, and validation that
fails the
document closed rather than the third turn. A continuation must be
session-scoped, `class: semantic`, `management: runtime`,
`release_boundary: session`, growing, and not also loop-carried (two
answers
about one value); it must name a declared `prompt_tokens` input and a
declared
`tokens` output whose contracts match; a workflow declares at most one,
because a
package has one conversation. Separately, **a session-scoped cell
binding a
`service_group` must resolve to a declared group that aliases it** — the
"claims continuity with no valid state groups/aliases" case.

**Runtime** — the conversation is bound before the pass and published
after it,
read from whichever way the pass published its tokens (`tokens`,
`tokens.row.N`,
or streamed `tokens.<n>` events; the real package publishes
`tokens.row.0`, which
is the bug the first draft of this had). Sessions are keyed and isolated
by id.
`reset_session` now works for an interpreted package and releases the
lease;
`close_session` already forgot it. Seeding no longer **errors** when a
request
carries no session — before this, a package could not declare session
state *and*
answer `Engine::generate`, which is a conversation's own first turn.
`session_token_count` reports the conversation rather than generated
tokens
alone, and `Engine::session_conversation` reports what it holds.

**Fail closed** — `create_session` refuses a package that publishes a
token
stream and declares no session state the interpreter can carry, naming
what the
package must declare. Against the old published metadata:

```
this package publishes a token stream but declares no session-scoped workflow
state, so a session could not continue a conversation: every turn would restart
from its own prompt. Declare the conversation in `pipeline.workflow.state` with
`scope: session` — either as a cell the loop carries, or with a
`session.continuation` naming the prompt input it rejoins — and re-validate the
package.
```

A package that publishes no token stream — speech, diffusion, video,
codec — has
no conversation to lose and keeps its session handle, which is what
`sessions_open_and_close_for_an_interpreted_package` already pins.

## The audit, and why the producer is not the site

Every `inference_metadata.yaml` the migration published, at its pinned
revision:

| package | onnx components | state groups | `scope: session` cells |
|---|---|---|---|
| `qwen3-0.6b-onnx-genai` | 11 | `decoder_cache` | **0** |
| `qwen2.5-0.5b-instruct-onnx-genai` | 11 | `decoder_cache` | **0** |
| `qwen2.5-14b-instruct-int4-zp-onnx` | 11 | `decoder_cache` | **0** |
| `deepseek-r1-distill-qwen-1.5b-onnx-genai` | 11 | `decoder_cache` |
**0** |
| `…-mistral-7b-v0-1-sliding-window` | 11 | `decoder_cache` | **0** |
| `…-qwen2-5-0-5b-{portable-f32,cuda-gqa-f16,static-cache-f32}` | 11 |
`decoder_cache` | **0** |
| `…-qwen2-5-1-5b-lora-selection` | 11 | `decoder_cache` | **0** |
| `…-qwen3-5-0-8b-hybrid-vlm-f32` | 14 | `decoder_cache` | **0** |
| `Muse-Glimmer-30B-ONNX-INT4-CUDA` | 14 | full+sliding | **0** |
| `moshiko-full-duplex-onnx-catalogue` | 23 | `temporal_cache` | **72**
|

Systematic — and **not from this repo's producer**. `decoder_workflow`,
which
backs both `migrate_model_io` and the `genai_config.json` importer,
emits the
`binding` token policy of rule 4, so its packages load onto the fused
decode core
whose paged KV sequence *is* the conversation; they continue one without
declaring one, and `multi_session.rs` / `one_runtime_e2e.rs` still pass
unchanged. The affected packages ship ONNX policy graphs instead and
have no such
core. So the fix lands in the schema, the runtime, and the package
metadata —
and in `tests/fixtures/onnx_genai_workflows/decoder`, the in-tree copy
of exactly
the affected shape, which now declares its conversation (+26 lines, no
other
edit).

## Published package

`justinchuby/qwen3-0.6b-onnx-genai`

* old `a89c02d343fc2b3c49d61b33d61f686698be1e4d`
* new
[`24da071476fdd1abaf7bb514424af8bc6604827f`](https://huggingface.co/justinchuby/qwen3-0.6b-onnx-genai/commit/24da071476fdd1abaf7bb514424af8bc6604827f)

A pure addition: one state cell and `session_state_lease` in the
manifest.
Diffed section by section against the old revision — the graphs, ports,
inputs,
outputs, state groups, serving contract, loop and **every existing state
cell**
are unchanged. `validate_metadata` passes on the full package, and the
revision
was re-downloaded from the hub and executed before this PR was opened.
No package
id appears anywhere in the code.

## Evidence

Against the **published new revision**, downloaded fresh
(`qwen3_0_6b_multi_turn_session.rs`, gated on
`ONNX_GENAI_QWEN3_WORKFLOW_DIR`):

```
turn 3      [1084, 594, 264, 1602, 2989, 949, 315, 847]
single shot [1084, 594, 264, 1602, 2989, 949, 315, 847]   ← equal
cold        [576, 3491, 374, 311, 1477, 279, 1372, 315]   ← what turn 3 used to be
```

`session_token_count` 16 → 42 across three turns of 8 tokens (34
conversation
tokens plus turn 3's 8). Independent session's first turn equals the
cold
request; `reset_session` returns the count to 0 and the next turn equals
the cold
request again.

Hermetic, in CI:

* `workflow_session_continuation.rs` (6 cases, Engine): turn 3 against
the
single-shot conversation; the conversation read back cell-by-cell after
each
turn — this fixture's synthetic weights decode a constant, so its
evidence is
  *what the decoder was asked to decode*, not what came out; isolation;
reset/close; stateless unchanged; and a package with the conversation
stripped
  out refusing a session.
* `chat_completions_continue_a_conversation_across_requests` and
  `chat_completions_without_a_session_stay_stateless` (server): three
  `POST /v1/chat/completions` sharing an `X-Session-Id`, with
`session_token_count` strictly increasing and turn 3's context at least
three
  times turn 1's; a different id starting its own conversation; `DELETE
  /v1/sessions/{id}` releasing it and a second delete returning 404.
* `session_continuity.rs` (10 cases, metadata): every fail-closed rule
above.

## Validation

`cargo fmt --all --check`; `clippy -D warnings --all-targets` on
metadata,
engine, server, CLI and C API. **onnx-genai-metadata** 23/23 suites,
**onnx-genai-engine** 478 lib + 75 integration binaries,
**onnx-genai-server**
243 lib + 40 http, CLI / C API / genai-config green. The JSON schema is
regenerated and `schema_sync` passes.

---------

Signed-off-by: justinchuby <justinchu@microsoft.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants