Repository navigation
feat(engine): one interpreter executes every package's declared workflow - #1723
Conversation
Make `EngineDecodeBackend::Native` a first-class backend for the one universal `pipeline.workflow` interpreter instead of another parallel native workflow. The interpreter (SSA, loops, branches, emits, loop-carried and session state, adapters, checkpoints) is reused unchanged; only *component execution* is now backend-neutral. What changed - Add `PipelineEngine::invoke_onnx_component`, a backend-neutral seam the node interpreter calls for every ONNX component. The ORT branch is behavior-preserving (it keeps the `Session::run` and the CUDA-graph / `IoBinding` stable-island fast path); the native branch runs a pure-Rust `InferenceSession`. - Add `pipeline::native_component` (feature `native-backend`): one native session per declared component, loaded once and reused across recurring loop edges, bridging the interpreter's `onnx_genai_ort::Value` currency to/from native `Tensor` with a faithful, fail-closed dtype mapping. No tensor is routed through the host-only `ComponentTensor` seam. - Flip `validate_pipeline_backend_request` to accept Native when the native backend is compiled in, and fail closed (no silent ORT fallback) otherwise. Execution islands (an ORT CUDA-graph optimization) are not planned under Native, so the interpreter drives individual component nodes correctly but unfused. - Remove the redundant host-only `ComponentSession`/`ComponentTensor` (onnx-genai-metadata) and `OrtComponentSession` (onnx-genai-ort). They had no callers and duplicated the component-invocation concept this seam now owns device-resident (Rule 3, Rule 10). Why - `Value` is already a device-capable handle, so keeping it as the neutral currency lets native tensors and loop-carried/shared state stay backend/device-resident without a host round-trip, and reuses the whole interpreter rather than cloning workflow logic. Autoregressive decode runs through the one interpreter loop (the decoder Loop over policy components), reusing the single sampling/stopping/commit policy — no third token loop is added. Tests (feature `native-backend`) - `native_workflow_parity`: same package run under ORT and Native with identical outputs for single-pass, nested loop+branch, loop-carried state, read-only shared session state, diffusion loop, and autoregressive decode; a native-used counter proves native sessions ran (not an ORT fallback); a routing test proves recurring loop edges reuse resident native sessions; a fail-closed test proves an unsupported op errors with an actionable diagnostic instead of silently falling back to ORT. Design note and staged follow-up boundaries (CUDA zero-copy residency, specialized AR executor, direct-Engine facade, Gemma4 target+assistant) are in docs/architecture/NATIVE_WORKFLOW_BACKEND.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
… feat/native-workflow-backend
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1723 +/- ##
==========================================
+ Coverage 80.21% 80.62% +0.41%
==========================================
Files 413 413
Lines 202063 203061 +998
Branches 202063 203061 +998
==========================================
+ Hits 162091 163728 +1637
+ Misses 34435 33789 -646
- Partials 5537 5544 +7
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
… plan Phase 0 of collapsing the two execution engines into the one canonical `pipeline.workflow` runtime (see the new docs/architecture/WORKFLOW_RUNTIME_UNIFICATION.md). The `Engine` session/decode model is already unified across ORT and native: ORT vs native is a `DecodeLoopBackend` choice beneath the single `run_decode_loop` policy, and the `*_native_*` session methods were already deprecated shims that delegate to the backend-dispatched `create_session`/`close_session`/`generate_in_session`/`rewind_session_by`. Per Rule 3 (no compatibility shims for pre-release APIs), delete them: - `Engine::create_native_session` -> `create_session` - `Engine::close_native_session` -> `close_session` - `Engine::generate_native_in_session[_with_callback]` -> `generate_in_session[_with_callback]` - `Engine::rewind_native_session` -> `rewind_session_by` The internal `generate_native_in_session_with_callbacks` (the real native executor beneath the unified session path) is retained. The only external caller, `onnx-genai-bench`'s `multiturn`, is migrated to the unified API. Proven behavior-preserving: `generate_in_session_with_priority_and_callback` already dispatches to the native executor when `decode_backend == Native`; the native session tests (`native_engine`, `multi_session`) already use the unified API and pass, and the native workflow parity/smoke tests still pass. Also documents the full sole-runtime plan (one interpreter, one state/session model, one decode policy, ORT/native executors beneath) with a phased, independently-testable deletion order, and normalizes fmt drift inherited from the merged main under the pinned toolchain. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
… feat/native-workflow-backend
Speculative decode expressed as a workflow (draft → verify → accept/reject → correction) produces identical output on ORT and native. The accept/reject/ rollback semantics live in the backend-agnostic interpreter; the native executor only runs component forward passes and never re-derives speculative or state-transition semantics. This is the single-parity-case pattern the Gemma4 target+assistant speculative workflow (onnx-genai #1716/#1696) will rely on. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…ase 1 Record the concrete Gemma4-driven requirement from onnx-genai #1696/#1716: the chained speculative proposer (folded_carry_output or recurrent) must be lifted from the direct-Engine Rust propose() (speculative/mod.rs:1273-1356) into a backend-agnostic interpreter construct keyed by SpeculativeProposalExecution:: Chained, so examples 22 (recurrent) and 24 (folded, cacheless drafter) become single ORT/native parity cases with only the per-step forward pass per-backend. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…oint Sibling (#1716) confirmed REPLACE, single implementation: no dependency on the old SpeculatorConfig / SharedKvProposer::propose path for examples 22/24. Record the two cross-PR Phase-1 preconditions: parent ratifies "replace + delete the old proposer path" and assigns single interpreter-seam ownership; and the pre-existing gates gemma4_assistant_full / chained_proposer_real are re-pointed at the interpreter Chained construct so the deletion drops no coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…dependency gate Parent ratified REPLACE + DELETE (no facade/no delegator) and assigned the full consolidation to this PR. Record the measured execution gate: Phases 1-4 are blocked on #1716 landing on main (a ~50k-insertion, still-evolving branch that edits the interpreter seam this PR owns) and on an executable chained-proposal fixture (none exists yet; gemma4-real-packages' package not built). Note that the AR/chained interpreter node is a real Engine-decode-core integration (SessionDecodeLoopBackend is Engine-tied), so it must land after #1716 to avoid redo against its pipeline/* seam edits. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
… + fail-closed bridge Rubber-duck of #1723 found two substantive, unblocked blockers; fix both now. Blocker 1 — no ORT sessions under Native. PipelineEngine::build built an ORT Session for every component even under EngineDecodeBackend::Native, pulling in the ORT runtime, double-loading weights, misreporting the EP, and letting ORT reject a native-only graph. Now build components with PipelineModels::load_with_ort_session_filter(.., |_| false) under Native: zero ORT sessions, the package's I/O contract stays available as backend-neutral graph_io_metadata, and execution_provider_status() reports the real native device. Tests: native_backend_builds_no_ort_sessions (no ORT sessions + native run + native-cpu EP), ort_backend_builds_ort_sessions (contrast). Blocker 2 — explicit device/provider + fail-closed residency. native_component no longer calls InferenceSession::load (auto-detects a CPU EP, ignores the requested device). It resolves the device via resolve_native_decode_device( config.native_device, ..) and builds each session with the explicit EP (CpuExecutionProvider / CudaExecutionProvider), mirroring native_decode/load.rs; CUDA without the feature and plugin EPs fail closed. The Value<->Tensor bridge gates on residency: host values/tensors round-trip through bytes (no device round-trip on CPU); a device-resident value/tensor fails closed with an actionable diagnostic instead of reading device memory as host bytes (Tensor::as_bytes is host-only) or silently host-copying. Root-caused: the native runtime has the device-tensor APIs, so the zero-copy device bridge is a wiring task (boundary A), not a fundamental CPU-only limit. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…H200-verified) Adds a native-cuda-gated test proving, on real CUDA hardware: the Native backend resolves and reports its CUDA device (Blocker 2 device/provider propagation, not an auto-detected CPU EP), still builds zero ORT sessions (Blocker 1), and at a device-resident boundary either bridges or fails closed with an actionable native diagnostic — never panics or silently host-copies. Verified passing on an H200 (cargo test --features native-cuda). The CUDA code path (CudaExecutionProvider::initialized + DevicePreference::Gpu) also compiles clean under --features native-cuda. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Rubber-duck blocker 2 required a real device-resident bridge, not a fail-closed intermediate. Implement it end-to-end behind `native-cuda` so a multi-component native CUDA workflow keeps intermediate and recurring/loop-carried/state tensors on the device with no host round-trip on recurring edges. onnx-genai-ort: - Add guard-carrying `Value::from_external_memory_with_owner` so a `Value` OWNS the native `Tensor`/`DeviceIoBinding` behind the device pointer it wraps. `TensorBacking::External` gains an `Option<Arc<dyn Any + Send + Sync>>` owner, freed by `Value::Drop` AFTER the `OrtValue` is released (no leak, no use-after-free). - `try_alias_clone` now aliases an owned `External` value by sharing the owner `Arc`, so a device value flows through the interpreter's `clone_value` fast path instead of failing a host read. engine (pipeline/native_component.rs): - CUDA path binds device-resident inputs zero-copy via `ExternalMemorySpec` + `device_binding_from_external_memory`, binds a device output buffer (`allocate_device_output_binding`) for every output whose concrete shape the interpreter resolved from bound input symbols, runs via `run_with_device_bindings`, and republishes each device output as an owning `Value`. Dynamic-shape outputs fall back to host materialization. Cross-session stream correctness via `Tensor::sync()` / allocator sync. Device/EP stay explicit; genuinely unsupported device/provider still fails closed. - Track `(device_input_bindings, device_outputs)` counters exposed via `native_device_residency_counts()` to prove no host round-trip. engine (pipeline/workflow.rs): - `resolve_component_output_shapes` resolves output contract shapes from the invocation's bound symbols; threaded into the native seam. ORT path ignores it. tests: - `native_cuda_device_resident_multicomponent` (H200): ORT/native bit-for-bit parity on a two-step static-cache decode AND non-zero device-residency counters proving the recurring KV stayed on-device. - Pin the CPU native device in the backend-agnostic parity helpers so they are deterministic under `native-cuda`; add `native_cuda_engine`. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Make the smoke run deterministic regardless of build features or a GPU being present (device-residency is covered in native_workflow_parity), consistent with the CPU-pinned parity helpers. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…he gate gemma4-real-packages committed the tiny executable gemma4_chained target+assistant fixture on the #1716 schema (branch justinchuby/gemma4-chained-fixture, commit c9c62bc), reusing the proven tiny-gemma4-assistant graphs. Static contract review (this PR): accepted, no regen — ABI/state_service/speculative match the agreed chained contract. So the second execution-gate condition (an executable chained fixture) is now satisfied; only "#1716 on main" remains. Record the fixture, the Phase-1 wire-in plan (required ORT/native + native-CUDA device-resident parity case in native_workflow_parity.rs), and the one open contract item flagged to gemma4-metadata-audit: the source-vs-destination meaning of port_bindings.target_hidden_context for the folded-carry case (carry_0 is the target hidden_states.0 [H], not inputs_embeds [2H]). No code change; consolidation remains gated on #1716 landing on main. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
gemma4-real-packages confirmed from the fixture graph ground truth: carry_0 = target.hidden_states.0 (H=16); each proposer step inputs_embeds[2H] = concat(embed(last_token)[leading H], carry[trailing H]) (order pinned by the assistant Slice starts=[16] ends=[32]); input_embedding.f32 = [vocab=32, hidden=16]. steps = single-pass body, speculative: = loop driver. Record the coverage boundary: this tiny drafter reads only the carry half (embed[0:16] is present for the 2H ABI but unused), so the fixture isolates folded-carry threading + borrowed-KV + rollback and does NOT validate embed-gather correctness (Phase 1 covers that separately). Narrow the open items to the two gemma4-metadata-audit field rulings: (1) direction of port_bindings.target_hidden_context; (2) where the interpreter sources the embedding table for real models (shared_weights omitted). Both remain the field owner's call; consolidation stays gated on #1716 landing on main. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…fields resolved gemma4-real-packages updated the fixture to the final #1716 schema (commit 8a66e2c, merging #1716 tip 9f47bb6) and made the folded carry explicit. Both previously-open contract items are now RESOLVED by ratified, mandatory proposal_execution fields (verified in crates/onnx-genai-metadata/src/validation.rs, required whenever folded_carry_output is set): - folded_carry_seed {component, output} names carry_0's source (fixture: target.hidden_states.0) — supersedes the target_hidden_context ambiguity. - token_embedding {component, table} names the embed-gather source (fixture: target.hidden_table) — generalizes to real models, no model-name gate. The Phase-1 chained driver reads both fields directly (no convention inference). Bump the recorded fixture SHA and point Phase 1 at the resolved fields. Consolidation stays gated on #1716 landing on main. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…ined contract
gemma4-real-packages published the faithful E2B fix (parent kept the
schema strict: folded_carry_seed names a real target output, no
request-input escape hatch). mobius gemma4 now emits the post-final-norm
hidden as hidden_states.{idx} (PR #546 @ 710d4927, backward-compatible),
and the real E2B packages carry the SAME folded_carry_seed/token_embedding
contract at scale with real ports (target hidden_states.34,
model.embed_tokens.weight [262144,1536] fp16).
This empirically confirms the Phase-1 field-reading chained driver
generalizes from the tiny fixture to real models with no model-name gate.
Record the real-package coordinates as an optional post-parity scale case;
the tiny gemma4_chained @ 8a66e2c stays the required hermetic parity
fixture (unchanged, revalidated). No seam change; still gated on #1716.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…backend Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com> # Conflicts: # crates/onnx-genai-engine/src/pipeline/mod.rs
Tiny target+assistant Gemma4 fixture for native_workflow_parity.rs, driven by the onnx-genai #1716 `proposal_execution: chained` contract (metadata-driven), not the legacy SharedKvProposerConfig. - Reuses the proven tiny-gemma4-assistant graphs (hidden 16, 2H=32, vocab 32, kv_heads 2, head_dim 8; target layer 0 = sliding owner, layer 1 = full owner). - Assistant is a borrowed-KV drafter: inputs_embeds [b,q,2H] + read_only shared_kv.{full,sliding}_attention.{key,value} -> logits + projected_state (folded carry); no own KV, no present.*. - Combined inference_metadata.yaml: components {target, assistant}; assistant read_only aliases keyed on the owner state cell (no output/no layer); proposal_execution {kind: chained, token_embedding_input: inputs_embeds, logits_output: logits, folded_carry_output: projected_state}; port_bindings.target_hidden_context: inputs_embeds; shared_state [full_attention, sliding_attention]; vocabulary identical; distribution_preserving: true; rollback_state = 4 target KV cells. - Greedy / logits-emit (external argmax); standard ops only (ORT and native CPU EP); FLOAT graphs. No RNG uint64 Cast. Rust-valid: `cargo run -p onnx-genai-metadata --bin validate_metadata` passes. Both graphs execute on ORT CPU with the declared contract shapes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add the explicit folded-carry fields introduced by #1716 @ 9f47bb6 to the chained proposal_execution, using actual graph names (no placeholders): - folded_carry_seed: {component: target, output: hidden_states.0} (the tiny target exposes hidden_states.0 = Gather(hidden_table, input_ids)) - token_embedding: {component: target, table: hidden_table} target_hidden_context/folded_carry_output retained. Rust validate_metadata (built from the merged #1716 tip) + python jsonschema both pass; both graphs still execute on ORT CPU with the declared shapes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…terpreter
A `chained` proposer emits one distribution per invocation, so materializing a
proposal block means driving the same typed component repeatedly while threading
its recurrence forward. That loop lived in a direct-`Engine` Rust `propose()`
bound to ORT `Session`s, which made it a second execution engine the
backend-neutral component seam could not drive.
It now lives in `pipeline::speculative`, on the interpreter's own seam:
* `PipelineEngine::propose_chained` drives the chain through
`invoke_component_values` -> `invoke_onnx_component`, so ORT and native run the
identical loop with no second "invoke a component" implementation.
* Everything the loop needs is read from `speculative.proposal_execution`:
`token_embedding_input`, `logits_output`, `recurrent[]`, `folded_carry_output`,
`folded_carry_seed` ({component, output}), and `token_embedding`
({component, table}). No model name, port spelling, or shape heuristic.
* `accept_chained_proposal` and `rollback_speculative_state` complete the round:
the accepted prefix plus the target's free correction, and truncation of every
declared `rollback_state` cell along its serving group's own `sequence_axis`.
The folded carry owns no state cell and is deliberately never restored.
* The embedding half of the fused input is gathered from the initializer the
contract names, read out of the target artifact.
Two interpreter facts had to become explicit for this to work at all:
* `run_pipeline_retained` keeps the pass's whole SSA map. A proposal binds the
proposer's borrowed read-only shared KV and masks from the values the workflow
itself bound, rather than a caller's reconstruction of them.
* Island fusion elides values no later step reads, which swallowed exactly the
tensors a proposal needs. `plan_execution_islands` now takes the externally
used set, computed from the declared contract, so a speculative package keeps
its folded-carry seed and proposer bindings live (and a package without a
chained proposer contributes nothing).
Ports sharing the fused input's position symbol narrow to the step's single
position, derived from the declared port contracts, so a borrowed-KV drafter's
`kv_sequence`-keyed mask is left alone while its `sequence`-keyed position ids
are not.
Proven on the hermetic `gemma4_chained` package (imported here from
gemma4-real-packages @ 8a66e2c): contract reading, chain driving, embedding
gather against the package's own second copy of the table, acceptance,
rejection + rollback, width bounding, and a full propose/verify/accept/rollback
decode that reproduces plain greedy decoding token for token.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
… its gates The interpreter now owns chained proposal driving, so the direct-`Engine` shared-KV proposer is a second implementation of the same contract. Rule 3: it is deleted, not kept behind a facade. Removed: `SpeculativeMode::SharedKv`, `SharedKvProposerConfig`, `SharedKvBinding`, `validate_shared_kv_proposer_config`, `SharedKvProposer` and its `propose()`, `SharedKvProposerModel`, `load_shared_kv_proposer`, the paged-KV slicer `shared_kv_slices_from_materialized`, and the native `NativeSharedKvProposerModel` / `NativeSpeculationKind::SharedKv` dispatch — which was already unreachable, since `native_shared_kv_proposer` was hard-wired to `None` at load and its arm could only ever return "requested without a loaded proposer session". Coverage moves rather than disappearing. Both legacy gates asserted the same property — speculative decoding is token-identical to plain greedy — through the deleted KV slicer: * `gemma4_assistant_full` -> `gemma4_chained_workflow::speculative_decode_equals_greedy_decode` (plus the ORT/native/CUDA parity case), on the same tiny graphs driven through the declared `proposal_execution: chained` contract. * `gemma4_assistant_mixed` -> `gemma4_chained_workflow::mixed_head_dim_speculative_decode_equals_greedy_decode`, on a new hermetic `gemma4_chained_mixed` package built from the same heterogeneous graphs (layer 0 sliding head_dim 8, layer 1 full head_dim 16). The property is stronger where it now lives: there is no global KV geometry to get wrong, because each shared-KV group's ports declare their own head width and the interpreter binds them straight through. Also fixes two merge-fallout parity cases against the #1716 tip: the authored nested-loop package must declare the `linear_effects` capability it uses, and the `speculative` fixture now has an asymmetric `verifier.past_key_values.0.value`. Verified: `gemma4_chained_workflow` (8), `native_workflow_parity` (10 on native-backend, 12 on native-cuda incl. the H200 device-resident chained parity case). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
With the interpreter owning chained proposal driving, nothing reaches the
supporting legacy surface, so it goes rather than lingering as a facade:
* `onnx-genai-ort::shared_kv_proposer` — `SharedKvProposerSession`, its
signature detection, and its step contract. Its only caller was the deleted
`SharedKvProposer`.
* `onnx-genai-ort::gemma4_assistant` — the model-name-gated Gemma4
`Gemma4SharedKvSpec`/`Gemma4AssistantSignature` path. It was an orphan file,
never declared as a module and therefore never compiled; the chained contract
supersedes it with no model-name conditional at all.
* `native_decode::proposer` (`NativeProposerSession`) and
`NativeDecodeSession::shared_kv_inputs`, plus the two unit tests pinning them.
* `SpeculativeProposerContext::shared_kv_slices` / `CandidateProposalInputs`, a
field every remaining caller now passes as `None`.
* `SharedKvProposerSpec` and `resolve_shared_kv` in the metadata parser.
A legacy HuggingFace `proposal_type: shared_kv` block still parses — configs in
the wild carry it — but now degrades to `Unknown` with a diagnostic naming the
contract that replaced it (`speculative.proposal_execution: {kind: chained}`),
so a package author is pointed at the right block instead of at a resolution
path nothing can run. The parser's six shared-KV resolution tests collapse into
one that pins exactly that.
Verified: workspace builds clean (the `onnx-genai-bench` compare warnings are
pre-existing on main); `onnx-genai-engine` lib 568, `onnx-genai-metadata` lib 19,
`native_workflow_parity` 10, `gemma4_chained_workflow` 8, `native_workflow_smoke`
2, `onnx_genai_workflow_conformance` 14, `workflow_policy_e2e` 20,
`static_cache_workflow` 7, `native_speculative_driver` 3, `native_engine` 14.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`items_after_test_module` fires on `debug_shapes_enabled`, which #1716 left below `captured_step_retry_tests`. It blocks `cargo clippy --all-targets -p onnx-genai-ort -- -D warnings`, so move the function above the module rather than leave the gate red for the next change. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…mains The plan document was written as a gate; both gate conditions are now met and Phase 1 is code, so it records what executes rather than what is proposed: * a table mapping each `speculative.proposal_execution` field to the driver's use of it, so the "no model-name gate" claim is checkable field by field; * the coverage boundary the tiny drafter imposes (it slices only the carry half, so the parity case cannot prove embedding-gather) and the separate test that does prove it; * the exact deletions, and the gates each was re-pointed to; * a new §6 listing the remaining work by symbol — the decode core to lift (`run_decode_loop` / `DecodeLoopBackend`, tied to `Engine` state), the missing plain-decoder `WorkflowSpec` synthesizer, the server's `if handle.pipeline` dispatch sites, the bench `--pipeline` mode, and the `SpeculativeProposer` implementations that still need to become contract-keyed interpreter constructs. `NATIVE_WORKFLOW_BACKEND.md` boundary D is marked done, and the CLI capability inventory no longer lists a `SpeculativeMode::SharedKv` that does not exist. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
… real-package gate Three things the chained lift depends on that nothing was proving: * **Liveness is scoped.** `externally_used_values` force-keeps values so island fusion cannot elide a proposal's inputs. If that ever stopped being conditional on a *chained* contract, every workflow in the repo would silently lose fusion with no test failing. Two unit cases pin it: no contract and a `block` proposer both contribute nothing, and a chained contract keeps exactly its own bindings plus the folded-carry seed — not the target logits the workflow emits. * **The folded-carry contract is required, not inferred.** A folded carry with no `folded_carry_seed`, no `token_embedding`, or a chain with neither a recurrence nor a folded carry, must fail to resolve with a diagnostic naming the missing field — otherwise a runtime would be free to guess, which is the behavior this whole design removes. * **Real packages, not just the fixture.** `chained_proposer_real` gains an opt-in `ONNX_GENAI_CHAINED_WORKFLOW_PACKAGE` case that resolves a real workflow package's chained contract through the interpreter and gathers its declared embedding table from the real target artifact — the check a per-model heuristic would pass vacuously and a contract read cannot. The published Gemma4-E2B speculative package carries the identical field shape, so the same driver serves both. Verified against the hermetic package (`REAL_CHAINED_WORKFLOW_EVIDENCE proposer=assistant target=target width<=4 rollback_cells=4`); it skips cleanly when unset. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
The user requirement is one runtime and no caller-side dispatch. Two public
runtime types made that impossible by construction: every caller had to decide
which one it held, and each decided separately.
`PipelineEngine` is gone. The workflow interpreter is now internal state
(`pipeline::WorkflowRuntime`, `pub(crate)`) held by the one public `Engine`, and
the operations that used to exist twice exist once:
* `Engine::from_dir` resolves the package shape itself
(`PipelineModelDirectory::load_if_declared`) and loads either half. Callers
stop running that probe and picking a constructor — the server, the CLI, and
the benchmarks each used to do it separately.
* `Engine::generate*` is one entry point: a package declaring
`pipeline.workflow` is driven by the interpreter over its declared `tokens`
output, a bare decoder by the decode core. The caller does not choose.
* `tokenize`, `execution_provider_status`, `effective_max_context`,
`resource_snapshot`, `governor`, and the island/adapter diagnostics all resolve
the same way, so a report reads the same regardless of package shape.
* Exactly one resource governor per runtime. `Engine::governor` is `Option`
precisely so a workflow package does not get a second one that would
double-count every reservation; the accessor returns whichever half owns it.
* `Engine::tokenizer` is `Option` because a workflow package may legitimately
ship none (an image-generation pipeline). That is a load-time fact with an
actionable error, not a panic in the decode loop.
Server, CLI, C ABI, benchmarks:
* `EngineBackend::{Single,Pipeline}` in the server driver collapses to one
owned runtime, and `run_engine_driver` asks `engine.is_workflow()` instead of
destructuring a variant.
* `ModelHandle::pipeline` is deleted. The completions route's four
`if handle.pipeline` branches become one `submit_generation` call; the
sessions and admin routes ask the runtime. `EngineDriver::warmup` no longer
takes a caller-supplied `pipeline: bool` — it reads the flag it derived from
the engine at startup, so a route and the driver cannot disagree.
* The CLI, `run_comfyui`, `profile_native`, the workflow example, and 40+ tests
now name one type.
Dead API removed rather than left as a second surface: `WorkflowRuntime`'s
`from_dir`, `generate`, `detokenize`, `resource_snapshot`, `device_authority`,
`set_vram_limit`, and `clear_prefix_cache` are superseded by `Engine`'s and are
deleted.
New `one_runtime_e2e` (5 cases) pins the property against regression to two
runtimes: one constructor loads both shapes, autoregressive text generation and
session lifecycle are one API each, shape-specific diagnostics degrade uniformly
rather than erroring, and tokenization is served from whichever half owns the
tokenizer.
Verified: full `onnx-genai-engine` suite under `native-backend` (lib 572 + every
integration target, zero failures), server 228 / CLI / C ABI / bench suites,
`cargo fmt --all --check`, and
`cargo clippy --locked --all-targets -D warnings` on engine, server, cli, capi.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…y case `Engine::models()` returns a `Result` now that one type serves both package shapes, and this assertion sits behind `--features native-cuda` so the non-CUDA build never type-checked it. Verified on H200: `native_workflow_parity` 12/12 twice in a row, including `chained_speculative_proposal_parity_native_cuda` and `native_cuda_device_resident_multicomponent`. (Running three CUDA test binaries concurrently on one device can make the residency case flake on contention; it passes deterministically when the file runs on its own.) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
… blocker Phase 3 is done and §6's caller items are struck through with what replaced them. §7 is new and states the blocker precisely, with the three facts that make it a decision rather than an implementation: * `validate_metadata` rejects `model.io` beside `pipeline.workflow` by ratified rule, and every bare-decoder package declares `model.io`; * `PipelineModelDirectory::load` errors without a `pipeline` section; * `decoder_io_is_legacy()` — surfaced on `/v1/debug/config` — flips for every existing package the moment one is synthesized. Three options are laid out with their costs, including the one that needs no schema change (a `binding` component carrying `contract.id: onnx-genai.autoregressive-decode`, dispatched the way adapters already are). The second, independent gate is recorded too: the bare-decoder generate path is covered by real-model tests that skip without weights, so the rewire cannot be proven green here. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…anonical workflow Parent decision A: keep authored bare-decoder metadata unchanged and valid, and lower it at load into an internal canonical WorkflowSpec. `onnx_genai_metadata::canonical` is that compiler. It is a pure function of the declared ABI -- same `ModelIoSpec` in, byte-identical document out -- which is what makes the lowered form *derived* rather than a second authored answer: it cannot drift from `model.io`, because it is recomputed from it on every load. Nothing writes it back, so `validate_model_io_against_workflow` never sees a pair and no published package needs re-authoring. It emits canonical YAML and parses it through the ordinary schema path. That is deliberate: the lowered form is exactly what an author would have had to write, it is diffable, and a field this module gets wrong is a loud error at lowering time instead of a subtly malformed spec the interpreter trips over later. It caught two on the first run -- a missing `increment` on `growing` recurrence and a missing `contract` on the loop induction value. No schema change. Both canonical components are `binding` components identified by contract id -- `onnx-genai.autoregressive-decode` for one decoder forward pass, `onnx-genai.token-policy` for next-token selection and stop detection -- dispatched the way workflow adapters already are. The published JSON schema is untouched. KV is declared `management: runtime`, the schema's existing word for "the runtime owns these buffers", so the paged / share-buffer / CUDA-graph executors keep owning their KV and no per-step host round-trip is introduced. `/v1/debug/config` gains `workflow_provenance` (authored | lowered | none). `pipeline` keeps its meaning -- does the package *serialize* a workflow -- so a lowered decoder reports `pipeline: false` with `workflow_provenance: lowered`, and no report claims the file contains something it does not. Tests: 7 unit cases (determinism, schema round-trip, KV becomes runtime-managed state, an authored workflow is never lowered beside itself, the package is never mutated and stays valid, unpaired KV refused, a missing port named) plus `canonical_lowering_corpus` against the real corpus through the engine's own loader -- 7 packages covered (phi35-mini int4 shared-buffer, qwen3-0.6b int4, qwen2.5-0.5b CUDA, qwen2.5-coder-7b, phi4-mini CUDA, gpt-oss-20b MoE, qwen3.5-2b) -- asserting deterministic lowering and that the lowered decoder's ports mirror the resolved ABI exactly: no port invented, none dropped. That corpus test immediately earned its keep. Routing `Engine::from_dir` through `load_if_declared` had turned a shape guess into a hard load error for a package recognized as multi-component that ships no native metadata. A resolution failure now means "not a workflow package" and the decoder loader handles it, as it always did. Greedy goldens over 5 real models (phi35-mini, phi4-mini CUDA, qwen2.5-0.5b CUDA, qwen3-0.6b, gpt-oss-20b) are byte-identical before and after. `.goldens/` holds the harness and the c58eb5b baseline so later work diffs against the previous runtime rather than only against itself. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
Delete temporary commit-message and local golden-evidence files that were unintentionally included with PR #1723. Runtime code, tests, and maintained documentation are unchanged. Signed-off-by: justinchuby <justinchu@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CI status, and what is and is not this branch's
The CUDA job's failing step is Reproduced in a clean What this branch's own CUDA surface does. The step that is about this work —
Fixed here after CI caught it. The first
One further pre-existing failure worth recording, unrelated to CI: |
Continuous batching and speculative decoding were the two loops #1723 left outside the interpreter. Both took their bound and their stop from Rust, and which one ran was decided beside the workflow rather than by it. They are now declared steps: a body naming `onnx-genai.continuous-batch` is run by the row-major batch executor, a body naming `onnx-genai.speculative-block` by the propose-verify-accept executor, and the interpreter's `Loop` owns the iteration in both cases exactly as it does for a single token. ## How a body is chosen `decoder_workflow::iteration_variant` re-authors a package's own declared generation loop into the body for a named `IterationPolicy`. The bound, the liveness predicate, the carried cells, the termination and the emits stay byte-identical; only the body's node changes. It sits in the module that authors the single-token body, so the three cannot drift into three statements of one loop. `WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is selected, and it selects by naming a policy. Nothing reads a package's decode path, its file layout or its model identity to decide which iteration runs. Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer, native CUDA persistent batch) is still a question about a resource this process holds -- the same family as `holds_decode_core()` -- and all three implement one `BatchedDecodeSession` behind the one authored body. ## A batch publishes the batch's liveness, not each row's The interpreter reads a loop's predicate *before* the body runs and preserves an inactive row's carried state afterwards, so a slot that published `false` cannot come back within the pass. A continuous batch recycles slots, so a per-slot bit would permanently retire a slot the scheduler is about to backfill -- dropping the backfilled request's tokens out of the declared stream and ending the pass while that request was still running. The batch answers one question for every slot: will this batch decode again? The per-row question the emit needs -- did this slot contribute a token this iteration -- is `accepted_len`, which is zero for a slot that sat the step out. `a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and fails with a per-slot bit. ## Measured `authored_iteration_executors.rs` drives one runtime as a single-token generation, a continuous batch and a speculative request, and reads the per-contract counter recorded inside the interpreter's node dispatch. The three produce **disjoint** contract sets. A batched generation that merely produced the right tokens would keep passing if Rust went back to picking the row-major loop from `ModelDecodePath::StaticCache`; a count against `onnx-genai.continuous-batch` cannot be produced by a shape `match`. `authored_iteration_e2e.rs` checks the executors still do the job: a batched row stops on a declared end token that is **not** the first listed, rows with different budgets finish independently, and a wider declared draft width proposes more candidates per block while the rejected suffix is rolled back -- with the stream byte-identical to plain greedy. The manager's scripted-decode unit tests now run on the authored body, so multi-EOS stops, per-row device/host sampling parity, priority admission and occupancy are exercised through the declared document. ## Deleted `generate_batched_static` had its own row loop *and* its own copy of the per-row token policy; a fixed batch is the continuous-batch algorithm with all-at-once admission, so it now runs the same authored body and the copy is gone (with it, the stale "past/present batching is deferred" refusal -- shared-buffer packages batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one `WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }` is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is `advance_block` -- both implementing the *same* contract, so the ORT and native paths are two implementations of one declared step rather than two loops. ## The emit carries a length now A batch row may sit out a step and a verified block may contribute several tokens, so a re-authored emit publishes a per-row prefix bounded by the step's declared `accepted_len`. That also makes the token output row-wise, giving each row its own accumulated stream. Each row's stream is held against what the executor committed for that row when the pass ends -- `emitted_token_rows` reads every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on the run-to-completion routes and in the server's driver, so the check runs where batches are actually driven. `tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with `aliasing: permitted`: one declaration, and the package offers the shared buffer a row-major batch decodes into. It exists so the authored batch has a hermetic package to run on. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Continuous batching and speculative decoding were the two loops #1723 left outside the interpreter. Both took their bound and their stop from Rust, and which one ran was decided beside the workflow rather than by it. They are now declared steps: a body naming `onnx-genai.continuous-batch` is run by the row-major batch executor, a body naming `onnx-genai.speculative-block` by the propose-verify-accept executor, and the interpreter's `Loop` owns the iteration in both cases exactly as it does for a single token. ## How a body is chosen `decoder_workflow::iteration_variant` re-authors a package's own declared generation loop into the body for a named `IterationPolicy`. The bound, the liveness predicate, the carried cells, the termination and the emits stay byte-identical; only the body's node changes. It sits in the module that authors the single-token body, so the three cannot drift into three statements of one loop. `WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is selected, and it selects by naming a policy. Nothing reads a package's decode path, its file layout or its model identity to decide which iteration runs. Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer, native CUDA persistent batch) is still a question about a resource this process holds -- the same family as `holds_decode_core()` -- and all three implement one `BatchedDecodeSession` behind the one authored body. ## A batch publishes the batch's liveness, not each row's The interpreter reads a loop's predicate *before* the body runs and preserves an inactive row's carried state afterwards, so a slot that published `false` cannot come back within the pass. A continuous batch recycles slots, so a per-slot bit would permanently retire a slot the scheduler is about to backfill -- dropping the backfilled request's tokens out of the declared stream and ending the pass while that request was still running. The batch answers one question for every slot: will this batch decode again? The per-row question the emit needs -- did this slot contribute a token this iteration -- is `accepted_len`, which is zero for a slot that sat the step out. `a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and fails with a per-slot bit. ## Measured `authored_iteration_executors.rs` drives one runtime as a single-token generation, a continuous batch and a speculative request, and reads the per-contract counter recorded inside the interpreter's node dispatch. The three produce **disjoint** contract sets. A batched generation that merely produced the right tokens would keep passing if Rust went back to picking the row-major loop from `ModelDecodePath::StaticCache`; a count against `onnx-genai.continuous-batch` cannot be produced by a shape `match`. `authored_iteration_e2e.rs` checks the executors still do the job: a batched row stops on a declared end token that is **not** the first listed, rows with different budgets finish independently, and a wider declared draft width proposes more candidates per block while the rejected suffix is rolled back -- with the stream byte-identical to plain greedy. The manager's scripted-decode unit tests now run on the authored body, so multi-EOS stops, per-row device/host sampling parity, priority admission and occupancy are exercised through the declared document. ## Deleted `generate_batched_static` had its own row loop *and* its own copy of the per-row token policy; a fixed batch is the continuous-batch algorithm with all-at-once admission, so it now runs the same authored body and the copy is gone (with it, the stale "past/present batching is deferred" refusal -- shared-buffer packages batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one `WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }` is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is `advance_block` -- both implementing the *same* contract, so the ORT and native paths are two implementations of one declared step rather than two loops. ## The emit carries a length now A batch row may sit out a step and a verified block may contribute several tokens, so a re-authored emit publishes a per-row prefix bounded by the step's declared `accepted_len`. That also makes the token output row-wise, giving each row its own accumulated stream. Each row's stream is held against what the executor committed for that row when the pass ends -- `emitted_token_rows` reads every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on the run-to-completion routes and in the server's driver, so the check runs where batches are actually driven. `tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with `aliasing: permitted`: one declaration, and the package offers the shared buffer a row-major batch decodes into. It exists so the authored batch has a hermetic package to run on. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…l-closed same-device checks Fix-forward revision of rejected PR #1854 (author Deckard locked out of this artifact). Addresses all 9 Cycle-17 blocking findings from Roy's review: 1. Range-level (not value-level) rollback tracking: if range N within one ValueId commits and a LATER range in that same value hits Fatal, the earlier committed range is now correctly included in the rollback. 2. BoundaryApplicationOutcome.fatal_progress surfaces committed_count and poisoned_range straight from TransitionOutcome::Fatal (no more '..' discard). 3. Rollback-of-rollback failures are now explicit: a new RollbackFailure struct (value, range, detail, committed_count, poisoned_range, quarantined) replaces the previous silent all_ok=false. 4. New deterministic PLAN-ENTRY fault injection: a test-only apply_residency_plan_at_boundary_with_phase8_faults entry point accepts a per-ValueId DriverFaultPlan map, enabling 4 new GPU tests that exercise partial-commit-then-fatal, rollback-of-rollback, same-device rejection, and expert-group atomicity end to end (not just at the primitive level). 5. Same-device fail-closed: check_same_device validates allocator.device_key() and both pools' device_ordinal against the requested device_ordinal before any mutation. 6. Corrected the module doc's false 'one production consumer called at model load' claim; #1723 is now honestly described as technical current-main context only, with zero production callers today. 7. Cross-ValueId atomic expert-group tiering: new graph-derived ExpertWeightGroup type + expert_weight_groups(graph) (onnx-runtime-ep-api), purely structural from QMoe/BlockQuantizedMoe node inputs (no tensor-name heuristics). apply_residency_plan_at_boundary now takes an expert_groups parameter and enforces all-or-none group fallback via a validation-only precheck phase that completes for an entire group before any member is mutated. 8. Renamed misleading telemetry: device_bytes_before/after -> single truthful device_bytes_released; COARSE_RESIDENCY_PROFILE_ENV -> COARSE_RESIDENCY_ENABLE_ENV (deprecated alias kept for compatibility). Also fixes a pre-existing (unrelated) Cargo feature-forwarding bug: onnx-runtime-ep-cuda's gpu-tests feature did not forward to onnx-runtime-cuda-memory/gpu-tests, hiding CudaVmmAllocator test-only items needed to compile these tests. Tests: 7/7 coarse_residency_plan_gpu (3 pre-existing + 4 new) pass on A100 (CUDA_VISIBLE_DEVICES=4), 4/4 coarse_residency_plan (non-GPU) pass, 3/3 + 91/91 unit tests (onnx-runtime-ep-cuda coarse_residency + onnx-runtime-ep-api) pass, 530/530 onnx-runtime-ep-cuda lib tests pass (no regressions), and all 29 Slice2-4 GPU regressions (composable_vmm_production_gpu, content_preserving_transition_gpu, expert_bank_remap_cost_gpu) pass unchanged. fmt/clippy clean on all touched crates. Independent code review (not Deckard) found no blocking issues. PR remains in draft; Roy's explicit re-review is still required before merge. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Continuous batching and speculative decoding were the two loops #1723 left outside the interpreter. Both took their bound and their stop from Rust, and which one ran was decided beside the workflow rather than by it. They are now declared steps: a body naming `onnx-genai.continuous-batch` is run by the row-major batch executor, a body naming `onnx-genai.speculative-block` by the propose-verify-accept executor, and the interpreter's `Loop` owns the iteration in both cases exactly as it does for a single token. ## How a body is chosen `decoder_workflow::iteration_variant` re-authors a package's own declared generation loop into the body for a named `IterationPolicy`. The bound, the liveness predicate, the carried cells, the termination and the emits stay byte-identical; only the body's node changes. It sits in the module that authors the single-token body, so the three cannot drift into three statements of one loop. `WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is selected, and it selects by naming a policy. Nothing reads a package's decode path, its file layout or its model identity to decide which iteration runs. Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer, native CUDA persistent batch) is still a question about a resource this process holds -- the same family as `holds_decode_core()` -- and all three implement one `BatchedDecodeSession` behind the one authored body. ## A batch publishes the batch's liveness, not each row's The interpreter reads a loop's predicate *before* the body runs and preserves an inactive row's carried state afterwards, so a slot that published `false` cannot come back within the pass. A continuous batch recycles slots, so a per-slot bit would permanently retire a slot the scheduler is about to backfill -- dropping the backfilled request's tokens out of the declared stream and ending the pass while that request was still running. The batch answers one question for every slot: will this batch decode again? The per-row question the emit needs -- did this slot contribute a token this iteration -- is `accepted_len`, which is zero for a slot that sat the step out. `a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and fails with a per-slot bit. ## Measured `authored_iteration_executors.rs` drives one runtime as a single-token generation, a continuous batch and a speculative request, and reads the per-contract counter recorded inside the interpreter's node dispatch. The three produce **disjoint** contract sets. A batched generation that merely produced the right tokens would keep passing if Rust went back to picking the row-major loop from `ModelDecodePath::StaticCache`; a count against `onnx-genai.continuous-batch` cannot be produced by a shape `match`. `authored_iteration_e2e.rs` checks the executors still do the job: a batched row stops on a declared end token that is **not** the first listed, rows with different budgets finish independently, and a wider declared draft width proposes more candidates per block while the rejected suffix is rolled back -- with the stream byte-identical to plain greedy. The manager's scripted-decode unit tests now run on the authored body, so multi-EOS stops, per-row device/host sampling parity, priority admission and occupancy are exercised through the declared document. ## Deleted `generate_batched_static` had its own row loop *and* its own copy of the per-row token policy; a fixed batch is the continuous-batch algorithm with all-at-once admission, so it now runs the same authored body and the copy is gone (with it, the stale "past/present batching is deferred" refusal -- shared-buffer packages batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one `WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }` is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is `advance_block` -- both implementing the *same* contract, so the ORT and native paths are two implementations of one declared step rather than two loops. ## The emit carries a length now A batch row may sit out a step and a verified block may contribute several tokens, so a re-authored emit publishes a per-row prefix bounded by the step's declared `accepted_len`. That also makes the token output row-wise, giving each row its own accumulated stream. Each row's stream is held against what the executor committed for that row when the pass ends -- `emitted_token_rows` reads every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on the run-to-completion routes and in the server's driver, so the check runs where batches are actually driven. `tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with `aliasing: permitted`: one declaration, and the package offers the shared buffer a row-major batch decodes into. It exists so the authored batch has a hermetic package to run on. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
…r the unified interpreter (#1883) ## Context Follow-up from the merged S3 capacity-emission slice (#1860): the user directed re-validating that work's integration against the newly-merged unified-workflow runtime (#1723, "one runtime, one interpreter, one drive"). ## What broke #1723 made package loading eagerly resolve **every** declared `pipeline.workflow` component's on-disk artifact at load time (`crates/onnx-genai-ort/src/loader.rs`). Two fixtures added earlier in #1832 — `tests/fixtures/tiny-deepseek-v4-qmoe/` and `tests/fixtures/tiny-glm52-full-attention/` — hand-authored a full workflow document referencing ten auxiliary policy components (`cache_length_update`, `token_sampler`, `termination`, `token_state_update`, `last_token_logits`, `decoder_state_initializer`, `decoder_step_update`, `termination_batch_initializer`, `token_to_slot`, `generated_length_update`) whose `policies/*.onnx` artifact files were never committed. Under the old canonical-dispatch loader this apparently went unnoticed; under the new interpreter it hard-fails: ``` Failed to resolve workflow package: IO error: workflow ONNX component file not found: .../tiny-deepseek-v4-qmoe/policies/cache_length_update.onnx ``` This broke 3/4 tests in `deepseek_v4_tiny_qmoe_e2e.rs` and 3/4 in `glm_tiny_full_attention_e2e.rs` — both are correctness gates for already-merged Sapper work (S3 KV capacity-emission, GLM-5.2 enablement). ## Fix Re-emitted both fixtures' `inference_metadata.yaml` with the repository's own `migrate_model_io --reemit` tool — the exact remedy #1723 used for the 14 fixtures it found in the same stale state (`tiny-llm`, `tiny-mtp-full`). Metadata-only: `model.onnx.textproto`, `manifest.json`, and `tokenizer.json` are byte-identical/untouched, so DeepSeek-V4's dense-CSA/QMoE graph and GLM-5.2's full-attention graph are completely unaffected. No Rust source touched. Confirmed the regenerated documents still: - declare `onnx-genai.autoregressive-decode` - use `bounded` (not `invariant`) sequence recurrence — the exact latent-bug class #1723 called out - correctly omit MTP (matching `num_nextn_predict_layers: 0` / no MTP field) - validate cleanly against the schema (`validate_metadata`) ## Verification (fresh `origin/main` @ `7a35dba77`, real A100 GPU) | Test | Result | |---|---| | `deepseek_v4_tiny_qmoe_e2e` | 4/4 pass, `captures=1 replays=10 fallbacks=0`, anchor IDs identical across native CPU/CUDA/stock-ORT | | `glm_tiny_full_attention_e2e` | 4/4 pass, anchor IDs identical across native CPU/CUDA/stock-ORT | | `glm_tiny_qmoe_native_cuda_e2e` | 4/4 pass (unaffected, sanity re-check) | | `deepseek_v2_tiny_qmoe_native_e2e` | 1/1 pass (unaffected, sanity re-check) | | `authored_body_selects_executor` | 4/4 pass | | `one_runtime_e2e` | 6/6 pass | | `native_workflow_parity` | 10/10 pass | | `native_workflow_smoke` | 2/2 pass | | `canonical_execution_parity` | 7/7 pass | Also ran `cargo test -p onnx-genai-engine --features native-backend` (full lib+test suite): 613/614 pass; the one failure (`native_generate_rejects_over_kv_byte_budget_before_backend_run`) is **pre-existing on main without this change** (confirmed via `git stash` bisection) — unrelated to this fix, out of scope here. Independent review: **APPROVE**, no blocking findings (contract/MTP/EOS semantics verified unchanged, `model.onnx.textproto` confirmed byte-identical). ## Scope Metadata-only fixture fix. No workflow-runtime refactor, no native-vs-ORT branching, no side decode loop, no changes to the S3 capacity-emission mechanism itself (`rewrite_kv_capacity_appends`) — that remains verified compatible with the unified `run_workflow_node`/contract-dispatch model as-is, since it operates purely at ONNX-graph-build time below the workflow-interpreter layer. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Sapper <sapper@squad.local> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Continuous batching and speculative decoding were the two loops #1723 left outside the interpreter. Both took their bound and their stop from Rust, and which one ran was decided beside the workflow rather than by it. They are now declared steps: a body naming `onnx-genai.continuous-batch` is run by the row-major batch executor, a body naming `onnx-genai.speculative-block` by the propose-verify-accept executor, and the interpreter's `Loop` owns the iteration in both cases exactly as it does for a single token. ## How a body is chosen `decoder_workflow::iteration_variant` re-authors a package's own declared generation loop into the body for a named `IterationPolicy`. The bound, the liveness predicate, the carried cells, the termination and the emits stay byte-identical; only the body's node changes. It sits in the module that authors the single-token body, so the three cannot drift into three statements of one loop. `WorkflowRuntime::iteration_runtime(policy)` is the single place an algorithm is selected, and it selects by naming a policy. Nothing reads a package's decode path, its file layout or its model identity to decide which iteration runs. Which decode *session* a batch runs on (ORT static-cache, ORT shared-buffer, native CUDA persistent batch) is still a question about a resource this process holds -- the same family as `holds_decode_core()` -- and all three implement one `BatchedDecodeSession` behind the one authored body. ## A batch publishes the batch's liveness, not each row's The interpreter reads a loop's predicate *before* the body runs and preserves an inactive row's carried state afterwards, so a slot that published `false` cannot come back within the pass. A continuous batch recycles slots, so a per-slot bit would permanently retire a slot the scheduler is about to backfill -- dropping the backfilled request's tokens out of the declared stream and ending the pass while that request was still running. The batch answers one question for every slot: will this batch decode again? The per-row question the emit needs -- did this slot contribute a token this iteration -- is `accepted_len`, which is zero for a slot that sat the step out. `a_slot_freed_into_an_empty_queue_is_reused_by_a_later_submission` pins it, and fails with a per-slot bit. ## Measured `authored_iteration_executors.rs` drives one runtime as a single-token generation, a continuous batch and a speculative request, and reads the per-contract counter recorded inside the interpreter's node dispatch. The three produce **disjoint** contract sets. A batched generation that merely produced the right tokens would keep passing if Rust went back to picking the row-major loop from `ModelDecodePath::StaticCache`; a count against `onnx-genai.continuous-batch` cannot be produced by a shape `match`. `authored_iteration_e2e.rs` checks the executors still do the job: a batched row stops on a declared end token that is **not** the first listed, rows with different budgets finish independently, and a wider declared draft width proposes more candidates per block while the rejected suffix is rolled back -- with the stream byte-identical to plain greedy. The manager's scripted-decode unit tests now run on the authored body, so multi-EOS stops, per-row device/host sampling parity, priority admission and occupancy are exercised through the declared document. ## Deleted `generate_batched_static` had its own row loop *and* its own copy of the per-row token policy; a fixed batch is the continuous-batch algorithm with all-at-once admission, so it now runs the same authored body and the copy is gone (with it, the stale "past/present batching is deferred" refusal -- shared-buffer packages batch). `ContinuousBatchManager::step`'s hand-rolled iteration is one `WorkflowGenerationCursor` advance. `generate_speculative_loop`'s `loop { .. }` is `advance_speculative_block`, and `NativeSpeculativeDriver::generate`'s is `advance_block` -- both implementing the *same* contract, so the ORT and native paths are two implementations of one declared step rather than two loops. ## The emit carries a length now A batch row may sit out a step and a verified block may contribute several tokens, so a re-authored emit publishes a per-row prefix bounded by the step's declared `accepted_len`. That also makes the token output row-wise, giving each row its own accumulated stream. Each row's stream is held against what the executor committed for that row when the pass ends -- `emitted_token_rows` reads every row rather than row 0, and `ContinuousBatchManager::drain` ends the pass on the run-to-completion routes and in the server's driver, so the check runs where batches are actually driven. `tests/fixtures/tiny-llm-batched` is `tiny-llm-sharedbuffer` with `aliasing: permitted`: one declaration, and the package offers the shared buffer a row-major batch decodes into. It exists so the authored batch has a hermetic package to run on. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
… the matrix (main is red) (#1888) `main` is red on a **required** check and has been since #1883, so nothing in the repository can merge. ``` the_matrix_holds_for_every_maintained_workflow tests/fixtures/tiny-deepseek-v4-qmoe/inference_metadata.yaml: graph component count left: 1 right: 11 ``` `crates/onnx-genai-metadata/tests/decoder_recognizer_agreement.rs` runs in `Fast (Linux x86_64)`, which with `Rust quality` is one of only two required status checks on `main`. Same class of repo-wide block as the `op_rules` catalog pin (#1860 → #1872/#1882): a pin that fell behind a deliberate change. ## Bisected | commit | result | |---|---| | `182d1f776` (#1864, parent) | **ok**, 14/14 | | `7adfae901` (#1883) | **FAILED**, 13/14 | | `7a0cb6c39` (tip) | **FAILED**, 13/14 | Both re-emitted fixtures are affected — DeepSeek-V4 fails first, and GLM-5.2 has the identical 11 → 1 collapse behind it. ## Why the fixtures are right and the matrix is wrong #1832 hand-authored both documents with ten auxiliary policy components (`cache_length_update`, `token_sampler`, `termination`, …) whose `policies/*.onnx` artifacts **were never committed**. #1723 made package loading resolve every declared component eagerly, turning that latent staleness into a hard load failure, and #1883 re-emitted both with `migrate_model_io --reemit` — the same remedy #1723 applied to the fourteen fixtures it found in the same state. So the ten components were never real. The re-emission is the fix; the matrix is what fell behind it. ## Delta accounted for, not made to match Each field justified from the re-emitted document rather than read off the failure: | field | was | now | why | |---|---|---|---| | components | 11 | 1 | the document declares a single `decoder`; the ten policies are gone because their artifacts never existed | | cardinality | `Composite` | `SingleGraph` | direct consequence of the above | | decoder | `Some("model")` | `Some("decoder")` | the re-emit tool names it `decoder`; the hand-written document said `model` | | layer 1 | `false` | `true` | exactly one decoder component, no competing roles | | layer 2 | `None` | `Some("decoder")` | both layers now resolve | ## Checked for silently deleted coverage first Updating a pin can quietly retire the only witness to a classification path, so I checked before touching the rows rather than after: - the `11` / `Composite` / `Some("model")` shape these two used to pin is **still pinned** by `tests/fixtures/onnx_genai_workflows/decoder/inference_metadata.yaml`; - **44 of 64** rows still carry the `false` / `None` two-layer split that motivates reporting the layers separately. No path loses its only witness here. That reasoning is recorded in a comment beside the rows so the next person does not have to redo it. ## Verification `cargo test -p onnx-genai-metadata` — **14/14** on the agreement suite, all 22 targets green. `cargo fmt --all -- --check` clean. `cargo clippy --locked --all-targets -p onnx-genai-metadata -- -D warnings` clean. Not my area — I found this because it blocked #1877. If the intent was for these fixtures to keep an eleven-component workflow, the correct fix is to commit the missing `policies/*.onnx` artifacts instead and this PR should be closed. On the evidence in #1883 that is not the intent. Auto-merge armed, waiting on required checks. No `--admin`, no bypass. Refs #1832, #1723, #1883. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#1893) ## Summary While running a read-only tiny-MoE E2E baseline (3 fixtures) through the unified workflow/runtime architecture from #1723, `cargo build --release -p onnx-genai-bench --features native-cuda --bin profile_native` **did not compile at all** on `origin/main`. `onnx-genai-bench` has zero CI coverage (absent from `.github/workflows/ci.yml`), so this — plus two separate, independent, pre-existing bugs — went undetected. This PR is the tooling-only repair; the E2E baseline results themselves are reported separately (see below). **Scope: `crates/onnx-genai-bench` only. No `onnx-genai-engine` (production) changes.** ## What broke, and why 1. **`Engine::from_dir_with_config` doesn't exist post-#1723** → `Engine::from_dir`. 2. **`Pipeline::clear_prefix_cache()` doesn't exist post-#1723** (no public replacement was ever wired into this bench crate) → replaced both call sites with the public, per-request `GenerateOptions::cold_start`, the documented #1723 successor. Scoped to one request instead of mutating shared engine state — strictly better. 3. **`generate_with_callback(pipeline_request(..)?, ..)` type mismatch** (needs a `GenerateRequest`, got a `PipelineGenerateRequest`) at two call sites → `generate_with_pipeline_callbacks`. One of the two was hidden by a rustc diagnostic-suppression quirk on the first error and only surfaced after fixing it. 4. **`NativeDecodeSession::generate` and family became `pub(crate)`** → added a local greedy-only `native_session_generate()` reimplementation (with matching `loop.next_logits`/`loop.sampling` `prof_span!` instrumentation so the `--synthetic` profiling regression test keeps passing) and a manual `token_logprob()` log-softmax/top-k helper for `--dump-logprobs`. 5. **Separate, pre-existing bug (unrelated to #1723):** `NativeDecodeSession::from_session(..)` requires shape/dtype-unambiguous ports. The bench crate's synthetic decoder's 3 `Int64 [-1,-1]` inputs (`input_ids`/`attention_mask`/`position_ids`) are structurally indistinguishable, so this deterministically failed ("N ports match") — by design, confirmed via an existing engine-crate test. Fixed with the documented `from_session_with_io(..)` escape hatch and a new `synthetic_decoder_io()` `DecoderAbi` (adapted from `fused_batch_prefill.rs`'s `synthetic_io()` helper — which is itself independently bit-rotted, see "Found but not fixed" below). 6. **Separate packaging bug:** `onnx-genai-metadata` (needed for `DecoderAbi`) was `[dev-dependencies]`-only, so no regular binary target could ever call that public API surface. Promoted to `[dependencies]`. 7. Two pre-existing clippy lints inside `profile_native.rs` (`manual_contains`, `needless_borrows_for_generic_args`) and one on `tests/profile_native.rs`'s `assert_floor` helper (`too_many_arguments`), fixed to keep a scoped clippy gate green. ## Independent review An independent code-review pass on the initial repair found two substantive issues, both fixed in this PR: - **`token_logprob()` could panic** on an overflowed (`+inf`) logit: `max` became infinite, `v - max` on that entry produced `NaN`, which propagated through every ranked log-probability, and the old `partial_cmp(..).expect("logits must not be NaN")` sort comparator panicked on the resulting NaNs. Now mirrors `decode_loop::logprob_for_token`'s NaN/Inf-safe approach (filters non-finite logits, uses `total_cmp`). Added regression tests. - **`native_session_generate`'s bail condition used `!chain.is_empty()`**, an incidental proxy. Replaced with `!chain.preserves_argmax()` — the actual eligibility test the engine's own device-greedy fast path uses — and updated the `--repetition-penalty` doc comment, which had drifted out of sync (it read as though non-default values worked in raw/default mode, when in fact — like `--top-p`/`--top-k`/`--min-p` — they require `--steady`/`--pipeline`). ## Found but not fixed (out of scope, reported for follow-up) - `crates/onnx-genai-bench/tests/fused_batch_prefill.rs` references a nonexistent `kv_update` field on `DecoderAbi` — independently broken, unrelated to #1723 or this repair. - `bench_generic.rs`/`compare.rs` carry their own pre-existing clippy failures under `--all-targets` (const-assertion lint, collapsible-if) — left alone; this PR's clippy gate is deliberately scoped to `--bin profile_native --test profile_native`, not `--all-targets`. - Two of the three tiny-MoE fixtures (`tiny-deepseek-v4-qmoe`, `tiny-glm52-full-attention`) currently fail to load at all (`Engine::from_dir`) because their committed `inference_metadata.yaml` declares 10 `policies/*.onnx` workflow-component artifacts that were never committed to the fixture directories. Confirmed via the *unmodified* authoritative engine-crate regression tests (`deepseek_v4_tiny_qmoe_e2e.rs`, `glm_tiny_full_attention_e2e.rs`) — this is a pre-existing fixture-generation gap, not caused by this PR or #1723. `generate.py --mobius-root <path>` (a local Mobius checkout exists) can presumably regenerate the missing artifacts; tracked as a follow-up, not attempted here (out of scope: this PR does not touch fixtures). ## Verification - `cargo build --release -p onnx-genai-bench --features native-cuda --bin profile_native` — clean, no warnings. - `cargo test --release -p onnx-genai-bench --features native-cuda --bin profile_native --test profile_native` — 8 unit + 4/5 non-ignored integration tests pass (1 correctly `#[ignore]`d, needs real multi-node hardware). - `cargo fmt -p onnx-genai-bench -- --check` — clean. - `cargo clippy --release -p onnx-genai-bench --features native-cuda --bin profile_native --test profile_native -- -D warnings` — clean. - Re-verified end-to-end correctness on an idle A100 (GPU 1) after all fixes: `profile_native --model tests/fixtures/tiny-glm52-qmoe-indexshare --ep cuda --steady --tokens 12 ...` still produces `generated_token_ids: [33, 221, 186, 142, 162, 143, 148, 18, 18, 18, 18, 18]`, byte-identical to the locked `ANCHOR_IDS` regression constant in `crates/onnx-genai-engine/tests/glm_tiny_qmoe_native_cuda_e2e.rs`. Closes: tooling blocker for issue #82's E2E baseline effort (see separate report for the actual 3-fixture baseline measurements). Working as Sebastian (Performance Engineer). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ntly (#1899) ## Why `profile_native` is the tool every decode benchmark under `docs/benchmarks/` is produced with — and **nothing in CI compiles it**. - The Fast job's `cargo build --locked -p …` list cannot reach it: it is a **bin** that needs `--features bench-native`. - `benchmark.yml` runs only `--bench no_model` and `--bench model`, neither of which builds it. - A root `cargo build` does not enable the feature either. So #1723 broke it across five API changes and nobody noticed until I tried to use the tool by hand. That was #1878 (fixed by #1893) — but the *reason it landed* is this gap, and without closing it the next engine API change breaks it again just as quietly. ## What One step in the Fast (Linux) job that **builds, not runs**, the bin. Cheap: it compiles crates the job already builds, plus the bin itself. ```yaml - name: Build the native profiling bin run: >- cargo build --locked -p onnx-genai-bench --features bench-native --bin profile_native ``` ## Verification - The command itself is verified green on current `main` (`1be9f2cc2`): `Finished dev profile in 19m 08s` — that is how I confirmed #1893 actually fixed #1878 before closing it. - `ci.yml` parses as valid YAML and the step lands in the `fast-linux` job (checked by loading the workflow and inspecting the parsed step list, not by eyeballing the indentation). ## Note on cost This adds a compile of `onnx-genai-bench` with `bench-native` to the Fast job. That feature pulls in the native decode stack, which the job already builds for other crates, so the marginal cost is the bench crate itself rather than a new dependency tree. If it turns out to lengthen the job more than is wanted, the honest alternative is a separate scheduled job — but a silent breakage that reaches a benchmarking tool is worse than a slower Fast job, and this is the cheapest form of the check. Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
…scheduler (#1900) ## Root cause #1723 ("one runtime, one interpreter, one drive") changed `generate_with_callbacks`'s top-level dispatch from checking the `decode_backend` field to checking `holds_decode_core()` (`native_session.is_some() || session.is_some()`). Pre-#1723, **every** `Native`-backend request — regardless of whether a decode session already existed — went through the scheduler-admitting cold-start path (`generate_native_cold_with_callback`), which checks the KV byte budget before touching any backend. Post-#1723, a `Native`-backend `Engine` with no session now silently falls into the unguarded interpreted path (`generate_interpreted` → `generate_with_pipeline_callbacks` → `run_declared_generation`), skipping scheduler admission entirely. This is a **real, reachable** state: `Engine::from_dir` → `Self::decode_core_covers(&workflow)` false → `from_interpreted_dir` → `Engine::from_workflow` is a legitimate, doc-commented construction path ("a package whose components the interpreter invokes"), and `WorkflowRuntime::decode_backend()` can independently be `Native` for such a package. Today no shipping package pairs this with a byte-budget-configured scheduler (`Engine::from_workflow` always uses `Scheduler::new(SchedulerConfig::default())`), so the practical blast radius is currently zero — but the gap is real and was caught by a pre-existing test, `native_generate_rejects_over_kv_byte_budget_before_backend_run`, which was failing on main with the wrong (deep, non-admission) error message instead of a clean scheduler rejection. ## Fix - Added `admit_interpreted_generate_request` (`engine/runtime.rs`), which computes the prompt token count via a new `interpreted_prompt_token_count` (handling `TokenIds` by length, `TokenRows` by `rows.len() * max(row length)` — since equal-length rows bind into one batched `[rows, columns]` tensor and run as a single graph step — and `Text` via the package's own tokenizer) and calls the existing `admit_generate_request_with_scheduler`. - Called this from `generate_with_pipeline_callbacks`'s (`engine/workflow_api.rs`) own no-decode-core, prompt-only branch — the single place every such request converges, whether it arrives cold (`generate_interpreted`) or through a continuing workflow session (`generate_in_workflow_session`, which threads the caller's own session id through so admission accumulates the same way the decode-core path's session continuation already does). - Tensor-bound (multimodal) requests are deliberately left unchanged, matching pre-existing, unrelated precedent (CLI `interactive.rs`/`transcribe.rs`, server `driver.rs`) that skips admission regardless of decode-core status. ## Before / after evidence Before the fix, `native_generate_rejects_over_kv_byte_budget_before_backend_run` failed with: ``` workflow component 'decoder' input 'past.key' references unavailable value 'decoder_kv.0000' ``` (a deep node-execution failure, not a scheduler rejection). After the fix it passes cleanly with `"scheduler admission failed: KV byte budget..."`. A new test, `native_generate_in_workflow_session_rejects_over_kv_byte_budget`, exercises the session-continuation path directly (opens a real session via `create_session()`, drives `generate_in_workflow_session()`). I confirmed by temporarily reverting `workflow_api.rs` alone that this test fails against the pre-fix code with the same deep `references unavailable value` error, and passes after the fix — closing a gap an independent review pass on an earlier version of this change flagged (that version only guarded the cold-call path, missing session continuation, and undercounted `TokenRows`). ## Review Independently reviewed twice: once on an initial version (which lived inline in `generate_interpreted` and only covered the cold-call path — REQUEST CHANGES, two findings), and once on this final restructured version, which addresses both findings by moving the admission call site to the shared no-decode-core branch of `generate_with_pipeline_callbacks` and fixing the `TokenRows` formula. Second pass: **APPROVE**, after tracing control flow, confirming `scheduler.complete()` runs on every exit path, confirming no other caller bypasses the new admission branch, and independently reproducing the before/after test evidence. ## Tests (all local; no CI wait per current session directive) - `cargo test -p onnx-genai-engine --features native-backend --lib`: 618 passed, 0 failed, 1 ignored - `cargo test -p onnx-genai-engine --features native-cuda --lib`: 648 passed, 0 failed, 4 ignored - `authored_body_selects_executor` (4), `one_runtime_e2e` (6), `native_workflow_parity` (13), `native_workflow_smoke` (2), `canonical_execution_parity` (7): 32/32 passed - `cargo fmt --check -p onnx-genai-engine`: clean - `cargo clippy -p onnx-genai-engine --features native-cuda --all-targets -- -D warnings`: clean - GPU e2e on idle A100 (`CUDA_VISIBLE_DEVICES=1`, `--test-threads=1`): - `deepseek_v4_tiny_qmoe_e2e`: 4/4, `captures=1 replays=10 fallbacks=0` - `glm_tiny_qmoe_native_cuda_e2e`: 4/4 - `glm_tiny_full_attention_e2e`: 4/4 - `deepseek_v2_tiny_qmoe_native_e2e`: 2/2 ## Scope No workflow-unification refactors, no CUDA residency/paging/scheduler changes, no DeepSeek/GLM export/model-name branching. Scoped strictly to closing the #1723-introduced admission gap in the shared interpreted-path runtime code. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…l-closed same-device checks Fix-forward revision of rejected PR #1854 (author Deckard locked out of this artifact). Addresses all 9 Cycle-17 blocking findings from Roy's review: 1. Range-level (not value-level) rollback tracking: if range N within one ValueId commits and a LATER range in that same value hits Fatal, the earlier committed range is now correctly included in the rollback. 2. BoundaryApplicationOutcome.fatal_progress surfaces committed_count and poisoned_range straight from TransitionOutcome::Fatal (no more '..' discard). 3. Rollback-of-rollback failures are now explicit: a new RollbackFailure struct (value, range, detail, committed_count, poisoned_range, quarantined) replaces the previous silent all_ok=false. 4. New deterministic PLAN-ENTRY fault injection: a test-only apply_residency_plan_at_boundary_with_phase8_faults entry point accepts a per-ValueId DriverFaultPlan map, enabling 4 new GPU tests that exercise partial-commit-then-fatal, rollback-of-rollback, same-device rejection, and expert-group atomicity end to end (not just at the primitive level). 5. Same-device fail-closed: check_same_device validates allocator.device_key() and both pools' device_ordinal against the requested device_ordinal before any mutation. 6. Corrected the module doc's false 'one production consumer called at model load' claim; #1723 is now honestly described as technical current-main context only, with zero production callers today. 7. Cross-ValueId atomic expert-group tiering: new graph-derived ExpertWeightGroup type + expert_weight_groups(graph) (onnx-runtime-ep-api), purely structural from QMoe/BlockQuantizedMoe node inputs (no tensor-name heuristics). apply_residency_plan_at_boundary now takes an expert_groups parameter and enforces all-or-none group fallback via a validation-only precheck phase that completes for an entire group before any member is mutated. 8. Renamed misleading telemetry: device_bytes_before/after -> single truthful device_bytes_released; COARSE_RESIDENCY_PROFILE_ENV -> COARSE_RESIDENCY_ENABLE_ENV (deprecated alias kept for compatibility). Also fixes a pre-existing (unrelated) Cargo feature-forwarding bug: onnx-runtime-ep-cuda's gpu-tests feature did not forward to onnx-runtime-cuda-memory/gpu-tests, hiding CudaVmmAllocator test-only items needed to compile these tests. Tests: 7/7 coarse_residency_plan_gpu (3 pre-existing + 4 new) pass on A100 (CUDA_VISIBLE_DEVICES=4), 4/4 coarse_residency_plan (non-GPU) pass, 3/3 + 91/91 unit tests (onnx-runtime-ep-cuda coarse_residency + onnx-runtime-ep-api) pass, 530/530 onnx-runtime-ep-cuda lib tests pass (no regressions), and all 29 Slice2-4 GPU regressions (composable_vmm_production_gpu, content_preserving_transition_gpu, expert_bank_remap_cost_gpu) pass unchanged. fmt/clippy clean on all touched crates. Independent code review (not Deckard) found no blocking issues. PR remains in draft; Roy's explicit re-review is still required before merge. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…1892) ## Turn 3 was turn 3's prompt `justinchuby/qwen3-0.6b-onnx-genai@a89c02d3` declares all 65 of its state cells `scope: invocation`, and its 56 cache cells `release_boundary: invocation`. The interpreter honours that. Measured on the published package, before this branch: | | tokens | |---|---| | turn 3 of a session | `[576, 3491, 374, 311, 1477, 279, 1372, 315]` | | the same prompt sent cold, no session | `[576, 3491, 374, 311, 1477, 279, 1372, 315]` | | one request carrying the whole conversation | `[358, 1079, 264, 5458, 315, 279, 12103, 315]` | Turn 3 was byte-identical to the cold request. This is the regression #1723 called out and named two possible fixes for. The other one — *prefer the decode core when a package's workflow is decoder-shaped despite shipping several components* — is the runtime supplying a step the package declares a graph for, which is the thing #1723 removed. This is the first one: **the package declares its conversation, and the runtime carries it.** ## Root cause: the document, not the interpreter `scope: session` says *keep this*. It never said **how the next invocation reaches what was kept**, and there are only two answers: * **the graph reads it** — a loop carries the cell, or a step consumes the value its `initializer` names. `moshiko-full-duplex` does this with 72 cells, and the runtime already carried them correctly. * **the request binding rejoins it** — and nothing could say so. A decoder like this one has no reader of the second kind and cannot grow one. Its `decoder_state_initializer` takes `(prompt_tokens, prompt_lengths, max_iterations)` and emits empty `past_key_values.*`; its setup then prefills the turn's prompt from that empty cache. There is no port for a prior session length, so a runtime handing it turn 1's cache would produce a mask and a cache that disagree. Declaring `scope: session` on the existing cells would not have fixed multi-turn — it would have produced *garbage*, which is worse. That is why this is not a one-line metadata patch. ## The declaration ```yaml conversation: contract: {dtype: int64, rank: 2, shape: [batch_size, conversation_length], batch_layout: {kind: request_aligned, axis: 0}} class: semantic scope: session initializer: request.input_ids recurrence: {kind: bounded, axis: 1, max: package.max_context} management: runtime release_boundary: session session: policy: exclusive continuation: kind: prompt_prefix prompt_input: request.input_ids # must carry role prompt_tokens tokens_output: tokens # must carry role tokens ``` A turn's prompt is the cell's value followed by the caller's tokens; the cell then absorbs both that prompt and what the turn published. It is bound **before the request is bound**, so every input derived from the prompt — the token tensor, `prompt_lengths`, the mask the initializer builds — is derived from the whole conversation, exactly as for a caller who sent it in one request. Rewriting the bound tensors afterwards would leave the derived ones describing only the turn. Nothing model-specific, no `model.io`, no branch on package kind: the only questions asked are *does this workflow declare a continuation* and *does this request carry a session*. ## Three layers, each changed where it was wrong **Metadata** — `SessionLeaseContract.continuation`, and validation that fails the document closed rather than the third turn. A continuation must be session-scoped, `class: semantic`, `management: runtime`, `release_boundary: session`, growing, and not also loop-carried (two answers about one value); it must name a declared `prompt_tokens` input and a declared `tokens` output whose contracts match; a workflow declares at most one, because a package has one conversation. Separately, **a session-scoped cell binding a `service_group` must resolve to a declared group that aliases it** — the "claims continuity with no valid state groups/aliases" case. **Runtime** — the conversation is bound before the pass and published after it, read from whichever way the pass published its tokens (`tokens`, `tokens.row.N`, or streamed `tokens.<n>` events; the real package publishes `tokens.row.0`, which is the bug the first draft of this had). Sessions are keyed and isolated by id. `reset_session` now works for an interpreted package and releases the lease; `close_session` already forgot it. Seeding no longer **errors** when a request carries no session — before this, a package could not declare session state *and* answer `Engine::generate`, which is a conversation's own first turn. `session_token_count` reports the conversation rather than generated tokens alone, and `Engine::session_conversation` reports what it holds. **Fail closed** — `create_session` refuses a package that publishes a token stream and declares no session state the interpreter can carry, naming what the package must declare. Against the old published metadata: ``` this package publishes a token stream but declares no session-scoped workflow state, so a session could not continue a conversation: every turn would restart from its own prompt. Declare the conversation in `pipeline.workflow.state` with `scope: session` — either as a cell the loop carries, or with a `session.continuation` naming the prompt input it rejoins — and re-validate the package. ``` A package that publishes no token stream — speech, diffusion, video, codec — has no conversation to lose and keeps its session handle, which is what `sessions_open_and_close_for_an_interpreted_package` already pins. ## The audit, and why the producer is not the site Every `inference_metadata.yaml` the migration published, at its pinned revision: | package | onnx components | state groups | `scope: session` cells | |---|---|---|---| | `qwen3-0.6b-onnx-genai` | 11 | `decoder_cache` | **0** | | `qwen2.5-0.5b-instruct-onnx-genai` | 11 | `decoder_cache` | **0** | | `qwen2.5-14b-instruct-int4-zp-onnx` | 11 | `decoder_cache` | **0** | | `deepseek-r1-distill-qwen-1.5b-onnx-genai` | 11 | `decoder_cache` | **0** | | `…-mistral-7b-v0-1-sliding-window` | 11 | `decoder_cache` | **0** | | `…-qwen2-5-0-5b-{portable-f32,cuda-gqa-f16,static-cache-f32}` | 11 | `decoder_cache` | **0** | | `…-qwen2-5-1-5b-lora-selection` | 11 | `decoder_cache` | **0** | | `…-qwen3-5-0-8b-hybrid-vlm-f32` | 14 | `decoder_cache` | **0** | | `Muse-Glimmer-30B-ONNX-INT4-CUDA` | 14 | full+sliding | **0** | | `moshiko-full-duplex-onnx-catalogue` | 23 | `temporal_cache` | **72** | Systematic — and **not from this repo's producer**. `decoder_workflow`, which backs both `migrate_model_io` and the `genai_config.json` importer, emits the `binding` token policy of rule 4, so its packages load onto the fused decode core whose paged KV sequence *is* the conversation; they continue one without declaring one, and `multi_session.rs` / `one_runtime_e2e.rs` still pass unchanged. The affected packages ship ONNX policy graphs instead and have no such core. So the fix lands in the schema, the runtime, and the package metadata — and in `tests/fixtures/onnx_genai_workflows/decoder`, the in-tree copy of exactly the affected shape, which now declares its conversation (+26 lines, no other edit). ## Published package `justinchuby/qwen3-0.6b-onnx-genai` * old `a89c02d343fc2b3c49d61b33d61f686698be1e4d` * new [`24da071476fdd1abaf7bb514424af8bc6604827f`](https://huggingface.co/justinchuby/qwen3-0.6b-onnx-genai/commit/24da071476fdd1abaf7bb514424af8bc6604827f) A pure addition: one state cell and `session_state_lease` in the manifest. Diffed section by section against the old revision — the graphs, ports, inputs, outputs, state groups, serving contract, loop and **every existing state cell** are unchanged. `validate_metadata` passes on the full package, and the revision was re-downloaded from the hub and executed before this PR was opened. No package id appears anywhere in the code. ## Evidence Against the **published new revision**, downloaded fresh (`qwen3_0_6b_multi_turn_session.rs`, gated on `ONNX_GENAI_QWEN3_WORKFLOW_DIR`): ``` turn 3 [1084, 594, 264, 1602, 2989, 949, 315, 847] single shot [1084, 594, 264, 1602, 2989, 949, 315, 847] ← equal cold [576, 3491, 374, 311, 1477, 279, 1372, 315] ← what turn 3 used to be ``` `session_token_count` 16 → 42 across three turns of 8 tokens (34 conversation tokens plus turn 3's 8). Independent session's first turn equals the cold request; `reset_session` returns the count to 0 and the next turn equals the cold request again. Hermetic, in CI: * `workflow_session_continuation.rs` (6 cases, Engine): turn 3 against the single-shot conversation; the conversation read back cell-by-cell after each turn — this fixture's synthetic weights decode a constant, so its evidence is *what the decoder was asked to decode*, not what came out; isolation; reset/close; stateless unchanged; and a package with the conversation stripped out refusing a session. * `chat_completions_continue_a_conversation_across_requests` and `chat_completions_without_a_session_stay_stateless` (server): three `POST /v1/chat/completions` sharing an `X-Session-Id`, with `session_token_count` strictly increasing and turn 3's context at least three times turn 1's; a different id starting its own conversation; `DELETE /v1/sessions/{id}` releasing it and a second delete returning 404. * `session_continuity.rs` (10 cases, metadata): every fail-closed rule above. ## Validation `cargo fmt --all --check`; `clippy -D warnings --all-targets` on metadata, engine, server, CLI and C API. **onnx-genai-metadata** 23/23 suites, **onnx-genai-engine** 478 lib + 75 integration binaries, **onnx-genai-server** 243 lib + 40 http, CLI / C API / genai-config green. The JSON schema is regenerated and `schema_sync` passes. --------- Signed-off-by: justinchuby <justinchu@microsoft.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
One runtime, one interpreter, one drive
Every package — a bare decoder, a composite pipeline, a speculative block —
serializes
pipeline.workflow, and every generated token now comes out ofrun_workflow_nodewalking that document. The iteration bound, the livenesspredicate, the carried state, the emit that publishes the token stream and the
stop are read from the package rather than supplied by a Rust loop written
beside it.
What varies between packages is which executor implements a declared node, and
that is answered by the contract the node names. A component declaring
onnx-genai.autoregressive-decodeis run by the fused decode session — paged KV,the device sampling fast paths, the captured CUDA graph — when this process
holds one, and invoked from its own artifact when it does not.
holds_decode_core()asks whether that session exists; nothing asks what kindof package was loaded.
Measured, not asserted
authored_body_selects_executor.rsdrives three packages declaring threedifferent bodies through one entry point and reads a per-contract execution
counter recorded inside the interpreter's node dispatch:
tiny-llm(runtime decode + policy)decoderfixture (in-graph sampler)model, plus a fused policy islandspeculativefixtureThe zero matters more than the counts: a package that declares a graph for every
step must never have the runtime supply one, or the declared document is
decorative.
Deleted
pipeline/canonical_decode.rs(771 lines),Engine::lowered_workflow,canonical_workflow()and its nine guard call sites,WorkflowShape/declared_workflow_shape,is_workflow(),workflow_shape(),WorkflowShapeReport,from_pipeline_dir*,generate_pipeline_with_callback(s),workflow_generate,forget_canonical_workflow_for_test, the server'sstart_pipeline/run_pipeline_driver/run_pipeline_generation/GeneratePipeline, and theCLI's
Backend::{Text,Pipeline}enum.The shape-dispatch gate's allowance went from 23 files to zero (the gate
itself aside). Those symbols do not exist anywhere in the tree.
Three things the convergence surfaced
AUTOREGRESSIVE_DECODE_CONTRACT, but the fourteen converted fixtures predatedit — nothing in the tree declared the step it was executing. Re-emitted through
a new
migrate_model_io --reemit, which reads the ABI back out of the workflowthe package already declares.
sequencecell's recurrence was wrong. It saidinvariantwhilecarrying the prompt on the first iteration and one token thereafter. Nothing
could honour that, and nothing had to, because no interpreter was walking the
loop. It now declares
bounded.max_iterations_onlyoptimization is sound for an in-graph EOS predicate acaller disabled; a predicate a runtime executor writes carries the context
bound and stop strings too. Without this the context-limit stop was ignored and
generation ran into a model-window overflow.
Real finish reasons, and a load-bearing emit
The composite drive no longer fabricates
MaxTokens,prefix_cache_hit_len: 0and
logprobs: None. It reports the core's stop when a core reached one, theinterpreter's own record of whether the predicate ended the loop otherwise, and
the prefix/logprob facts of whoever scored the tokens —
Nonefor a packagewhose sampler is an ONNX component, which is a fact about the package rather
than a missing feature. The emitted stream and the committed one are compared,
and a mismatch refuses.
Sessions
A session is the conversation a client is having; where the runtime keeps it is
not something a client can act on.
/v1/completionsand/v1/sessionsno longerreject workflow packages, and
create_session/session_token_count/close_sessionmean the same thing for every package.Real-model evidence
.goldens/REAL_MODEL_EVIDENCE.md— one probe, two builds (cleanorigin/main@cb81745b0worktree versus this branch), same machine, pinnedpackage revisions.
qwen3-0.6b-onnx-genai@a89c02d3decodes byte-identically on both, throughcompletely different execution paths:
mainran it on the decode core from itsgenai_config.json, head runs the eleven-component workflow it declares.gemma4-e2b@b1189a9bandgemma4-e2b-speculative@3e5d3b6acannot run onmainat all (one fails before generating, one fails to load). Both load onhead.
differs (head uses the sampler the package declares), and multi-turn
continuation is lost for
qwen3_0_6bbecause its declared workflow has nosession-scoped state. That is called out as a regression with both fixes named.
cannot seed, so composite speculative widths 1/2/4 against those packages
have no measurement. The hermetic
gemma4_chainedCUDA gate is not substitutedfor it.
Two defects the pinned packages found are fixed here:
generate_in_sessionfailed on a package with no decode core because only one wrapper had been routed
to the interpreted session path, and a text prompt against a package declaring a
prompt_tokensinput was refused even when the package shippedtokenizer.json.CUDA residency — partial, and measured
WorkflowRuntimecounts every device→host materialization, and the H200 chainedparity gate reads it across a propose/verify/reject/rollback decode.
Landed: borrowed proposer bindings are narrowed in place (a pointer view via
the new
Value::alias_with_offset) instead of being copied down in full andnarrowed after; the per-draft-token argmax reads four bytes off the device
instead of a ≈300 KiB logits row; the embedding table is read once rather than
per proposal.
tensor_ops_foris fail-closed — a device value with no deviceimplementation is an error naming both remedies, never a quiet copy.
Not landed, and the assertion says so: the count is a ratchet at 16, not a
zero. The folded carry seed and the folded carry itself still cross the bus,
because the fused proposer input is assembled on the host. Removing them needs a
device-side scatter this seam does not have, and faking it with a staging copy
would satisfy the counter and change nothing.
Validation
465 engine lib tests, every engine integration test, 241 server tests, metadata /
CLI / C API suites,
cargo fmt,clippy -D warningson every crate touched, andall 12
native_workflow_paritycases on an H200 including both chainedspeculative gates.
origin/mainwas merged (never rebased); the zero-allowlistIR11/opset24 fixture guard passes.
Known pre-existing failure, verified on
origin/main@cb81745b0in a cleanworktree and unrelated to this branch:
onnx-genai-ort::loader::tests::selected_non_dense_candidate_fails_explicitly.