Repository navigation
feat(server): refuse overlapping turns on one session at the routing layer - #2056
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2056 +/- ##
==========================================
+ Coverage 80.32% 80.57% +0.24%
==========================================
Files 426 428 +2
Lines 204769 209700 +4931
Branches 204769 209700 +4931
==========================================
+ Hits 164484 168960 +4476
- Misses 34664 35014 +350
- Partials 5621 5726 +105
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
…strict Review of #2056 found the lease was correct for one model and wrong for two. A `SessionPlacement` is a worker id plus an engine session id, and both are per-engine: every engine numbers its own sessions from its own counter and every worker pool starts at worker 0. Two loaded models therefore name two unrelated conversations with the identical placement. The lease map lived on `EngineDriver`, so it also answered "is this session busy?" only for whichever engine the caller happened to be holding — while the `SessionRegistry` that has to ask that question spans every loaded model. Key the lease by `ModelSessionPlacement { model: ModelKey, placement }` and give the one `SessionLeases` map to the `SessionRegistry`, which owns the bindings the lease is about. An engine learns its own `ModelKey` once, in `ModelHandle::new`, so a handle's id and its driver's are the same string by construction; a close whose lease names another model is refused rather than performed. `SessionEntry` stores the model-qualified binding, so eviction and close land on the engine that opened the session, and a session id presented on a different model is refused with the same typed 409 instead of generating into a stranger's conversation. Three further invariants the lease is only correct under: `max_sessions` is now strict. When every binding is mid-turn there is no evictable victim, and the new session is refused with a typed `AtCapacity` mapped to the existing 429 `resource_limit_error` rather than admitted over the bound. Admitting it made the limit advisory *permanently*: nothing walks the registry back down, because the next insert evicts one and adds one, so a server sized for n conversations could be pushed to n+k and stay there. The refusal is transient and clears when any turn in flight ends. Close is one decision, not three. `DELETE` no longer reads a binding, takes its lease, and then removes it — those are three decisions about a binding that can change between them, and the middle one is where a rebind slips in and the close destroys a conversation it never leased. `SessionRegistry::take_for_close` holds the registry lock across the find, the acquire and the remove and returns the guard naming the owner, so what is leased, what is unbound and what is closed are the same binding on the same engine. It also removes the `registry.resolve("")` default-model close. LRU eviction obeys the same rule. Insert is exact. An id that is already bound is refused with `AlreadyBound` rather than silently rebound, so the active-session gauge moves by exactly one per insert and one per close, and the registry's own `Drop` returns it to baseline. Tests load two models whose first sessions have provably identical placements — the fixture asserts the collision rather than assuming it — and pin that a busy session on one model cannot be evicted or closed by the other, that a `DELETE` of a non-default model's session closes it on that model's engine, that an insert at full capacity with every conversation busy is refused rather than overshot, and that the next insert after a release evicts rather than grows. Two thread-and-barrier regressions cover the close race: racing deletes unbind exactly one binding once, and a delete racing a rebind never orphans a conversation. Pre-enqueue acquisition and the typed 409 are unchanged. This is still routing exclusion only — no turn runs in parallel with another. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
438f821 to
dc366c9
Compare
Review fixes — all five blocking items addressedRebased onto latest 1. Globally unique lease key, one shared mapThe review is right that the collision is real, and there's now a test that asserts it rather than assuming it: two loaded models each number sessions from their own counter and each
2.
|
| Command | Result |
|---|---|
cargo test -p onnx-genai-server --lib |
302 passed, 0 failed, 2 ignored |
cargo test -p onnx-genai-server --test http |
40 passed, 0 failed, 1 ignored |
| full server suite ×4 | 3/4 fully green; 1 environmental latency flake (above) |
cargo test -p onnx-genai-engine --lib -- session |
21 passed |
cargo test -p onnx-genai-engine --test multi_session |
2 passed |
cargo test -p onnx-genai-engine --test onnx_genai_workflow_conformance |
15 passed |
cargo fmt --all -- --check |
clean |
cargo clippy -p onnx-genai-server --all-targets -- -D warnings |
clean |
Pre-enqueue acquisition and the typed 409 are unchanged. Still routing exclusion/refusal only — no turn executes in parallel with another, W is still 1, no intra-worker multiplexing, no global engine mutex, no backend Send/Sync changes.
SESSION_CONCURRENCY.md §4.2, §5, §5.1, §12.1 and §13 updated for what actually landed — in particular §5's paragraph documenting the old "run one over the bound" LRU behaviour is replaced with the fail-closed refusal.
Not merging.
…strict Review of #2056 found the lease was correct for one model and wrong for two. A `SessionPlacement` is a worker id plus an engine session id, and both are per-engine: every engine numbers its own sessions from its own counter and every worker pool starts at worker 0. Two loaded models therefore name two unrelated conversations with the identical placement. The lease map lived on `EngineDriver`, so it also answered "is this session busy?" only for whichever engine the caller happened to be holding — while the `SessionRegistry` that has to ask that question spans every loaded model. Key the lease by `ModelSessionPlacement { model: ModelKey, placement }` and give the one `SessionLeases` map to the `SessionRegistry`, which owns the bindings the lease is about. An engine learns its own `ModelKey` once, in `ModelHandle::new`, so a handle's id and its driver's are the same string by construction; a close whose lease names another model is refused rather than performed. `SessionEntry` stores the model-qualified binding, so eviction and close land on the engine that opened the session, and a session id presented on a different model is refused with the same typed 409 instead of generating into a stranger's conversation. Three further invariants the lease is only correct under: `max_sessions` is now strict. When every binding is mid-turn there is no evictable victim, and the new session is refused with a typed `AtCapacity` mapped to the existing 429 `resource_limit_error` rather than admitted over the bound. Admitting it made the limit advisory *permanently*: nothing walks the registry back down, because the next insert evicts one and adds one, so a server sized for n conversations could be pushed to n+k and stay there. The refusal is transient and clears when any turn in flight ends. Close is one decision, not three. `DELETE` no longer reads a binding, takes its lease, and then removes it — those are three decisions about a binding that can change between them, and the middle one is where a rebind slips in and the close destroys a conversation it never leased. `SessionRegistry::take_for_close` holds the registry lock across the find, the acquire and the remove and returns the guard naming the owner, so what is leased, what is unbound and what is closed are the same binding on the same engine. It also removes the `registry.resolve("")` default-model close. LRU eviction obeys the same rule. Insert is exact. An id that is already bound is refused with `AlreadyBound` rather than silently rebound, so the active-session gauge moves by exactly one per insert and one per close, and the registry's own `Drop` returns it to baseline. Tests load two models whose first sessions have provably identical placements — the fixture asserts the collision rather than assuming it — and pin that a busy session on one model cannot be evicted or closed by the other, that a `DELETE` of a non-default model's session closes it on that model's engine, that an insert at full capacity with every conversation busy is refused rather than overshot, and that the next insert after a release evicts rather than grows. Two thread-and-barrier regressions cover the close race: racing deletes unbind exactly one binding once, and a delete racing a rebind never orphans a conversation. Pre-enqueue acquisition and the typed 409 are unchanged. This is still routing exclusion only — no turn runs in parallel with another. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
dc366c9 to
b0c71cc
Compare
Rebased onto latest
|
| Base before | eccc824f6 |
| Base now | 089309588 |
| Commits | 60f2a9764 (lease) → b0c71cc95 (review fixes) — linear, 2 commits, no merges |
Interdiff: empty
The replay was clean — no conflicts, and nothing needed semantic resolution:
$ diff <(git diff eccc824f6..dc366c996) <(git diff 089309588..b0c71cc95) # modulo blob hashes
PATCH CONTENT IDENTICAL
Both patches are 4135 lines. The branch's own diff is byte-for-byte what it was before the rebase, so there is no interdiff to review and no re-review needed on the code.
That was checked rather than assumed, because upstream did touch this crate. Its server-crate changes (image_generation.rs, multimodal.rs, routes/images.rs) are image-generation dtype/shape metadata plumbing following the onnx-genai-metadata schema work — routes/images.rs contains zero occurrences of session, and none of the three files touch the registry, the driver's command surface, or any lease type. The struct changes in that series (VisionInputSpec, ImageInputBinding, the metadata schema/version types) are upstream of the routing layer this PR changes and share no field with it. All four completion handlers still bracket their turn with lease_bound_session → open_session_after_admission, and SessionRegistry still owns the one SessionLeases map.
Post-rebase verification
| Command | Result |
|---|---|
cargo test -p onnx-genai-server --lib |
302 passed, 0 failed, 2 ignored |
cargo test -p onnx-genai-server --test http |
40 passed, 0 failed, 1 ignored |
cargo test -p onnx-genai-engine --lib -- session |
21 passed |
cargo test -p onnx-genai-engine --test multi_session |
2 passed |
cargo test -p onnx-genai-engine --test onnx_genai_workflow_conformance |
15 passed |
cargo fmt --all -- --check |
clean |
cargo clippy -p onnx-genai-server --all-targets -- -D warnings |
clean |
cargo clippy -p onnx-genai-engine --lib -- -D warnings |
clean |
Fully green on this pass, including fim_stream_returns_headers_before_generation_finishes (the wall-clock latency test that flaked once under load in the previous round).
Behaviour is unchanged by the rebase: still routing exclusion/refusal only, W = 1, pre-enqueue acquisition, typed 409.
Not merging.
…layer
Implements Phase 2 of docs/architecture/SESSION_CONCURRENCY.md: a
routing-layer exclusive turn lease, keyed by the typed `SessionPlacement`
and acquired before a turn becomes work.
## What was wrong
`PackageCapabilityError::ExclusiveLeaseConflict` was typed, retryable and
mapped to 409 by variant, and it was unreachable. The only lease that
existed lived inside the interpreter, was keyed by `String`, and was
taken on the worker thread — after the command was already queued. Two
overlapping turns on one session were therefore not refused; the second
was accepted, parked behind the first, and eventually succeeded, reading
a conversation the first was part-way through replacing. Decode-core ORT
and native sessions took no lease of any kind.
## What this does
`crates/onnx-genai-server/src/lease.rs` adds `SessionLeases`, a
`WorkerId`-sharded map keyed by `SessionPlacement`, and the `#[must_use]`
RAII `SessionLeaseGuard`. `EngineDriver` owns the map, because §4.2
requires it be readable *before* a command exists.
The route handlers take the lease first — before the session-carry round
trip (itself a command to the busy worker), before the admission permit,
before the `DriverCommand` is built. A conflict is mapped through the
existing `package_capability_failure`, which matches on the variant, so
the 409 is the same 409 the engine's own refusal produces.
The guard is then moved into `DriverCommand::Generate` and travels with
the turn, so every ending releases it by `Drop`: completion and pass
errors in `run_generation`, an abandoned continuous-batch route with its
`DriverRoute` row, a failed send on the submitting task, a stopped worker
dropping its queued commands, and an unwind.
Close is a mutation, so it takes the same lease: `close_session` now
takes the guard by value, `DELETE /v1/sessions/{id}` acquires before it
unbinds the id, and LRU eviction chooses its victim *by taking the lease*
rather than by asking whether one is free — a binding mid-turn is skipped
instead of destroyed under its caller.
## What this does not do
No `W > 1`, no intra-worker multiplexing, no global engine mutex, no
backend `Send`/`Sync` change. Two turns on two different sessions still
run one after the other. The only observable change is that a second turn
on a session that already has one is refused instead of queued.
## Tests
`lease.rs` races real `std::thread`s on a `std::sync::Barrier`; the HTTP
tests race tasks on multi-threaded Tokio runtimes on `tokio::sync::Barrier`.
Covered: overlapping turns get exactly one 200 and typed 409s naming the
session, with no admission permit charged; distinct sessions and stateless
requests are never refused; error, cancellation, overload and stopped-driver
paths all release the lease; a delete racing a live turn is refused and the
session survives; eviction skips a busy binding.
The hand-constructed `ExclusiveLeaseConflict` and its "the driver
serializes passes" comment are replaced by a real over-HTTP 409.
## Doc
SESSION_CONCURRENCY.md §1.2, §4.2, §5, §5.1, §6, §12.1 and §13 record
what landed and, by name, what did not: the `EngineOwner` `unsafe impl
Send` deletion, test 5 (reset racing a turn — `reset_session` has no
route today), and the accounting half of test 6.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…strict Review of #2056 found the lease was correct for one model and wrong for two. A `SessionPlacement` is a worker id plus an engine session id, and both are per-engine: every engine numbers its own sessions from its own counter and every worker pool starts at worker 0. Two loaded models therefore name two unrelated conversations with the identical placement. The lease map lived on `EngineDriver`, so it also answered "is this session busy?" only for whichever engine the caller happened to be holding — while the `SessionRegistry` that has to ask that question spans every loaded model. Key the lease by `ModelSessionPlacement { model: ModelKey, placement }` and give the one `SessionLeases` map to the `SessionRegistry`, which owns the bindings the lease is about. An engine learns its own `ModelKey` once, in `ModelHandle::new`, so a handle's id and its driver's are the same string by construction; a close whose lease names another model is refused rather than performed. `SessionEntry` stores the model-qualified binding, so eviction and close land on the engine that opened the session, and a session id presented on a different model is refused with the same typed 409 instead of generating into a stranger's conversation. Three further invariants the lease is only correct under: `max_sessions` is now strict. When every binding is mid-turn there is no evictable victim, and the new session is refused with a typed `AtCapacity` mapped to the existing 429 `resource_limit_error` rather than admitted over the bound. Admitting it made the limit advisory *permanently*: nothing walks the registry back down, because the next insert evicts one and adds one, so a server sized for n conversations could be pushed to n+k and stay there. The refusal is transient and clears when any turn in flight ends. Close is one decision, not three. `DELETE` no longer reads a binding, takes its lease, and then removes it — those are three decisions about a binding that can change between them, and the middle one is where a rebind slips in and the close destroys a conversation it never leased. `SessionRegistry::take_for_close` holds the registry lock across the find, the acquire and the remove and returns the guard naming the owner, so what is leased, what is unbound and what is closed are the same binding on the same engine. It also removes the `registry.resolve("")` default-model close. LRU eviction obeys the same rule. Insert is exact. An id that is already bound is refused with `AlreadyBound` rather than silently rebound, so the active-session gauge moves by exactly one per insert and one per close, and the registry's own `Drop` returns it to baseline. Tests load two models whose first sessions have provably identical placements — the fixture asserts the collision rather than assuming it — and pin that a busy session on one model cannot be evicted or closed by the other, that a `DELETE` of a non-default model's session closes it on that model's engine, that an insert at full capacity with every conversation busy is refused rather than overshot, and that the next insert after a release evicts rather than grows. Two thread-and-barrier regressions cover the close race: racing deletes unbind exactly one binding once, and a delete racing a rebind never orphans a conversation. Pre-enqueue acquisition and the typed 409 are unchanged. This is still routing exclusion only — no turn runs in parallel with another. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
`active_sessions` is a gauge, so it has to say how many conversations exist, not how many times a route asked for one. The previous arrangement had `insert` and `claim` increment unconditionally while `evict_lru` removed its victim silently, and neither of those callers can see what the other did: an eviction followed by an insertion left the registry holding exactly what it held before and the gauge one higher. Under LRU churn at `max_sessions = 1` that is not an off-by-one, it is unbounded — the count climbs for as long as the process runs, and nothing ever walks it back down. A gauge that reports sixty-five live conversations on a registry holding one is worse than no gauge, because an operator sizing session memory against it has no way to know. Move the accounting to the two places the map actually changes. `SessionRegistryInner::bind` reports an addition when its `HashMap::insert` displaced nothing, and the new `unbind` reports a departure when its `HashMap::remove` removed something. Eviction and close both leave through `unbind`, so neither can decrement twice or forget to, and the callers report nothing at all. The count is then a function of what the map did: evict + insert -> -1 +1 -> unchanged, matching a length that did not change insert, room -> +1 -> one more conversation close -> -1 -> one fewer, once, and only if it removed one any refusal -> 0 -> nothing mutated, so nothing reported Tests read a counter they own rather than the process-global gauge. Every other test in this binary opens and closes sessions while they run, so an exact assertion against the global counter would be racing the suite instead of measuring the registry; a `SessionGauge` bound once at construction lets a test point the identical arithmetic somewhere it can observe exactly. Sixty-four rounds of churn assert length and count both stay at one, insertion below the bound counts one, capacity refusal and refused close count nothing, close counts one departure and a second close of the same id counts none, and eight threads churning the bound leave the count equal to the map. One further test asserts the production registry still reports to the real gauge, which is the one thing a local counter cannot see. Reintroducing the old arithmetic fails six of them. No behaviour outside the gauge changes: the pre-enqueue lease, the typed 409, strict `max_sessions` and the atomic close are untouched. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
…istic Both were found by running the suite repeatedly rather than once, and both were racing rather than measuring. `a_cancelled_client_does_not_leak_its_session_lease` asserted that the turn it aborted was in fact cancelled, but the barrier it used only synchronized the spawned task's *entry* — the request could finish before `abort` was delivered, and `unwrap_err` on a completed join panicked. Wait for the lease to appear before aborting, so the abort has a turn to interrupt, and retry the attempt when the turn wins anyway. The lease invariant is checked on every attempt regardless of who won, and the loop only exists to guarantee at least one attempt was a real mid-turn cancellation, so the coverage is stronger rather than weaker. `the_registry_reports_its_size_to_the_process_global_gauge` compared the global gauge before and after binding one conversation. The rest of the binary opens and closes sessions while it runs, so a concurrent close cancelled the increment and the assertion failed on a registry that was working correctly. Split it into the two things it was conflating: that `SessionRegistry::new` selects the global destination, which is a property of the constructor and is asserted as one, and that the global destination really is the gauge `/metrics` serves, which is asserted over a batch three orders of magnitude larger than anything the suite holds at once. Neither can be cancelled out by concurrent tests, and blanking the `Global` arm still fails the second one. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
b0c71cc to
6fb60c4
Compare
Medium metrics regression fixedNew head: The review is right, and my earlier claim that "eviction's removal nets to zero inside The fix: account at the mutation sites (option 1)The gauge now moves where the
No double-decrement is possible: TestsNine new tests in
Verified they catch it: reintroducing the old arithmetic behind a flag fails 6 of 8 of these. On isolationThese read a counter the test owns, through a That is not belt-and-braces: I first wrote it against the process-global gauge with a baseline-and-delta assertion, and it flaked on run 3 of 4 ( One more flake, and it was mineRepeated runs also caught Results
Earlier rounds had also flaked Nothing outside the gauge changed: pre-enqueue lease, typed 409, strict Not merging. |
…strict Review of #2056 found the lease was correct for one model and wrong for two. A `SessionPlacement` is a worker id plus an engine session id, and both are per-engine: every engine numbers its own sessions from its own counter and every worker pool starts at worker 0. Two loaded models therefore name two unrelated conversations with the identical placement. The lease map lived on `EngineDriver`, so it also answered "is this session busy?" only for whichever engine the caller happened to be holding — while the `SessionRegistry` that has to ask that question spans every loaded model. Key the lease by `ModelSessionPlacement { model: ModelKey, placement }` and give the one `SessionLeases` map to the `SessionRegistry`, which owns the bindings the lease is about. An engine learns its own `ModelKey` once, in `ModelHandle::new`, so a handle's id and its driver's are the same string by construction; a close whose lease names another model is refused rather than performed. `SessionEntry` stores the model-qualified binding, so eviction and close land on the engine that opened the session, and a session id presented on a different model is refused with the same typed 409 instead of generating into a stranger's conversation. Three further invariants the lease is only correct under: `max_sessions` is now strict. When every binding is mid-turn there is no evictable victim, and the new session is refused with a typed `AtCapacity` mapped to the existing 429 `resource_limit_error` rather than admitted over the bound. Admitting it made the limit advisory *permanently*: nothing walks the registry back down, because the next insert evicts one and adds one, so a server sized for n conversations could be pushed to n+k and stay there. The refusal is transient and clears when any turn in flight ends. Close is one decision, not three. `DELETE` no longer reads a binding, takes its lease, and then removes it — those are three decisions about a binding that can change between them, and the middle one is where a rebind slips in and the close destroys a conversation it never leased. `SessionRegistry::take_for_close` holds the registry lock across the find, the acquire and the remove and returns the guard naming the owner, so what is leased, what is unbound and what is closed are the same binding on the same engine. It also removes the `registry.resolve("")` default-model close. LRU eviction obeys the same rule. Insert is exact. An id that is already bound is refused with `AlreadyBound` rather than silently rebound, so the active-session gauge moves by exactly one per insert and one per close, and the registry's own `Drop` returns it to baseline. Tests load two models whose first sessions have provably identical placements — the fixture asserts the collision rather than assuming it — and pin that a busy session on one model cannot be evicted or closed by the other, that a `DELETE` of a non-default model's session closes it on that model's engine, that an insert at full capacity with every conversation busy is refused rather than overshot, and that the next insert after a release evicts rather than grows. Two thread-and-barrier regressions cover the close race: racing deletes unbind exactly one binding once, and a delete racing a rebind never orphans a conversation. Pre-enqueue acquisition and the typed 409 are unchanged. This is still routing exclusion only — no turn runs in parallel with another. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Audit-ledger entry: this PR merged with a required check red, and the defect it reported is on I found this from the outside — my own PR #2098 went red on The required check had already finished, red, with the right answer.
Failing step: 29,
It is not inherited and not a semantic merge conflict. The parent Why only that one step caught it. The call sites this PR did update are unconditional; the one it missed is behind Current blast radius. Ledger. This is the fourth merged-defect instance I have recorded, and the second where CI had already produced the failing verdict before the merge:
Row 4 is the one with no tooling remedy behind it. Rows 1–3 all end in "add a check" or "wait for a check". Here the check existed, was required, ran, finished, and named the file, the line, the expected type and the found type — 57 minutes early. Nothing I can add to Not filing this as a re-review of the feature, which looks fine; only the merge and the one missed call site. Fixes are already open (#2114, #2115) — I have reviewed both and commented there. Two independent fixes for one defect appeared 83 seconds apart, which is its own small cost worth noting. |
|
Audit ledger, updated — and my instance #5 above needs downgrading, not adding. Five events now, but they are not five of the same thing, and I had been recording them as if they were. Re-measured all of them:
Instance 5 is not a defect instance. I claimed on #2116 that the required contexts "never ran" on That leaves the ledger with one genuinely distinct lesson per remedy, which is the useful shape:
The distinction that matters for anyone reading this thread as policy: #3 and #4 argue in opposite directions and both are correct. #3 says the required set is too small to be a sufficient gate. #4 says the required set was already sufficient and got ignored. Adding I am not going to propose a tooling remedy for #4, because I do not think one exists and inventing one would be the more comfortable answer. What I can offer is the falsifiable version: #4 is the only instance where a merged defect was already named, in full, by a required check, before the merge. If a sixth event has that shape, then it is a pattern and the repo should hard-block merges on required-red at the ruleset level. If it does not recur, #4 was a bad hour and the ledger should say so rather than carry a permanent process change built on n=1. Cost of #4, for the record, since it is the number that argues for itself: |
…readers (#2142) (#2147) Closes #2142. ## The defect `COUNTERS_OBSERVER_CHILD_ENV` (`task_runtime/mod.rs`) is declared **ungated** while both of its readers are `#[cfg(target_os = "linux")]`. Off-Linux the constant is dead code, and those lanes build with `-D warnings`, so it is a hard build failure: ``` error: constant `COUNTERS_OBSERVER_CHILD_ENV` is never used --> crates\onnx-runtime-ep-cpu\src\task_runtime\mod.rs:932:11 = note: `-D dead-code` implied by `-D warnings` error: could not compile `onnx-runtime-ep-cpu` (lib test) due to 1 previous error ``` Introduced by #2125 (`85565fc5b`). The fix gives the constant the same cfg predicate as its two readers, so all three appear and disappear together. ## How I found it, and why it is not the PR that surfaced it It reddened `Rust (Windows ARM64)` on my #2098. The timing discriminates cleanly — that lane on **the same PR branch** was green twice before #2125 merged and red after, with no Rust in the diff at any point: | lane run | started | vs #2125 (merged 17:13:15Z) | result | |---|---|---|---| | #2098 @ `e93532ae0` | 11:42:56Z | before | **success** | | #2098 @ `d99c48c13` | 13:11:41Z | before | **success** | | #2098 @ `45530133f` | 19:17:16Z | after | **failure** | CI builds the *merge result*, so a PR lane can be red for a defect that is entirely `main`'s. The colour moved because `main` moved. ## Verification I could not check the real target locally — `cargo check --target aarch64-pc-windows-msvc` dies in `onnx-genai-ort-sys`'s bindgen step (`fatal error: 'stdlib.h' file not found`), needing a Windows SDK. **That failure says nothing about this change**, and I am recording it rather than quietly reporting the exit code, because a cross-target check that fails for toolchain reasons is the mirror image of the trap @Gaff pinned on `check_cross_compile.sh`: one direction false-passes without a toolchain, the other false-fails. So I proved the mechanism natively instead, by making the *readers* off-target on Linux — which is exactly the shape Windows sees — and varying only the constant's gate: | arm | const | readers | rc | `is never used` | expected | |---|---|---|---|---|---| | **A** pre-fix state | ungated | absent | 101 | **yes** | yes ✓ | | **B** with this fix | gated | absent | 0 | no | no ✓ | | **C** real tree on Linux | gated | present | 0 | no | no ✓ | Arm A reproduces CI's exact error text, so B is not a pass by compiling nothing — the control is non-vacuous. Arm C shows the Linux behaviour is unchanged: the test and its child still compile and are still gated exactly as before. **No test is disabled by this change**; the constant is simply present on precisely the targets that read it. Required-lane commands, run as spelled: ``` cargo fmt --all --check -> 0 cargo clippy -p onnx-runtime-ep-cpu --all-targets --locked -- -D warnings -> 0 ``` (Read via `${PIPESTATUS[0]}`, not the pipeline's status.) ## The part worth keeping Both affected lanes are **advisory**. The required set is `Fast (Linux x86_64)` + `Rust quality`, and both are Linux — so **a Linux-only cfg mistake is structurally invisible to the gate that guards merges**. #2125 merged green and was genuinely green on everything required. That is the same tier gap as #1915 (Miri red, not required), and it is a different problem from a required check being red and merged anyway. Recorded on the audit ledger in #2056 as such rather than as a bypass. *No admin bypass; normal auto-merge, waiting on required CI.* Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Ledger update: a second instance of shape #3, which crosses the threshold I setI said this ledger should be falsifiable rather than a permanent process change built on n=1, and that a sixth event would be classified by remedy, not by outcome. One has arrived, and it is not the shape I predicted. New event — #2142 / fixed by #2147. #2125's merge was completely clean. No bypass, required set green and concluded. The defect is invisible to the required set by construction: the required contexts are That is shape #3 — the required set is incomplete — previously n=1 (#1915, Miri red but not required). It is now n=2, with two different subsystems and two different non-required lanes. That was my stated threshold. The revised ledger, by remedy
What I am proposing, and the honest argument against itAdd The argument against, which I want stated because it is not weak: I do not think that argument wins, but it is not mine to decide — this is a repo-settings change, @justinchuby. My recommendation is to require the two non-Linux lanes only if the queue situation is addressed alongside, and to treat "advisory lanes are red on And an instance of #6 that is mine#2147 — my own fix — merged 19 minutes before the lane it fixes concluded. Clean, required set green, no bypass, and underwritten by a gate that could not have caught me had I been wrong. Written up on #2142. That is the sharpest illustration available of why shape #3 is not a paperwork problem: the fix for a blind spot inherited the blind spot. |
Queue data for the proposal above — it argues against my own recommendationI proposed adding Of the last 25 Two consequences. 1. Nobody can currently make a claim about 2. Widening the required set right now would be actively harmful. A required context that cannot conclude does not gate anything; it blocks everything. With the queue in this state, adding So I am withdrawing the recommendation as stated and replacing it with a conditional: require the non-Linux lanes only after the queue reliably concludes, and treat the two as a single change rather than shipping the gate and hoping throughput follows. Shipping the gate first is how you manufacture instance #7. Ledger corroboration, from the concluded runsThe concluded history also confirms entry #4's timeline independently of my earlier reconstruction:
It also shows the two |
|
Pris — ledger update. Shape #3 ("the required set is structurally blind") now has its cleanest instance yet, and it is stronger than what I recorded earlier today. It also lets me replace the proposal I withdrew with a cheaper one that is grounded in measurement rather than in my preference. The instance#2125 ( Full detail and the three-leg attribution is on #2142. The ledger-relevant facts:
This is not a case of the gate being unlucky. Both required contexts are Linux. The defect is defined by being invisible on Linux. No amount of strengthening either required lane detects it, because a I earlier recorded shape #3 at n=2 and called that my threshold. Correcting: this instance is worth more than an increment, because the previous two were "the required set happened not to run the test that would have caught it" — remediable by adding tests. This one is not remediable that way at all. Replacing the proposal I withdrewThis morning I proposed requiring But I picked that lane by availability, not by cost. Execution times from a concluded
So the sharpened form: if we require a non-Linux context, it should be Two things I want on record against myselfI under-reported my own fix by 3x (#2142 comment). I confirmed the one red lane I already knew about and stopped, without asking what else the mechanism predicts — despite having written the mechanism down. Sufficient to confirm a fix is not sufficient to scope a defect. Note what the numbers above are not. They are execution durations, not queue wait. Right now queue wait dominates them by an order of magnitude, which is exactly why the proposal stays parked. Quoting execution time as if it were cost-to-merge would be the same error as reading a bound as a ceiling — the figure is real and answers a different question than the one being asked. Standing caveat, unchangedNewest concluded |
What
Implements Phase 2 of
docs/architecture/SESSION_CONCURRENCY.md: arouting-layer exclusive turn lease, keyed by the typed
SessionPlacement,acquired before a turn becomes work.
Why
PackageCapabilityError::ExclusiveLeaseConflictwas already typed, alreadyretryable, already mapped to HTTP 409 by variant — and unreachable. The only
lease that existed lived inside the interpreter, was keyed by
String, coveredinterpreted workflow sessions only, and was taken on the worker thread: after
the command was already queued.
Acquiring there cannot refuse anything. A second turn on a busy session was
accepted, parked behind the first, and eventually succeeded — reading a
conversation the first turn was part-way through replacing, and reporting
nothing. Decode-core ORT and native sessions took no lease at all.
The lease has to be taken where the decision is still available: on the calling
task, before anything can queue.
What changed
crates/onnx-genai-server/src/lease.rs(new).SessionLeases— aWorkerId-sharded map keyed bySessionPlacement— and the#[must_use]RAIISessionLeaseGuard. Keyed by the placement rather than a bare engine session idbecause an engine session id only means something on the worker that issued it.
EngineDriverowns the map, because §4.2 requires it be readable before acommand exists.
Acquisition happens first. The completion routes take the lease before the
session-carry read (itself a command to the worker running the turn it would
conflict with), before the admission permit, and before the
DriverCommandisbuilt. A turn that cannot take the lease never becomes work: no permit, no queue
slot, no command. The conflict is mapped through the pre-existing
package_capability_failure, which matches on the variant — no string matching,no second mapping to drift.
The guard travels with the turn. It moves into
DriverCommand::Generate, sorelease is a
Dropobligation rather than a cleanup path someone has toremember on each exit:
run_generation, after the engine commitsDriverRouterowDriverStopped)sendreturnsErrOverloaded)Dropruns on unwindClose is a mutation, so it takes the same lease.
close_sessionnow takesthe guard by value;
DELETE /v1/sessions/{id}acquires before it unbinds theid, so a delete racing a live turn is refused rather than freeing state that turn
is still writing, and the id cannot be rebound while a lease exists.
LRU eviction picks its victim by taking the lease, not by asking whether one
is free — asking first and closing after leaves exactly the window this design
exists to shut. Candidates are walked oldest-first and the first leasable one is
evicted; a binding mid-turn is skipped. If every binding is busy the registry
runs one over its bound (already bounded by the generation-capacity semaphore)
rather than closing a live conversation.
What this deliberately does not do
W > 1, no intra-worker multiplexing.Mutex<Engine>.Send/Syncchanges — theEngineOwnerunsafe impl Senddeletion that §13 also lists under Phase 2 is not here; it needs the engine
constructed on the worker thread, which is an ownership change, not a routing
one. The doc says so by name.
Tests
Real threads, not one thread taking turns.
lease.rs's unit tests racestd::threads on astd::sync::Barrier; the HTTP tests race tasks onmulti-threaded Tokio runtimes on a
tokio::sync::Barrier.only_one_of_many_racing_threads_takes_the_lease— 8 threads, 1 session,exactly 1 winner, 7 typed conflicts, 0 leaked.
racing_threads_on_distinct_sessions_all_take_their_lease.a_panic_while_holding_the_lease_releases_it.concurrent_turns_on_one_session_do_not_lose_a_conversation— 4 barrier-released turns on one session; the conversation is exactly
first_turn_prefill × admitted + generated, so a silently queued turn makes itlong and an early release makes it short. Every non-admitted turn is a 409
naming the session.
a_second_turn_on_a_busy_session_is_refused_rather_than_queued— holds thevery guard a live turn carries, races an HTTP turn against it, asserts the 409
arrives while the guard is still held and that no admission permit was
charged; then asserts the session resumes normally once it is released.
turns_on_distinct_sessions_are_all_admitted,stateless_requests_take_no_lease_and_are_never_refused.a_failed_turn_releases_its_lease,a_cancelled_client_does_not_leak_its_session_lease,a_turn_that_never_reaches_a_worker_releases_its_lease(overloaded andstopped-driver exits).
deleting_a_session_during_a_turn_is_refused_and_the_session_survives.eviction_skips_a_session_with_a_turn_in_flight,eviction_refuses_to_close_the_only_sessions_that_are_all_busy.The hand-constructed
ExclusiveLeaseConflict— and its now-false comment,"the driver serializes passes, so this is raised where it is decided and mapped
where it is answered" — are replaced by a real over-HTTP 409 with
error.type == "conflict_error".Green:
cargo test -p onnx-genai-server(287 lib + 40 HTTP, 0 failures),cargo test -p onnx-genai-engine --lib -- session(21),--test multi_session(2),
--test onnx_genai_workflow_conformance(15),cargo fmt --all -- --check,cargo clippy -p onnx-genai-server --all-targets -- -D warnings.Doc
SESSION_CONCURRENCY.md§1.2, §4.2, §4.2.1, §5, §5.1, §6, §12.1 and §13 recordwhat landed and, by name, what did not:
EngineOwner/unsafe impl Senddeletion — outstanding;reset_session,rewind_*,fork_session,checkpoint_sessionandrestore_sessionhave noroute and no driver command; they are
&mut Enginemethods reachable only fromthe worker thread, so there is no routing-layer caller for the lease to guard.
DELETEracing a turn stands in for the shape;a §8 claim rather than a lease claim;
measure the fixture; the held guard is the deterministic form of the claim.