Repository navigation
fix(server): give the native-backend test the session lease the driver now requires - #2115
Conversation
|
Review: Your comment says the lease release is imminent rather than done when the result arrives. That is exactly right, and the mechanism is one line further than the comment goes, so recording it here — if let Some(mut route) = routes.remove(&handle.id) {
route.metrics.result(...);
let _ = deliver_driver_event(&route.events, DriverEvent::Finished(result), DELIVERY_GRACE);
} // <-- `route` drops HERE, and `_lease` (driver.rs:273) with it
The bounded poll keeps the assertion's force — it still fails, with I have posted the same finding on #2114, which fixes the same defect with a single-shot Two notes, neither blocking:
Context for anyone arriving here: the defect is #2056's, which merged at 12:03:16Z with Verified statically only: I did not compile |
…r now requires Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
9e87996 to
7a7a76f
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2115 +/- ##
==========================================
+ Coverage 80.29% 80.70% +0.40%
==========================================
Files 426 429 +3
Lines 205244 215150 +9906
Branches 205244 215150 +9906
==========================================
+ Hits 164801 173628 +8827
- Misses 34792 35748 +956
- Partials 5651 5774 +123
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
|
Correction to my review, with measurements. The race is real but I described the wrong widener, and I could not reproduce it in 120 runs. I said the window "widens on a loaded 2-core runner". I went and measured instead of leaving that as an assertion, and it does not hold. Recording it because #2114 was closed partly on my say-so. What I ran. Built the test CI never runs (see below), then reverted the poll to #2114's single-shot
Why my mechanism was wrong. Single-CPU binding makes failure less likely, not more: the driver thread runs to the end of the The real widener is a full output channel, and it is much bigger than I claimed. let mut pending = match events.try_send(event) { Ok(()) => return Ok(()), ... };
let deadline = Instant::now() + grace;
loop {
if Instant::now() >= deadline { return Err(DriverDeliveryError::Stalled); }
thread::sleep(DELIVERY_RETRY_INTERVAL);
...
}The So the corrected verdict: the ordering hazard I described is real and the poll is still the right code — but the reachability argument I gave for it was wrong, and a reader of my review would have gone looking for the flake on a busy runner and not found it. A test with a slower consumer or a longer generation is where a single-shot The finding that outlasts this. While setting the experiment up I checked what CI does with this test: No lane runs Same family as #2058: a code path exactly one step can see. I will open this separately rather than bury it in a merged PR thread. Everything above was run under |
## Summary - migrate CPU/plugin `NonMaxSuppression` to #2101's host-only `KernelSizedOutput` contract - refactor selection into one owned result shared by native `execute` and plugin Compute - claim NMS in the plugin only after adding it to the exact kernel-sized strategy census - tighten NMS mixed-edge dtype advertisement and slot constraints ## Semantics Implements the existing ONNX NonMaxSuppression opset-10 behavior without changing selection ordering: - boxes `[batch, spatial_dimension, 4]`, Float32 - scores `[batch, classes, spatial_dimension]`, Float32 - optional positional scalars: Int64 `max_output_boxes_per_class`, Float32 `iou_threshold`, Float32 `score_threshold` - `center_point_box` 0 and 1 - output Int64 `[num_selected_indices, 3]` rows `[batch, class, box]` - strict score filter (`score > threshold`), stable lower-index tie ordering, complete per-batch/per-class ordering - empty boxes/classes/batches, zero max output, NaN scores, and threshold boundary behavior covered Absent and trailing optional inputs remain positional. Strided host boxes/scores are materialized by the existing CPU dense helpers. Device inputs fail at the host-accessibility gate before selection; there is no implicit payload D2H. ## Single execution and materialization `compute_owned_output` performs validation, dense materialization, selection, shape construction, and byte encoding once. Native execution validates/copies that owned result into its preallocated tensor. Plugin Compute receives the same owned `KernelSizedOutput`, asks ORT for the exact `[selected,3]` allocation, and performs one final byte copy. Representative test geometry: 1 batch × 2 classes × 3 boxes selected 4 rows. Selection counter: **exactly 1**. Materialized output: **4 × 3 × 8 = 96 bytes**, one copy. ## Validation - CPU NMS targeted suite — **8 passed** - legacy overlapping NMS regression — **1 passed** - generic KernelSizedOutput plugin unit suite — **4 passed** - plugin shape/strategy census — **5 passed** - real ORT plugin NMS E2E — **1 passed**; assignment `ours=["NonMaxSuppression"]`, dynamic shape `[4,3]`, exact Int64 rows - real ORT plugin Unique regression — **1 passed** - `cargo clippy -p onnx-runtime-ep-cpu -p onnx-runtime-ep-plugin -p onnx-runtime-ep-cpu-plugin --all-targets -- -D warnings` — passed - `cargo fmt -p onnx-runtime-ep-cpu -p onnx-runtime-ep-plugin -p onnx-runtime-ep-cpu-plugin` — passed ## Mutation evidence Each mutation was applied independently, caught, and reverted: - inverted IoU keep comparison -> overlapping-box regression failed - ignored score threshold -> strict threshold/NaN test failed - swapped center-box width/height conversion -> center-point test failed - ran selection twice -> once-only counter test failed (`2 != 1`) - skipped materialization copy -> real ORT E2E returned zero rows and failed exact output - removed NMS from the census -> exact census test failed - emptied the census -> nonempty guard failed ## Host-only limitation / CUDA follow-up This PR is deliberately CPU/plugin single-execution only. Kernel-sized outputs are host-only today; CUDA NMS must wait for the CUDA Unique device-sized-output policy to prove device allocation/materialization ownership. No CUDA NMS implementation or device payload copy is introduced here. ## Performance No throughput claim is made; no A/B benchmark was run. The structural improvement is removal of the formerly required double-algorithm path: the plugin now selects once and copies only the final owned output bytes once. ## Latest rebase / Rust quality attribution (2026-08-25) Rebased again onto latest `origin/main`, which now contains #2115 merge commit `0be2d23fe6f719a5cc70fe91e53a80d26f7f2739`. Verified directly: ```text git merge-base --is-ancestor 0be2d23 HEAD # exit 0 ``` The PR diff remains exactly these six NMS/plugin files and no server/engine files: - `crates/onnx-runtime-ep-cpu/src/kernels/{mod.rs,selection.rs}` - `crates/onnx-runtime-ep-plugin/src/compute.rs` - `crates/onnx-runtime-ep-cpu-plugin/tests/{plugin_ort_e2e.rs,shape_inference_coverage.rs}` - `crates/onnx-runtime-ep-cpu-plugin/tests/fixtures/non_max_suppression_kernel_sized/model.onnx.textproto` The exact required native-backend Rust-quality command now passes with `-D warnings`: ```text cargo clippy --locked --all-targets \ -p onnx-genai-engine -p onnx-genai-server \ --features onnx-genai-engine/native-backend,onnx-genai-server/native-backend \ -- -D warnings ``` Post-rebase targeted evidence is also green: CPU NMS 8 passed; generic KernelSizedOutput 4 passed; plugin census 5 passed; real ORT NMS E2E 1 passed; real ORT Unique regression 1 passed. Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Rust qualityis a required check and it is red onmain. Every open PR in the repo inherits it. This restores it.What broke
#2056 changed
EngineDriver::generateto takeOption<SessionLeaseGuard>andclose_sessionto takeSessionLeaseGuard. One call site was not updated:How it reached
main— the gate worked and was overriddenI want to state this precisely rather than call it a CI gap, because it wasn't one.
Rust qualityran on feat(server): refuse overlapping turns on one session at the routing layer #2056's own head as run32837618182, job97771768920, and failed at 11:06:27Z with the two errors above — byte-identical to the failure onmainnow.29668c96), ~57 minutes later, with that required check red.So no semantic conflict, no two-PRs-green-independently story: the check named the exact file and lines, and the merge proceeded past it.
The contributing factor worth naming is why that red was easy to discount: the site is behind
#[cfg(feature = "native-backend")], so a defaultcargo testnever compiles it and the failure looks like somebody else's inherited red. Worth knowing that this same call site was repaired once before — #1944, forgenerate's third argument. Second decay, same cause.The fix
generategets a real lease. For the close path I first wrote the obvious thing and it was wrong, so the comment in the diff records why:collect_generation_resultreturns on theDriverEvent::Finishedevent (routes/completions.rs:925). The lease lives inDriverRoute._lease, and the route is dropped by the driver thread after it sends that event (driver.rs:1489). A returned result therefore implies the release is imminent, not that it has happened — a bare re-acquire races the driver thread. Bounded polling instead, matching the existing pattern attests.rs:5449.Falsified in both directions
Not merely green:
1 passed; 0 failedin 0.15sthe finished turn never released its lease, 5.37sThe 5.37s matters: it is the full 200 x 25ms budget, so the loop demonstrably polls rather than succeeding on its first attempt and reporting a property it never tested.
Commands run
The test executes here rather than only compiling —
tests/fixtures/tiny-native-sub4-engineis present in-tree.Scope
One test function. No production code, no API change. All runs bounded to
taskset -c 24-31underscripts/hostlock.sh; these are compile/test runs, not measurements, and I make no claim about host quietness.