Repository navigation
Slice 6: device-side expert-route telemetry design + inert proof harness (#1810) - #1884
Conversation
…ess (#1810) Adapt FreeToken's device-side expert-route observation to onnx-genai's QMoE/BlockQuantizedMoE as an inert, observer-only telemetry contract, without touching PMM/VMM or CUDA-graph authorities. Design + test-only proof harness; no production residency/lifecycle wiring. Independent of PR #1854 (Slice 5): touches none of its files; the harness is a new integration test (its own compilation unit, no src/ edit). - docs/memory/EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md: 8-section design — where expert IDs live on device and why the host is blind; fixed-capacity GPU- resident telemetry contract (route bitmap + bounded dedup queue + u32[6] header {epoch,request,device,overflow,poison,count}); coarse-boundary state machine (produce at launch, consume after stream completion, fail closed; forbids cuMemMap/cuMemSetAccess during capture/replay); one-authority rule (PMM/VMM owns mapping; policy only chooses a desired set); FreeToken copied-vs-rejected comparison; cost model + falsifiable GO/NO-GO gates. - crates/onnx-runtime-ep-cuda/tests/expert_route_telemetry_probe_gpu.rs: CPU oracle + CUDA bitmap/dedup/overflow/poison/epoch/capture-replay/isolation tests + Trap-4-separated microbench. Uses only the public CudaRuntime API. Validated on idle A100 (GPU 5, CUDA 13.0, driver 580): all 8 tests pass; capture/replay re-accumulates each replay's real routes (VA stable, epoch 1->3, stale detected); overflow/poison/foreign-request/foreign-device fail closed. Independent code-review (not Deckard) found no high-confidence bugs. No speedup claimed (measurement-discipline). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1884 +/- ##
=======================================
Coverage 72.56% 72.56%
=======================================
Files 12 12
Lines 5231 5231
Branches 5231 5231
=======================================
Hits 3796 3796
Misses 1307 1307
Partials 128 128
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
Cycle 20 — PR #1884 Independent Review (Slice 6 route-observation harness)Reviewer: Roy (Lead) Note: measurements were run only after Cycle 19's #1854 GPU tests had fully Scope verification (blocking-if-violated — clean)
Source-citation accuracy (spot-checked, not trusted)Verified by direct file reads, not by trusting the doc: Technical contract verification
CPU oracle / non-vacuityEvery GPU assertion diffs against Measurement discipline (independently re-run, not trusted)Re-ran on idle GPU 5 (A100, confirmed 4 MiB/210 MHz before start via
GO/NO-GO gate framing (no overclaim)The design's own G1 threshold (§6.4: "≤2 µs at decode") is nominally exceeded Non-blocking observations (do not require revision)
Neither observation touches correctness, scope, or any claim this PR actually ConclusionAll items in the review checklist check out: strict scope (zero #1854 overlap, APPROVE HARNESS FOR MERGE. |
…eration step (#1902) ## Context This is a follow-up investigation, not a fix for a currently-broken state. On fresh `main` (post #1884/#1893), both `tests/fixtures/tiny-deepseek-v4-qmoe` and `tests/fixtures/tiny-glm52-full-attention` already load and decode correctly — the missing-`policies/*.onnx` problem I'd flagged in a prior cycle was already fixed by **#1883** (re-emitting `inference_metadata.yaml` with `migrate_model_io --reemit`, dropping ten hand-authored, never-real auxiliary policy components) and **#1888** (test-matrix follow-up). That fix is correct, already reviewed, and this PR does not revisit it. ## What this PR actually fixes Regeneration **determinism**. Re-running `generate.py` (and its `--full-attention` variant) against each fixture's pinned Mobius commit reproduces `model.onnx.textproto`/`tokenizer.json` byte-for-byte, but **not** `inference_metadata.yaml` — it reproduces the pre-canonicalization document `write_onnx_genai_config` emits, not the canonical single-`decoder`-component shape actually committed. This is true for **all three** tiny fixtures (not just the two #1883 touched), meaning a naive future regeneration would silently reintroduce the exact bug #1883 fixed. Fix: document the required two-step recipe (`generate.py`, then `migrate_model_io --reemit <output-dir>`) directly in each generator's docstring. **Docstring-only change** — no functional/behavior change, no fixture artifacts touched. ## Verification (2026-08-23, idle A100, `CUDA_VISIBLE_DEVICES` pinned, `hostlock`-gated) - `generate.py` alone vs committed: textproto/tokenizer byte-identical, metadata differs — confirmed for all 3 fixtures. - `generate.py` + `migrate_model_io --reemit` vs committed: metadata byte-identical — confirmed for all 3 fixtures. - `cargo test -p onnx-genai-metadata --test decoder_recognizer_agreement`: 14/14 pass (existing matrix from #1888 already guards the committed shape; no new test needed for a docstring-only change). - `cargo test -p onnx-genai-engine --features native-cuda --test deepseek_v4_tiny_qmoe_e2e --test glm_tiny_full_attention_e2e --test glm_tiny_qmoe_native_cuda_e2e`: **12/12 pass**, CPU/CUDA tokens byte-identical to `ANCHOR_IDS` for all 3 fixtures. - DeepSeek-V4 CUDA: `captures=1 replays=10 fallbacks=0` (native e2e test) / `captures=5 replays=310 fallbacks=0` (profile_native, n=3 steady). GLM full-attention CUDA: `captures=1 replays=10 fallbacks=0` in a small run / `captures=3 replays=186 fallbacks=0` (n=3 steady) — capture state recorded, not forced. GLM IndexShare CUDA: `captures=0 replays=0 fallbacks=0`, growing-KV-prefix decline (unchanged from prior cycle). Independent review (code-review agent) caught and this PR fixes: a `\` vs `\\` line-continuation bug that silently collapsed a documented shell command onto one line in the deepseek docstring, and confirmed the default GLM-5.2 IndexShare variant also needs the reemit step (verified directly via regenerate+reemit-vs-committed diff, since #1883 never touched that fixture so there's no "pre-1883 blob" to diff against for it). ## Scope No fixture regeneration performed or needed. No Mobius changes. No coarse-residency files touched. Independent of #1854. Closes the "regenerate missing policies" instruction as **superseded** — see full task report for the baseline-measurement portion of this cycle. Signed-off-by: Sebastian <sebastian@squad.local> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…f which was reported The `CUDA compile (Linux x86_64)` lane has been red since #1836 (`6e4b0ebb3`); last green was `cb81745b0`. The reported compile error was masking two further failures, so fixing it alone would have moved the lane from "red at build" to "red at the honesty script". Two more arrived on `main` while this PR was open, in #1884 and #1895. All five are fixed here; the lane cannot go green on any proper subset. 1. Unresolved import (the reported error). `content_preserving_transition_gpu.rs` imported `transition_granule_range_with_phase8_faults` unconditionally, but that item is gated `#[cfg(any(test, feature = "gpu-tests"))]`. The `test` arm does not cover an integration test: `tests/*.rs` are separate crates linking the library built *without* `cfg(test)`, so the item is genuinely absent in the `without-gpu-tests` configuration. Invisible to any local run that passes `--features gpu-tests`. Fixed with a two-arm local shim, mirroring `with_faults` in `crates/onnx-runtime-cuda-memory/tests/virtual_memory_gpu.rs`, which already solves this exact problem for `CudaVirtualBacking::with_driver_faults`. The gate itself is left alone: widening it would ship a driver fault injector that can force `cuMemUnmap`/`cuMemMap`/`cuMemSetAccess` to fail into production, and `required-features` is precisely what the honesty script exists to forbid ("exists only with gpu-tests enabled; CUDA tests must not hide from CPU inventory"). #1895 (`1be9f2cc2`) reached this file first, with `#![cfg(feature = "gpu-tests")]` on the whole target -- the option rejected above. It fixed the compile and traded it for an inventory failure on the same lane: `main` at `1be9f2cc2` reports "content_preserving_transition_gpu: Cargo reported no integration tests" plus 19 x "test exists only with gpu-tests enabled" (job `97266415620`). The target was not repaired, it was hidden. That line is removed here and the module doc records why, so the option is not re-tried a third time. 2. Five non-ignored tests in a `_gpu` target (4 honesty violations). `verify_cuda_test_honesty.py` requires every test in a `_gpu` target to be ignored, not passed, on a CPU-only runner -- that is how the suite is stopped from reporting green for a GPU it never touched. Three were genuinely CPU-only predicate tests over `verify_safe_point`. Moved to a new non-`_gpu` target rather than `#[ignore]`d: silencing them would have greened the lane by deleting coverage, and target naming is the escape hatch the script's own comment names for CPU-only tests in a CUDA crate. Moving them alone would *also* have stopped them running. A non-`_gpu` target is skipped by the honesty script, and `workspace_test_packages.py` deny-lists this crate from every offline lane, so the target would have been compiled and never executed. The CUDA lane therefore gains an explicit `--test content_preserving_transition` step, alongside the two CPU-only targets in this crate that already have one for the same reason. The other two asserted nothing -- `fault_injection_safe_point_recheck_rejects` is comments plus a `println!` and self-describes as "a documentation test"; `zero_len_is_committed_noop` `println!`s that the behaviour is "verified by implementation". Both passed unconditionally. Removed. The first has real sibling coverage in `fault_injection_recheck_safe_point_rejected_gpu`; the second does not -- the `len == 0` early return has no test anywhere. Deleting a test that asserts nothing loses no coverage, but the gap is pre-existing and real, and a genuine test needs a device (the early return still takes `&CudaRuntime`/`&mut CudaReservation`), so it is left for a GPU-capable change rather than papered over. Added `safe_point_accepts_a_clean_state`, because the three moved tests only ever assert `is_err()` and so all three survive a mutant that makes `verify_safe_point` reject unconditionally. Verified: under that mutant the inherited three pass and only the new test fails. 3. The manifest guard checked a proxy, not the property it claimed. The check was the literal substring `"gpu-tests = []"`, so #1860 turned it red by changing the value to `["onnx-runtime-cuda-memory/gpu-tests"]` -- a legitimate and necessary forwarding. Now matches the feature *key* whatever its value, extracted into `declares_gpu_tests_feature` so it is covered by the script's own `--self-test` fixtures, which it previously was not. 4. A GPU test in a `_gpu` target with no `#[ignore]` (#1895). `causal_conv_with_state_gpu::the_standard_domain_spelling_reaches_the_same_kernel` calls `require_cuda()` exactly like its two siblings but is missing their `#[cfg_attr(not(feature = "gpu-tests"), ignore = ...)]`, so it ran and failed on a CPU runner. Fixed by adding the sibling attribute verbatim; no design choice involved. 5. A CPU-only test inside a `_gpu` target (#1884). `expert_route_telemetry_probe_gpu::cpu_oracle_and_validator_self_consistent` passes in both configurations, which the script reports as "executed without gpu-tests". Its doc comment states the intent plainly: it "runs without a GPU so the reference cannot silently rot". `#[ignore]` would clear the checker by destroying exactly that property -- an ignored test runs nowhere -- so this takes the same route as defect 2. The pure-CPU oracle (`cpu_bitmap`, `cpu_dedup`, `consume_and_validate`, `synth_routes` and the header indices) moves verbatim to a shared `tests/expert_route_oracle/mod.rs`, which is a module and not a target: Cargo auto-discovers `tests/*.rs` and `tests/*/main.rs` only. Both the `_gpu` target and a new `expert_route_telemetry_probe` target declare it, so there is one copy of the oracle and the seven GPU tests keep diffing against the same code the CPU test checks. The new target gets its own `ci.yml` step, for the reason in defect 2. Verified live in its new home by mutation rather than by its own green: perturb the shift in `cpu_bitmap` -> FAILED, revert -> ok. "It compiles in the new file" is not evidence that it executes there, which is the defect an earlier review caught in defect 2's target. Each fix independently falsified by re-introducing it: import -> E0432; un-ignored test -> "must be ignored, not pass" (+3 inventory errors); manifest value -> "must define a gpu-tests feature"; missing `cfg_attr` -> "1 tests failed while checking ignored status"; relocated CPU test -> "executed without gpu-tests; CUDA tests must be ignored, not pass". Honesty script exits 0: 530 tests/79 targets identical in both configurations, 530 ignored without gpu-tests, 0 passed on this no-CUDA host. `cargo clippy -p onnx-runtime-ep-cuda --features cuda -- -D warnings` clean (the lane's exact command, both invocations of it); both explicit `ci.yml` test steps pass 4/4 and 1/1; fmt clean. Not fixed here: `cargo clippy -p onnx-runtime-ep-cuda --features cuda --all-targets -- -D warnings` fails on ~8 pre-existing lints in `_gpu` targets this PR does not touch. No lane runs that command -- both ep-cuda clippy steps omit `--all-targets` and `workspace_test_packages.py` deny-lists the crate -- so it is a real gap but a separate one, and folding it in would put unrelated files in a lane-restoration PR. Closes #1875 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…rtWeightGroup (Cycle-19) Roy's Cycle-19 review (#1854) found that ExpertWeightGroup members (fc1/fc2/fc3/scales tensors of one logical MoE expert bank) could each individually pass capability/alignment checks yet be assigned DIFFERENT hot/cold expert partitions, silently tiering only part of a logical expert across Device/Host. apply_residency_plan_at_boundary now adds a group-level preflight pass (4a) after the existing per-value precheck: for every multi-member group it derives one canonical (hot-expert-set, expert_count) from the first member that clears validation, and requires every other member to match it exactly (HashSet equality, after each member's own expert-count/index domain has already been validated). A group that disagrees on expert_count/index domain, hot-expert set, or contains an unsupported/missing-metadata member is rejected as a whole -- before any mutation -- exactly like the existing (approved) group capability fallback. For groups that agree, every member's byte ranges are then RE-DERIVED from the single group-authoritative hot set (via cold_ranges_for), never from that member's own independently-computed (already-verified equal) copy. This is a "cannot diverge by construction" fix: even if the equality check above were ever weakened, no two members of an agreeing group could still end up with different byte ranges, because there is only one hot set in play by the time ranges are computed. Scope, per the reviewer's narrowly-described gap: - No tensor-name heuristics; graph-derived expert_groups is unchanged. - No public API signature changes; ExpertWeightGroup, plan_residency, and validate_decision (weight.rs) are untouched -- the fix lives entirely in coarse_residency.rs's boundary-application logic. - Per-member byte geometry/offsets may still legitimately differ across fc1/fc2/fc3/scales; only the logical hot-expert-index set is required to be identical. - Default-off gate, same-device checks, Fatal/quarantine/progress reporting, and rollback fault behavior are unchanged; all previously approved tests (1-8) pass unmodified. New tests (coarse_residency_plan_gpu.rs, tests 9-13), one logical group with fc1/fc2-equivalent members per case: - equal hot sets in different order/with duplicates normalize and the whole group transitions together; - disagreeing hot sets between two supported members reject the whole group with zero side effects; - expert-count/index-domain mismatch rejects the whole group with a distinct reason, zero side effects; - an unsupported/missing-metadata member rejects the whole group before any mutation; - a successful group transition followed by a later member's Fatal (forced via the existing DriverFaultPlan phase-8 fault-injection machinery) rolls back the entire group, including the earlier member's already-committed range, with bit-identical content preserved throughout. Rebased onto latest main (through #1884/#1860); no conflicts. Local validation: build, clippy (-D warnings, filtered to touched files), fmt, ep-api lib (91/91), ep-cuda lib (543/543, 31 ignored), non-GPU coarse_residency_plan (4/4), and on an idle A100 (CUDA_VISIBLE_DEVICES=4): coarse_residency_plan_gpu (12/12, including all 5 new + all 7 previously-approved GPU tests), Slice2-4 GPU regressions composable_vmm_production_gpu/content_preserving_transition_gpu/ expert_bank_remap_cost_gpu (29/29), and qmoe_gpu correctness (46/46). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…f which was reported The `CUDA compile (Linux x86_64)` lane has been red since #1836 (`6e4b0ebb3`); last green was `cb81745b0`. The reported compile error was masking two further failures, so fixing it alone would have moved the lane from "red at build" to "red at the honesty script". Two more arrived on `main` while this PR was open, in #1884 and #1895. All five are fixed here; the lane cannot go green on any proper subset. 1. Unresolved import (the reported error). `content_preserving_transition_gpu.rs` imported `transition_granule_range_with_phase8_faults` unconditionally, but that item is gated `#[cfg(any(test, feature = "gpu-tests"))]`. The `test` arm does not cover an integration test: `tests/*.rs` are separate crates linking the library built *without* `cfg(test)`, so the item is genuinely absent in the `without-gpu-tests` configuration. Invisible to any local run that passes `--features gpu-tests`. Fixed with a two-arm local shim, mirroring `with_faults` in `crates/onnx-runtime-cuda-memory/tests/virtual_memory_gpu.rs`, which already solves this exact problem for `CudaVirtualBacking::with_driver_faults`. The gate itself is left alone: widening it would ship a driver fault injector that can force `cuMemUnmap`/`cuMemMap`/`cuMemSetAccess` to fail into production, and `required-features` is precisely what the honesty script exists to forbid ("exists only with gpu-tests enabled; CUDA tests must not hide from CPU inventory"). #1895 (`1be9f2cc2`) reached this file first, with `#![cfg(feature = "gpu-tests")]` on the whole target -- the option rejected above. It fixed the compile and traded it for an inventory failure on the same lane: `main` at `1be9f2cc2` reports "content_preserving_transition_gpu: Cargo reported no integration tests" plus 19 x "test exists only with gpu-tests enabled" (job `97266415620`). The target was not repaired, it was hidden. That line is removed here and the module doc records why, so the option is not re-tried a third time. 2. Five non-ignored tests in a `_gpu` target (4 honesty violations). `verify_cuda_test_honesty.py` requires every test in a `_gpu` target to be ignored, not passed, on a CPU-only runner -- that is how the suite is stopped from reporting green for a GPU it never touched. Three were genuinely CPU-only predicate tests over `verify_safe_point`. Moved to a new non-`_gpu` target rather than `#[ignore]`d: silencing them would have greened the lane by deleting coverage, and target naming is the escape hatch the script's own comment names for CPU-only tests in a CUDA crate. Moving them alone would *also* have stopped them running. A non-`_gpu` target is skipped by the honesty script, and `workspace_test_packages.py` deny-lists this crate from every offline lane, so the target would have been compiled and never executed. The CUDA lane therefore gains an explicit `--test content_preserving_transition` step, alongside the two CPU-only targets in this crate that already have one for the same reason. The other two asserted nothing -- `fault_injection_safe_point_recheck_rejects` is comments plus a `println!` and self-describes as "a documentation test"; `zero_len_is_committed_noop` `println!`s that the behaviour is "verified by implementation". Both passed unconditionally. Removed. The first has real sibling coverage in `fault_injection_recheck_safe_point_rejected_gpu`; the second does not -- the `len == 0` early return has no test anywhere. Deleting a test that asserts nothing loses no coverage, but the gap is pre-existing and real, and a genuine test needs a device (the early return still takes `&CudaRuntime`/`&mut CudaReservation`), so it is left for a GPU-capable change rather than papered over. Added `safe_point_accepts_a_clean_state`, because the three moved tests only ever assert `is_err()` and so all three survive a mutant that makes `verify_safe_point` reject unconditionally. Verified: under that mutant the inherited three pass and only the new test fails. 3. The manifest guard checked a proxy, not the property it claimed. The check was the literal substring `"gpu-tests = []"`, so #1860 turned it red by changing the value to `["onnx-runtime-cuda-memory/gpu-tests"]` -- a legitimate and necessary forwarding. Now matches the feature *key* whatever its value, extracted into `declares_gpu_tests_feature` so it is covered by the script's own `--self-test` fixtures, which it previously was not. The key is looked for only inside the `[features]` table. Review falsified the first version of this fix: a file-wide regex also accepts a *dependency* named `gpu-tests`, or one under `[target.'cfg(...)'.dependencies]` -- the same proxy mistake in a new costume, inside the fix for that mistake. Twelve fixtures now, including both false-positive shapes. 4. A GPU test in a `_gpu` target with no `#[ignore]` (#1895). `causal_conv_with_state_gpu::the_standard_domain_spelling_reaches_the_same_kernel` calls `require_cuda()` exactly like its two siblings but is missing their `#[cfg_attr(not(feature = "gpu-tests"), ignore = ...)]`, so it ran and failed on a CPU runner. Fixed by adding the sibling attribute verbatim; no design choice involved. 5. A CPU-only test inside a `_gpu` target (#1884). `expert_route_telemetry_probe_gpu::cpu_oracle_and_validator_self_consistent` passes in both configurations, which the script reports as "executed without gpu-tests". Its doc comment states the intent plainly: it "runs without a GPU so the reference cannot silently rot". `#[ignore]` would clear the checker by destroying exactly that property -- an ignored test runs nowhere -- so this takes the same route as defect 2. The pure-CPU oracle (`cpu_bitmap`, `cpu_dedup`, `consume_and_validate`, `synth_routes` and the header indices) moves verbatim to a shared `tests/expert_route_oracle/mod.rs`, which is a module and not a target: Cargo auto-discovers `tests/*.rs` and `tests/*/main.rs` only. Both the `_gpu` target and a new `expert_route_telemetry_probe` target declare it, so there is one copy of the oracle and the seven GPU tests keep diffing against the same code the CPU test checks. The new target gets its own `ci.yml` step, for the reason in defect 2. Verified live in its new home by mutation rather than by its own green: perturb the shift in `cpu_bitmap` -> FAILED, revert -> ok. "It compiles in the new file" is not evidence that it executes there, which is the defect an earlier review caught in defect 2's target. Each fix independently falsified by re-introducing it: import -> E0432; un-ignored test -> "must be ignored, not pass" (+3 inventory errors); manifest value -> "must define a gpu-tests feature"; missing `cfg_attr` -> "1 tests failed while checking ignored status"; relocated CPU test -> "executed without gpu-tests; CUDA tests must be ignored, not pass". Honesty script exits 0: 530 tests/79 targets identical in both configurations, 530 ignored without gpu-tests, 0 passed on this no-CUDA host. `cargo clippy -p onnx-runtime-ep-cuda --features cuda -- -D warnings` clean (the lane's exact command, both invocations of it); both explicit `ci.yml` test steps pass 4/4 and 1/1; fmt clean. Not fixed here: `cargo clippy -p onnx-runtime-ep-cuda --features cuda --all-targets -- -D warnings` reports 23 pre-existing errors and dies on `index_share_gpu` and `qmoe_zero_copy_cold_expert_spike_gpu` before reaching the rest, so 23 is a floor, not a total. None are in a file this PR touches -- no diagnostic location matches any of them. No lane runs that command: both ep-cuda clippy steps omit `--all-targets` and `workspace_test_packages.py` deny-lists the crate. A real gap, but a separate one; folding it in would put unrelated files in a lane-restoration PR. Closes #1875 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…f which was reported (#1881) Closes #1875. `CUDA compile (Linux x86_64)` has been red since **#1836 (`6e4b0ebb3`)**; last green was `cb81745b0`. Resch handed this off explicitly rather than guessing at another domain's intended API visibility — that was the right call, because **the reported compile error was masking two further failures on the same lane.** Fixing only the error would have moved it from "red at build" to "red at the honesty script". **Two more arrived on `main` while this PR was open** — #1884 and #1895, both today. Five independent defects from five different PRs, reviewable separately. The lane cannot go green on any proper subset of them. --- ### 1. Unresolved import — the reported error `content_preserving_transition_gpu.rs` imports `transition_granule_range_with_phase8_faults`, gated `#[cfg(any(test, feature = "gpu-tests"))]`. The `test` arm does not cover an integration test: `tests/*.rs` are separate crates linking the library built *without* `cfg(test)`, so in the `without-gpu-tests` configuration the item is genuinely absent. Invisible to any local run passing `--features gpu-tests`. **I rejected both options in the issue, with evidence:** - *Widen the gate* contradicts the function's own doc — "Not reachable from production" — and would ship an API that forces `cuMemUnmap` / `cuMemMap` / `cuMemSetAccess` to fail to every consumer. - *`required-features`* is exactly what the honesty lane exists to forbid: `compare_inventories` emits *"exists only with gpu-tests enabled; CUDA tests must not hide from CPU inventory"*. It trades a compile error for a lane failure and deletes the CPU-side inventory of 14 tests. **Fix:** a two-arm local shim, mirroring `with_faults` in `crates/onnx-runtime-cuda-memory/tests/virtual_memory_gpu.rs`, which already solves this identical problem for `CudaVirtualBacking::with_driver_faults`. The gate stays intact; the target compiles and lists its tests in both configurations. The `unreachable!` arm is genuinely unreachable — after change 2 every test in that file is `#[ignore]`d. > **#1895 reached this file first, and took the option rejected above.** > `1be9f2cc2` applied `#![cfg(feature = "gpu-tests")]` to the whole target. It fixed the compile and **traded it for an inventory failure on the same lane.** From the job log on `main@1be9f2c` (job `97266415620`), not inferred: > ``` > - content_preserving_transition_gpu: Cargo reported no integration tests > + 19 × "test exists only with gpu-tests enabled" > ``` > The target was not repaired, it was **hidden** — the symptom stopped being reported without the property becoming true. Its module doc argued that a compile failure "is not a useful signal about a suite that cannot run on a machine with no GPU anyway", which is the exact proposition `verify_cuda_test_honesty.py` exists to reject. This PR removes that line and records the reasoning in the module doc so the option isn't tried a third time. ### 2. Five non-ignored tests in a `_gpu` target — 4 honesty violations The script requires every test in a `_gpu` target to be **ignored, not passed** on a CPU-only runner. That rule is how the suite is stopped from reporting green for a GPU it never touched. **Three were genuinely CPU-only** predicate tests over `verify_safe_point` — pure logic over `ResizeSafePoint`, no driver. **Moved** to a new non-`_gpu` target, not `#[ignore]`d: silencing them would green the lane *by deleting coverage*, and target naming is the escape hatch the script's own comment names, so that "a genuinely CPU-only target is not policed as a device test merely because it lives in a CUDA crate". > **Correction after review — moving them was not sufficient, and my first commit shipped the same defect this PR is about.** > A non-`_gpu` target is skipped by the honesty script, and `.github/scripts/workspace_test_packages.py:26` deny-lists this crate from every offline lane. So the new target was **compiled and never executed** — the moved tests, and the mutation guard I added specifically to catch a bad `verify_safe_point`, would have run on **zero** CI runs. My claim that moving "keeps them executing on every run" was exactly backwards: before the move the honesty script *did* execute them (that was the violation); after it, nothing did. > This is the same failure mode as the defect being fixed — *a test that isn't on the leg you're citing* — and it would have produced a green check proving nothing. Fixed by adding an explicit `--test content_preserving_transition` step to the CUDA lane, mirroring the two CPU-only targets in this crate that already have one for precisely this reason (`ci.yml:989-998`). Verified by running the added command verbatim: 4 passed. **Two asserted nothing** and are removed: - `fault_injection_safe_point_recheck_rejects` — comments plus a `println!`; self-describes as "this test documents the contract". - `zero_len_is_committed_noop` — `println!`s that the behaviour is "verified by implementation". Both passed unconditionally. This touches another author's tests, so I flag it explicitly rather than burying it in the diff. Deleting a test that asserts nothing loses no coverage — but to be precise about what is and isn't covered: the first *does* have real sibling coverage in `fault_injection_recheck_safe_point_rejected_gpu`, and **the second does not** — the `len == 0` early return is untested anywhere. That gap is pre-existing, and a genuine test needs a device (the early return still takes `&CudaRuntime`/`&mut CudaReservation`), so I have left it for a GPU-capable change rather than claim coverage that doesn't exist. **Added `safe_point_accepts_a_clean_state`.** The three moved tests only ever assert `is_err()`, so they cannot distinguish "rejects the unsafe field" from "rejects everything". Empirically confirmed — with `verify_safe_point` mutated to reject unconditionally: ``` test result: FAILED. 3 passed; 1 failed failures: safe_point_accepts_a_clean_state ``` All three inherited tests survive the mutant. Only the new one kills it. ### 3. The manifest guard checked a proxy, not the property it claimed ```python if "gpu-tests = []" not in manifest: errors.append(f"crates/{crate.name}/Cargo.toml must define a gpu-tests feature") ``` A literal substring match. **#1860 turned this red** by changing the value to `["onnx-runtime-cuda-memory/gpu-tests"]` — a legitimate and necessary forwarding. The guard's message claims to check that the feature *is defined*; it actually checked that it was defined *with one particular empty value*. Now matches the feature **key** whatever its value, extracted into `declares_gpu_tests_feature()` so it is covered by the script's own `--self-test` fixtures — which it previously was not. Fixtures include the forwarding form, a multi-line list, a commented-out line, and a `features = ["gpu-tests"]` dependency mention that must *not* count. > Note for @resch: #1860 broke this lane as well as the shape-inference pin you fixed in #1870 — same commit, two unrelated pins that both encoded a value rather than the property. ### 4. A GPU test in a `_gpu` target with no `#[ignore]` — #1895 `causal_conv_with_state_gpu::the_standard_domain_spelling_reaches_the_same_kernel` calls `require_cuda()` exactly like its two siblings in the same file, but is missing their `#[cfg_attr(not(feature = "gpu-tests"), ignore = …)]`. So it ran on a CPU runner and failed: ``` - causal_conv_with_state_gpu: 1 tests failed while checking ignored status - causal_conv_with_state_gpu: Cargo inventory has 3 tests but libtest reported 2 ignored ``` **Fix:** add the sibling attribute verbatim. No design choice involved — a four-line omission. ### 5. A CPU-only test inside a `_gpu` target — #1884 `expert_route_telemetry_probe_gpu::cpu_oracle_and_validator_self_consistent` passes in **both** configurations, which the script reports as `executed without gpu-tests`. Its doc comment states the intent plainly: > *CPU-only: the oracle and the boundary validator are self-consistent. Runs without a GPU so the reference cannot silently rot.* That intent is correct and worth keeping. **`#[ignore]` would clear the checker by destroying exactly the property the test was written for** — an ignored test runs nowhere, so the reference could then rot precisely as its author feared, with a green check over it. Same reasoning as defect 2, so the same route. **Fix:** the pure-CPU oracle — `cpu_bitmap`, `cpu_dedup`, `consume_and_validate`, `synth_routes`, the `Decision` enum and the header indices — moves **verbatim** to a shared `tests/expert_route_oracle/mod.rs`, and the test moves to a new `expert_route_telemetry_probe` target that declares it. `tests/expert_route_oracle/mod.rs` is a module, not a target: Cargo auto-discovers `tests/*.rs` and `tests/*/main.rs` only. Both targets declare it, so there is **one** copy of the oracle and the seven GPU tests keep diffing against the same code the CPU test checks — the alternative, duplicating it, would let the two copies drift and quietly defeat the point. The new target gets its own `ci.yml` step, for the reason in defect 2's correction. **Verified live in its new home by mutation, not by its own green.** Perturbing the shift in `cpu_bitmap`: ``` test result: FAILED. 0 passed; 1 failed failures: cpu_oracle_and_validator_self_consistent ``` and `ok` again on revert. "It compiles in the new file" is not evidence that it *executes* there — which is the defect the review caught in defect 2, so this time it is checked rather than asserted. > Defects 4 and 5 are in other people's very recent files. No open PR touches either (checked across all open PRs), so there is no concurrent-edit hazard, and neither fix requires domain knowledge of the CUDA kernels or the telemetry design — only of where a test must live to be run honestly. --- ## Validation Every fix independently falsified by re-introducing it — each re-reds with a **distinct** error: | reverted | resulting failure | |---|---| | the shim | `error[E0432]: unresolved import ...with_phase8_faults` | | one un-ignored test | `1 tests executed without gpu-tests; CUDA tests must be ignored, not pass` + 3 inventory errors | | the manifest check | `must define a gpu-tests feature` | | the `cfg_attr` (4) | `1 tests failed while checking ignored status` | | the relocation (5) | `1 tests executed without gpu-tests; CUDA tests must be ignored, not pass` | ``` CUDA test honesty check passed: 530 tests/79 targets without gpu-tests (530 ignored), 530 tests/79 targets with gpu-tests (474 fail-loud, 56 ignored, 0 passed on this no-CUDA host) ``` - `cargo clippy -p onnx-runtime-ep-cuda --features cuda -- -D warnings` — clean (the lane's exact command, and both of the two places `ci.yml` invokes it) - `cargo test -p onnx-runtime-ep-cuda --features cuda --lib` — **538 passed, 0 failed** - both explicit CI test steps, run verbatim — `content_preserving_transition` 4 passed, `expert_route_telemetry_probe` 1 passed - `cargo fmt --all --check` — clean - Rebased onto `496597764`; honesty script re-run green on the exact tree being shipped (not on an earlier one — main moved four times during this work). - Independent Opus review — found the never-executed-target defect above; all other categories (shim signature/arg-order vs the real function across all 11 params and 7 call sites, error-string assertion, regex, target classification, and every honesty-script rule) checked and ruled out. **One observation I am not fixing here.** The lane's clippy step omits `--all-targets`, so this crate's test code is linted by **no** lane — `workspace_test_packages.py` also deny-lists it. Running it now yields 23 errors and dies on `index_share_gpu` and `qmoe_zero_copy_cold_expert_spike_gpu` before reaching the rest, all pre-existing and none in a file this PR touches (verified: no diagnostic location matches any of them). Unrelated and out of scope — folding it in would put a pile of unrelated files into a lane-restoration PR — but it is the same absent-coverage shape that let defects 2, 4 and 5 sit un-noticed. ### Credit @Gaff established that this is the **only** instance of the class in the workspace — 24 `cfg(any(test, feature = …))` sites, the other 23 either unreferenced from `tests/` or referenced only from gated positions — and stated the root cause more crisply than I had: an item so gated is broken **only when a `tests/` file names it from a position not itself gated**, in practice a top-level `use`, because a `use` resolves unconditionally regardless of how the functions below it are gated. A gated call site is safe. He also retracted his own grep-based falsifier for the class on #1817 after running the negative control on it. One correction to his note, since it changes what the precedent demonstrates: `onnx-runtime-cuda-memory` does **not** gate the consumers. `vmm_release_quarantine_gpu.rs:59-67` is a **two-arm shim** — a `#[cfg(feature = "gpu-tests")]` helper *and* a `#[cfg(not(feature = "gpu-tests"))]` `unreachable!` twin — with its callers ungated. That distinction is load-bearing here: genuinely gating the consumers would delete the tests from the base inventory and red this lane under `compare_inventories`. So the crate is still the right worked example; it is an example of the shim, which is what this PR adopts (making it the third in the repo, after `virtual_memory_gpu.rs::with_faults` and `vmm_release_quarantine_gpu.rs::install_faults`). 🤖 Generated with [Copilot CLI](https://githubnext.com/projects/copilot-cli) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… test it means (#1911) `CUDA compile (Linux x86_64)` is **red on `main` again**, one commit after #1881 restored it. ``` CUDA test honesty check failed: - activations_gpu: 1 tests failed while checking ignored status - activations_gpu: Cargo inventory has 4 tests but libtest reported 3 ignored ``` #1905 added `activations_gpu::mish_matches_cpu_including_the_saturating_tail` without the `#[cfg_attr(not(feature = "gpu-tests"), ignore = …)]` that its **three siblings in the same file** carry, so it runs and fails on a CPU runner. Fix is the sibling attribute, verbatim. No criticism of @-the-author intended, and I want to be explicit about why. **This is the fourth instance of the same omission in three days** — #1884, #1895, and now #1905. That is a property of the *signal*, not of the authors: while the lane is red for an unrelated reason, adding a `_gpu` test without an ignore is free, because the only check that would object is a job nobody can distinguish from already-broken. #1881 restored the lane; this keeps it restored. --- ## The second half: the checker didn't say which test The message above names a **count**, in a four-test target, and the run that produced it was a CI job on a merge commit. Finding out which test needed a rebase and a grep. A check whose entire job is to notice that a `_gpu` test was added without an ignore should **say which one**. `run_libtest` now parses libtest's per-test outcome lines, so the same failure reads: ``` activations_gpu: 1 tests failed while checking ignored status (mish_matches_cpu_including_the_saturating_tail) ``` **Verified on the real path, not only in fixtures.** With the `cfg_attr` reverted, the full script emits exactly that line; restored, it exits 0. Fixtures alone would only have proved the formatter works. Naming also applies to `executed without gpu-tests` and `passed with gpu-tests on a no-CUDA host`, which have the same problem — those are the two messages that fired for #1884 and #1895. `IgnoredResult` and `ActiveResult` now share a `LibtestResult` base for it. It **degrades to the bare count** if the per-test lines can't be parsed, rather than rendering an empty `()`. The count is still true, so a future change in libtest's output format must not turn a real failure into a confusing one. Four new `--self-test` fixtures cover the naming, the inventory-mismatch message, the empty-name degradation, and the outcome regex itself — because a guard whose own correctness is unverified is precisely the defect this checker exists to catch, and I'd rather not add one to it while fixing it. ## Validation ``` CUDA test honesty check passed: 531 tests/79 targets without gpu-tests (531 ignored), 531 tests/79 targets with gpu-tests (475 fail-loud, 56 ignored, 0 passed on this no-CUDA host) ``` - run on this branch, based on current `main` (`7e274a4e2`) — not on an older tree - `--self-test` passes - `cargo fmt --all --check` clean - mutation-effective both ways: revert the `cfg_attr` → the named failure above; revert the naming → the bare count Follow-up to #1881. Refs #1875. 🤖 Generated with [Copilot CLI](https://githubnext.com/projects/copilot-cli) Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ndary window (#1810 Slice 7A) Add default-disabled, producer-only expert-route telemetry (a device-side bitmap + 6-word header [epoch,request,device,overflow,poison,count]) fused into the existing QMoE and BlockQuantizedMoE route kernels as inert observability. Off/on route outputs are byte-identical; PRODUCE performs zero host readback/sync and zero VMM/map calls in steady-state replay. Cycle-22 revision (Roy NO-GO on the rejected HEAD 8ec9d82): the prior design launched a separate reset/epoch kernel on every execute/replay, which contradicts the approved Slice-6 *coarse-boundary window* contract (design EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md §2.3/§3: the epoch is bumped and the record consumed at a safe boundary, NOT per replay). This revision makes the fused route kernels *accumulate only* and moves reset/epoch-advance to an explicit, host-ordered coarse boundary. Invariant / API: - `arm(request,device,experts)` validates identity, allocates the stable-VA record through the existing runtime allocator, and opens window 1 (stamps identity, epoch=1, zeroes bitmap/counters). No kernel is compiled or launched. - On execute/replay the fused `route_telemetry_mark_row` inside `qmoe_route`/`bqmoe_route` ONLY accumulates into the stable record: atomicOr the routed bit into the bitmap, atomicAdd the in-range count, atomicOr the sticky poison bit on an out-of-range id, atomicOr the sticky overflow bit on count saturation. The epoch is fixed for the whole window; every eager call and captured replay in the window accumulates the routed-expert union + count. There is no reset/epoch kernel in the captured graph. - `reset_route_telemetry_boundary()` is the ONLY place the window advances: it is rejected while the EP stream is capturing/replaying (fail closed), else it drains prior stream work via the existing `drain_for_unmap` authority, bumps the host-side epoch, and re-stamps identity + zeroes the record so the next window starts empty with no stale carryover. It allocates nothing, moves no pointer, and touches no PMM/VMM/cache/global coordinator. - Overflow saturates and fail-closes into the sticky overflow bit (never wraps into a smaller "success" value); out-of-range routes poison. Snapshot/validate run on the host at a boundary against an already-copied record. - Multi-request/multi-device/kernel instances are isolated (distinct records); the bitmap VA is stable across windows and captures. Epoch is a host counter, so the record footprint drops the former 4-byte device epoch buffer. Scope (unchanged from Slice 7A): crate-internal / test-only / default-OFF. No lifecycle/policy/consume wiring. Ordinary inference is byte-identical and has zero overhead when disarmed. Tests (rewritten to the new semantics, same commit): - off/on output byte-identity; - eager calls accumulate the CPU-oracle union + count within a window at a fixed epoch, then a boundary reset opens an empty epoch-2 window; - >=3 graph replays with different routes accumulate the union at a fixed epoch and a stable VA (no per-replay reset); - boundary reset increments the epoch and starts an empty window; - boundary reset rejected during an active capture (epoch unchanged); - capacity/device mismatch stays inert and never fails inference; - multi-instance request/device isolation, footprint, teardown/accounting; - QMoE and BQMoE route match the oracle for M in {1,2,4,8}. Host unit tests cover poison/overflow fail-closed on the validator; the #1884 probe harness continues to cover device-level poison/overflow. Measurement (G1 gate, idle pinned A100, DeepSeek-V2-Lite decode layer, CUDA graph replay, BATCH=512, RUNS=5, ~8s ramp + first-shape recheck, n=3): - GPU route+layer overhead 0.948-1.054us (0.82-0.91% of the 116us layer); - host-enqueue delta ~0us; epoch fixed at 1 across the whole window (count accumulated to 30726); output byte-identical; clock drift < 0.5%. - GATE (<=2us AND <=2%): GO. Telemetry remains default-OFF regardless. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ndary window (#1810 Slice 7A) Add default-disabled, producer-only expert-route telemetry (a device-side bitmap + 6-word header [epoch,request,device,overflow,poison,count]) fused into the existing QMoE and BlockQuantizedMoE route kernels as inert observability. Off/on route outputs are byte-identical; PRODUCE performs zero host readback/sync and zero VMM/map calls in steady-state replay. Cycle-22 revision (Roy NO-GO on the rejected HEAD 8ec9d82): the prior design launched a separate reset/epoch kernel on every execute/replay, which contradicts the approved Slice-6 *coarse-boundary window* contract (design EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md §2.3/§3: the epoch is bumped and the record consumed at a safe boundary, NOT per replay). This revision makes the fused route kernels *accumulate only* and moves reset/epoch-advance to an explicit, host-ordered coarse boundary. Invariant / API: - `arm(request,device,experts)` validates identity, allocates the stable-VA record through the existing runtime allocator, and opens window 1 (stamps identity, epoch=1, zeroes bitmap/counters). No kernel is compiled or launched. - On execute/replay the fused `route_telemetry_mark_row` inside `qmoe_route`/`bqmoe_route` ONLY accumulates into the stable record: atomicOr the routed bit into the bitmap, atomicAdd the in-range count, atomicOr the sticky poison bit on an out-of-range id, atomicOr the sticky overflow bit on count saturation. The epoch is fixed for the whole window; every eager call and captured replay in the window accumulates the routed-expert union + count. There is no reset/epoch kernel in the captured graph. - `reset_route_telemetry_boundary()` is the ONLY place the window advances: it is rejected while the EP stream is capturing/replaying (fail closed), else it drains prior stream work via the existing `drain_for_unmap` authority, bumps the host-side epoch, and re-stamps identity + zeroes the record so the next window starts empty with no stale carryover. It allocates nothing, moves no pointer, and touches no PMM/VMM/cache/global coordinator. - Overflow saturates and fail-closes into the sticky overflow bit (never wraps into a smaller "success" value); out-of-range routes poison. Snapshot/validate run on the host at a boundary against an already-copied record. - Multi-request/multi-device/kernel instances are isolated (distinct records); the bitmap VA is stable across windows and captures. Epoch is a host counter, so the record footprint drops the former 4-byte device epoch buffer. Scope (unchanged from Slice 7A): crate-internal / test-only / default-OFF. No lifecycle/policy/consume wiring. Ordinary inference is byte-identical and has zero overhead when disarmed. Tests (rewritten to the new semantics, same commit): - off/on output byte-identity; - eager calls accumulate the CPU-oracle union + count within a window at a fixed epoch, then a boundary reset opens an empty epoch-2 window; - >=3 graph replays with different routes accumulate the union at a fixed epoch and a stable VA (no per-replay reset); - boundary reset increments the epoch and starts an empty window; - boundary reset rejected during an active capture (epoch unchanged); - capacity/device mismatch stays inert and never fails inference; - multi-instance request/device isolation, footprint, teardown/accounting; - QMoE and BQMoE route match the oracle for M in {1,2,4,8}. Host unit tests cover poison/overflow fail-closed on the validator; the #1884 probe harness continues to cover device-level poison/overflow. Measurement (G1 gate, idle pinned A100, DeepSeek-V2-Lite decode layer, CUDA graph replay, BATCH=512, RUNS=5, ~8s ramp + first-shape recheck, n=3): - GPU route+layer overhead 0.948-1.054us (0.82-0.91% of the 116us layer); - host-enqueue delta ~0us; epoch fixed at 1 across the whole window (count accumulated to 30726); output byte-identical; clock drift < 0.5%. - GATE (<=2us AND <=2%): GO. Telemetry remains default-OFF regardless. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ndary window (#1810 Slice 7A) (#1922) ## Slice 7A — inert, default-OFF, producer-only expert-route telemetry (QMoE/BQMoE) Closes #1810 (Slice 7A). **Draft. Do not merge.** Crate-internal / test-only / default-OFF. > **Cycle-22 revision by Sebastian (Performance Engineer), revision owner.** > The original author is under reviewer-protocol lockout and did not advise, pair, or contribute to this revision. Reworked independently from the PR diff, the approved Slice-6 design doc, and Roy's Cycle-22 review. ### What changed vs the rejected HEAD (`8ec9d8230`) Roy's Cycle-22 NO-GO was correct and is **addressed at the design level**: the rejected design launched a **separate reset/epoch kernel on every execute/replay**, which contradicts the approved Slice-6 **coarse-boundary window** contract (`docs/memory/EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md` §2.3/§3 — the epoch is bumped and the record consumed at a *coarse safe boundary*, **not per replay**). That per-call reset is why the fused path measured **2.526 us / 2.18 % -> NO-GO**. This revision makes the fused route kernels **accumulate-only** and moves reset/epoch-advance to an explicit, host-ordered **coarse boundary**. The earlier NO-GO was scoped to the wrong (per-call reset) design point; the boundary-window producer now measures **GO** (see below), independently reproduced. ### Invariant / API (producer-only boundary-window contract) - **`arm(request, device, experts)`** - validates identity, allocates the stable-VA record via the existing runtime allocator, opens **window 1** (stamps identity, `epoch = 1`, zeroes bitmap/counters). **No kernel is compiled or launched.** - **execute / replay** - the fused `route_telemetry_mark_row` inside `qmoe_route` / `bqmoe_route` **only accumulates** into the stable record: `atomicOr` the routed bit into the bitmap, `atomicAdd` the in-range count, `atomicOr` the sticky **poison** bit on an out-of-range id, `atomicOr` the sticky **overflow** bit on count saturation. The **epoch is fixed for the whole window**; every eager call and captured replay in the window accumulates the routed-expert **union + count**. **No reset/epoch kernel in the captured graph; no host sync/alloc/drain/VMM on this path.** - **`reset_route_telemetry_boundary()`** - the **only** place the window advances. **Rejected while the EP stream is capturing/replaying** (fail closed, checked *before* any drain); otherwise drains prior stream work via the existing `drain_for_unmap` authority, bumps the **host-side** epoch, and re-stamps identity + zeroes the record so the next window starts empty with **no stale carryover**. Allocates nothing, moves no pointer, touches no PMM/VMM/cache/global coordinator. - **Fail-closed** - overflow **saturates** into the sticky overflow bit (never wraps into a smaller "success"); out-of-range routes **poison** without touching the bitmap; snapshot/validate run on the host at a boundary against an already-copied record. - **Isolation / stability** - multi-request / multi-device / multiple kernel instances use distinct records; the bitmap **VA is stable** across windows and captures. Epoch is a host counter, so the record footprint **drops the former 4-byte device epoch buffer** (`footprint = 4*ceil(experts/32) + 24`). - **Reuses only existing authorities** - `alloc_raw`/`free_raw`, `htod`/`dtoh`, `is_capturing`, `drain_for_unmap`, `ordinal`. No new coordinator/allocator/cache/PMM/VMM and **no host sync in steady-state replay**. Scope unchanged: **default-OFF, crate-internal, test-only**. Ordinary inference is **byte-identical** and has **zero overhead when disarmed**. No lifecycle/policy/consume wiring in this slice. ### Tests (rewritten to the new semantics, same commit) QMoE (`tests/qmoe_gpu.rs`) and BQMoE (`tests/block_quantized_moe_gpu.rs`) `mod route_telemetry`: - off/on output **byte-identity**; - eager calls **accumulate the CPU-oracle union + count** within a window at a **fixed epoch**, then a boundary reset opens an empty **epoch-2** window; - **>=3 graph replays** with different routes accumulate the union at a **fixed epoch** and a **stable VA** (QMoE - proves no per-replay reset); - boundary reset **increments the epoch and starts an empty window**; - boundary reset **rejected during an active capture** (epoch unchanged); - capacity / device mismatch stays **inert** and never fails inference; - multi-instance **request/device isolation**, footprint, teardown/accounting; - **QMoE and BQMoE routes for M in {1, 2, 4, 8}** match the oracle. Host unit tests cover **poison/overflow fail-closed** on the validator; the #1884 probe harness continues to cover device-level poison/overflow. **Results** (idle A100, GPU 5, `--test-threads=1`): QMoE `route_telemetry` 10 passed; BQMoE 8 passed; telemetry host unit tests 7 passed; #1884 probe functional regressions 6 passed; **full QMoE suite 56 passed**, **full BQMoE suite 13 passed** (no regressions to #1788/#1800/#1884/#1854). ### Measurement - G1 gate (`microbench_fused_route_telemetry_g1_gate`) Idle **pinned A100**, serialized off/on back-to-back, whole QMoE layer **CUDA-graph captured** and measured via `replay_graph` (host-enqueue gaps excluded), **BATCH = 512, RUNS = 5**, ~8 s clock ramp + first-shape recheck, GPU-event and host-enqueue reported separately. Gate denominator is a **realistic** DeepSeek-V2-Lite decode layer (64 experts, top_k 6, hidden 2048); tiny synthetic is informational only. **n = 3 independent runs:** | Run | route+layer overhead (GPU) | % of layer | host-enqueue delta | epoch | count | output | |-----|---------------------------|-----------|----------------|-------|-------|--------| | 1 | 0.976 us | 0.84 % | -0.016 us | fixed = 1 | 30726 | identical | | 2 | 0.948 us | 0.82 % | -0.006 us | fixed = 1 | 30726 | identical | | 3 | 1.054 us | 0.91 % | -0.008 us | fixed = 1 | 30726 | identical | **Range: 0.948-1.054 us, 0.82-0.91 %. Clock drift < 0.5 %.** The epoch stays **fixed at 1** across the whole BATCH x RUNS replay window (count accumulates to 30726) - a moving epoch would have proven a forbidden per-replay reset survived. **G1 GATE (<= 2 us AND <= 2 %): GO.** Independently reproduced vs Roy's boundary-only experiment (0.884 us / 0.76 %); same order, same verdict - not a reused number. Telemetry remains **default-OFF** regardless of the GO. ### Review Independent review (excluding the locked-out author and the revision owner): **APPROVE** - no blocking or substantive findings; confirmed accumulate-only execute path, reject-before-drain boundary reset, fail-closed overflow/poison, stable VA, isolation, footprint, and that the tests assert the *new* semantics (they would fail under the rejected per-replay-reset design). One MINOR note: the device-side poison branch of the fused mark is defensive and unreachable-by-construction from the production route kernel (Roy's prior approved framing) - covered by the host validator unit tests and the #1884 probe harness. **Roy's explicit final re-review is requested** before this leaves draft. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… coarse residency (#1810 Slice 7B) (#1971) ## #1810 Slice 7B — boundary-time route-telemetry **consumer** **Status: DRAFT — awaiting independent review. Do not merge.** Closes the loop opened by the merged Slice-6/7A **producer** (expert-route telemetry, PR #1922 `e1ec495ee`) and the merged Slice-4/5 **coarse-boundary residency lifecycle** (PR #1854). This is the smallest production seam the Slice-6 design §8 specifies: > expose a boundary-time consumer that produces a per-expert desired-set, and feed that set to the **existing** Slice 4/5 coarse-boundary plan application (`coarse_residency.rs`) as its policy input — with **no** new allocator, **no** id→slot rewrite, and every mapping change still owned by PMM/VMM. ### What it does New module `crates/onnx-runtime-ep-cuda/src/route_residency.rs`: `consume_route_window_at_boundary(...)`: 1. **Gate** — default-off via the existing `COARSE_RESIDENCY_ENABLE_ENV` (`coarse_residency_profile_enabled()`). When off (shipped default) it returns `Disabled` *before* reading the snapshot or touching any allocator. 2. **Safe boundary** — re-reads the existing `CudaWeightResidency::resize_safe_point` and fails closed with `RejectedNotSafeBoundary { reason }` if a graph is capturing/replaying, an admission is in flight, a deferred release has not settled, execution is multi-device, or a routed-residency guard is live. Coarse-boundary only — never a per-token remap. 3. **Validate** — runs the producer's own `consume_and_validate` on the already-completed window snapshot; fail-closed (→ `WholeBank { reason }`) on poison / overflow / stale epoch / foreign request / foreign device, or when no in-range experts were recorded. 4. **Plan** — turns the routed-expert union into a desired **hot set** and asks the already-validated `StaticProfileResidencyPolicy` to shape a `ResidencyPlan` (the design's "record → desired set" `RouteObserverPolicy` role — **reused, not duplicated**, so exactly one validated policy emits `PerExpertCandidate`). 5. **Apply** — hands the plan to the existing `CudaWeightResidency::apply_coarse_residency_plan`, which remains the **sole** authority that maps / unmaps / accounts / quarantines / rolls back through PMM/VMM. `RouteWindowConsumeOutcome` carries the exact reason on every non-applied path — **no silent fallback**. ### Invariants held by construction - **Allocates nothing**, opens no stream, owns no VA, copies no device bytes — pure host glue between two existing authorities. - **No remap during capture/replay**; coarse-boundary only, never per-token. - **No new host sync** in steady state (only the producer's already-taken snapshot and the existing transition primitive's drain). - **Default-off & byte-identical** when disarmed (two independent default-off switches: telemetry disarmed *and* this gate off). - **Same-device fail-closed** and **PMM/VMM remains the sole mapping/accounting/quarantine/rollback authority**. ### Tests — `tests/route_residency_consume_gpu.rs` (6 GPU tests, all passing on an idle A100) - `disabled_gate_is_structural_no_op` — off path is a structural no-op (byte-identical). - `route_window_hot_set_transitions_cold_experts` — routed hot-set stays resident, cold set tiers to host, bytes identical. - `expert_group_transitions_atomically_from_window` — atomic expert-group transition driven from a window. - `active_capture_and_multi_device_reject_consume` — active-capture and multi-device boundaries reject the consume (fail-closed). - `foreign_identity_and_defective_windows_fail_closed` — foreign request/device and poison/overflow/stale windows fail closed to whole-bank. - `injected_fault_rolls_back_consumer_transition` — injected driver fault (`fail_nth(Unmap, 3)` over an isolated + merged cold range) rolls back range-precisely and quarantines; `rollback_count == 1`, `values_touched == 0`, `committed_values` empty, content bit-identical (mirrors the proven `coarse_residency_plan_gpu.rs` rollback fixture). Run: ``` env -u ONNX_GENAI_WEIGHT_OFFLOAD_COARSE_RESIDENCY_ENABLE CUDA_VISIBLE_DEVICES=<idle> ONNX_GENAI_CUDA_DEVICE=0 \ cargo test -p onnx-runtime-ep-cuda --features cuda,cuda-13000,gpu-tests --release \ --test route_residency_consume_gpu -- --ignored --test-threads=1 ``` ### Honest scope Like `coarse_residency::apply_residency_plan_at_boundary` when it shipped (Slice 5), this consumer has **no live decode-loop call site yet** — wiring it into a running session's request boundary is the next slice. It ships here as the production seam: reachable, default-off, and proven by the GPU tests. `cargo fmt` applied; `cargo clippy` on the crate lib is clean for the new file (pre-existing unrelated warnings/errors in `optimizer.rs` / `standard_attention.rs` are out of scope). ### Constraints honored Did not touch PagedAttention or IQ1 fusion files. Branch `squad/1810-slice7b-telemetry-residency-consume` based on latest `main` (`011fbb284`, includes #1922 `e1ec495ee`). #1788/#1800/#1884/#1854 lifecycle regressions preserved (reused, not forked). Refs #1810. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
What & why
Part of #1810 (composable sub-weight VMM for MoE experts). Slice 6 produces the concrete design for adapting FreeToken's device-side expert-route observation/admission to onnx-genai's QMoE/BlockQuantizedMoE, plus an inert (test-only) proof harness. It stops at a bounded design + proof harness — no production residency/lifecycle wiring.
Working as Deckard (DevOps/Performance).
Independent of PR #1854 (Slice 5)
This PR touches none of #1854's files (
vmm_allocator.rs,ep-api/{lib,weight}.rs,ep-cuda/{coarse_residency,lib,weight_paging}.rs,coarse_residency_plan{,_gpu}.rs). The harness is a new integration test — its own compilation unit, so it needs nosrc/edit and shares no code with #1854.Contents (2 new files + 1 decision note, all additive)
docs/memory/EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md— 8-section design:qmoe_route/bqmoe_routewriteint selected_experts[rows,top_k]into a persistentScratchPoolslot; consumed on-device) and why the host is blind before launch (routing is computed inside the fused op from on-device activations; the only readback is a debugdtoh).ceil(E/32)u32,atomicOr) + optional bounded dedup queue;u32[6]header{epoch, request, device, overflow, poison, count}. Zero steady-state host sync; graph capture/replay safe; preserves the contiguous expert-bank pointer ABI.cuMemMap/cuMemSetAccessduring capture/replay.decode_freqthat's "only accurate with graphs disabled").crates/onnx-runtime-ep-cuda/tests/expert_route_telemetry_probe_gpu.rs— CPU oracle + CUDA bitmap/dedup/overflow/poison/epoch/capture-replay/isolation tests + a microbench that reports GPU-event vs host-enqueue time separately. Uses only the publicCudaRuntimeAPI..squad/decisions/inbox/deckard-1810-slice6-route-telemetry.md— decision note for the Scribe.Validation (idle A100-SXM4-80GB, GPU 5, CUDA 13.0, driver 580.105.08)
All 8 tests pass (
--test-threads=1, GPU verified idle before run):WholeBank); poison fails closed; foreign request/device fail closed; owner accepts;qmoe_routefor ~0 extra launches).No speedup is claimed (measurement-discipline — wall clock cannot resolve this).
Independent review
The
code-reviewagent (not Deckard) reviewed the harness and found no high-confidence bugs — confirmed race-correct dedup atomics, a genuine CPU/GPU cross-check (HashSet vs seen-bitmap), a fail-closed validator, a correct capture region (no sync inside), correct Trap-4 microbench separation, alignment-safe byte casts, and no use-after-free (DeviceBufferhas noDrop; buffers outlive device use).How to run
Not in scope (next slice, GO-gated)
Slice 7 wires telemetry as new
ScratchPoolslots ofQMoEKernel/BlockQuantizedMoEKernel, writes it from the route kernel, and feeds a boundary-time desired-set to the existing coarse-boundary plan — with no new allocator, no id→slot rewrite, and every mapping change still owned by PMM/VMM. Gates G5 (queue sizing) and G6 (real-model byte-hit headroom) require the real routing corpus and remain open; whole-bank stays the safe default until G6 passes.