Skip to content

Slice 6: device-side expert-route telemetry design + inert proof harness (#1810) - #1884

Merged
justinchuby merged 1 commit into
mainfrom
squad/1810-slice6-route-observation-harness
Aug 23, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/1810-slice6-route-observation-harness

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What & why

Part of #1810 (composable sub-weight VMM for MoE experts). Slice 6 produces the concrete design for adapting FreeToken's device-side expert-route observation/admission to onnx-genai's QMoE/BlockQuantizedMoE, plus an inert (test-only) proof harness. It stops at a bounded design + proof harness — no production residency/lifecycle wiring.

Working as Deckard (DevOps/Performance).

Independent of PR #1854 (Slice 5)

This PR touches none of #1854's files (vmm_allocator.rs, ep-api/{lib,weight}.rs, ep-cuda/{coarse_residency,lib,weight_paging}.rs, coarse_residency_plan{,_gpu}.rs). The harness is a new integration test — its own compilation unit, so it needs no src/ edit and shares no code with #1854.

Contents (2 new files + 1 decision note, all additive)

  • docs/memory/EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md — 8-section design:
    1. Where expert IDs exist on device today (qmoe_route/bqmoe_route write int selected_experts[rows,top_k] into a persistent ScratchPool slot; consumed on-device) and why the host is blind before launch (routing is computed inside the fused op from on-device activations; the only readback is a debug dtoh).
    2. Telemetry contract — fixed-capacity, GPU-resident: route bitmap (ceil(E/32) u32, atomicOr) + optional bounded dedup queue; u32[6] header {epoch, request, device, overflow, poison, count}. Zero steady-state host sync; graph capture/replay safe; preserves the contiguous expert-bank pointer ABI.
    3. State machine at a coarse safe boundary — produce during launch, consume only after existing stream-completion authority; overflow/stale/foreign-request/foreign-device fail closed to the whole-bank proof; forbids cuMemMap/cuMemSetAccess during capture/replay.
    4. One-authority rule — policy only chooses a desired set; PMM/VMM alone owns mapping/accounting/quarantine/rollback. No second LRU/cache/allocator, no id→slot rewrite.
    5. FreeToken comparison — copied concepts (device-accumulated stats, graph-safe, no per-step host scatter) vs rejected assumptions (fixed slot table / id→slot remap, per-token remap, host decode_freq that's "only accurate with graphs disabled").
    6. Cost model + falsifiable GO/NO-GO gates using authoritative shapes (tiny/full DeepSeek-V4, GLM-5.2). No speedup quoted.
  • crates/onnx-runtime-ep-cuda/tests/expert_route_telemetry_probe_gpu.rs — CPU oracle + CUDA bitmap/dedup/overflow/poison/epoch/capture-replay/isolation tests + a microbench that reports GPU-event vs host-enqueue time separately. Uses only the public CudaRuntime API.
  • .squad/decisions/inbox/deckard-1810-slice6-route-telemetry.md — decision note for the Scribe.

Validation (idle A100-SXM4-80GB, GPU 5, CUDA 13.0, driver 580.105.08)

All 8 tests pass (--test-threads=1, GPU verified idle before run):

  • bitmap == CPU oracle (E=4/64/256, rows 1 & 37); dedup set == oracle (199 distinct, no overflow);
  • overflow fails closed (distinct=226 > cap=8 → WholeBank); poison fails closed; foreign request/device fail closed; owner accepts;
  • capture/replay: 3 replays re-accumulate each replay's real routes, buffer VA stable, epoch 1→2→3, stale (epoch 3 vs boundary 4) fails closed;
  • microbench (Trap-4 separated, ramped): decode telemetry launch GPU 2.48 µs ≈ host-enqueue 2.47 µs (launch-bound — pessimistic separate-launch upper bound; Slice 7 folds atomics into qmoe_route for ~0 extra launches).

No speedup is claimed (measurement-discipline — wall clock cannot resolve this).

Independent review

The code-review agent (not Deckard) reviewed the harness and found no high-confidence bugs — confirmed race-correct dedup atomics, a genuine CPU/GPU cross-check (HashSet vs seen-bitmap), a fail-closed validator, a correct capture region (no sync inside), correct Trap-4 microbench separation, alignment-safe byte casts, and no use-after-free (DeviceBuffer has no Drop; buffers outlive device use).

How to run

CUDA_VISIBLE_DEVICES=<idle> cargo test -p onnx-runtime-ep-cuda \
  --features cuda-13000,gpu-tests --release \
  --test expert_route_telemetry_probe_gpu -- --ignored --nocapture --test-threads=1

Not in scope (next slice, GO-gated)

Slice 7 wires telemetry as new ScratchPool slots of QMoEKernel/BlockQuantizedMoEKernel, writes it from the route kernel, and feeds a boundary-time desired-set to the existing coarse-boundary plan — with no new allocator, no id→slot rewrite, and every mapping change still owned by PMM/VMM. Gates G5 (queue sizing) and G6 (real-model byte-hit headroom) require the real routing corpus and remain open; whole-bank stays the safe default until G6 passes.

…ess (#1810)

Adapt FreeToken's device-side expert-route observation to onnx-genai's
QMoE/BlockQuantizedMoE as an inert, observer-only telemetry contract, without
touching PMM/VMM or CUDA-graph authorities. Design + test-only proof harness;
no production residency/lifecycle wiring.

Independent of PR #1854 (Slice 5): touches none of its files; the harness is a
new integration test (its own compilation unit, no src/ edit).

- docs/memory/EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md: 8-section design — where
  expert IDs live on device and why the host is blind; fixed-capacity GPU-
  resident telemetry contract (route bitmap + bounded dedup queue + u32[6]
  header {epoch,request,device,overflow,poison,count}); coarse-boundary state
  machine (produce at launch, consume after stream completion, fail closed;
  forbids cuMemMap/cuMemSetAccess during capture/replay); one-authority rule
  (PMM/VMM owns mapping; policy only chooses a desired set); FreeToken
  copied-vs-rejected comparison; cost model + falsifiable GO/NO-GO gates.
- crates/onnx-runtime-ep-cuda/tests/expert_route_telemetry_probe_gpu.rs:
  CPU oracle + CUDA bitmap/dedup/overflow/poison/epoch/capture-replay/isolation
  tests + Trap-4-separated microbench. Uses only the public CudaRuntime API.

Validated on idle A100 (GPU 5, CUDA 13.0, driver 580): all 8 tests pass;
capture/replay re-accumulates each replay's real routes (VA stable, epoch 1->3,
stale detected); overflow/poison/foreign-request/foreign-device fail closed.
Independent code-review (not Deckard) found no high-confidence bugs. No speedup
claimed (measurement-discipline).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 72.56%. Comparing base (5827df5) to head (308bed0).
⚠️ Report is 10 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main    #1884   +/-   ##
=======================================
  Coverage   72.56%   72.56%           
=======================================
  Files          12       12           
  Lines        5231     5231           
  Branches     5231     5231           
=======================================
  Hits         3796     3796           
  Misses       1307     1307           
  Partials      128      128           
Flag Coverage Δ
cli-ort-linux 72.51% <ø> (ø)
cli-ort-windows 72.01% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@justinchuby

Copy link
Copy Markdown
Owner Author

Cycle 20 — PR #1884 Independent Review (Slice 6 route-observation harness)

Reviewer: Roy (Lead)
Subject: #1884, commit
308bed085bb9a395cd08ba81fa9020adc227287b
Author: Deckard (Systems/Perf)
Verdict: APPROVE HARNESS FOR MERGE

Note: measurements were run only after Cycle 19's #1854 GPU tests had fully
finished (serialized, not concurrent), per task instruction.

Scope verification (blocking-if-violated — clean)

git diff of this PR's 3 files (.squad/decisions/inbox/deckard-1810-slice6-route-telemetry.md,
crates/onnx-runtime-ep-cuda/tests/expert_route_telemetry_probe_gpu.rs,
docs/memory/EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md) against origin/main
confirms zero overlap with #1854's touched files
(vmm_allocator.rs, ep-api/{lib,weight}.rs,
ep-cuda/{coarse_residency,lib,weight_paging}.rs,
coarse_residency_plan{,_gpu}.rs) and zero src/ changes anywhere. This is
a new integration-test compilation unit only — no mod registration, no
production call sites (grep confirms nothing outside this test file references
its types). Matches Cycle 18's handoff spec exactly (worktree/branch used was
the one I prepared).

Source-citation accuracy (spot-checked, not trusted)

Verified by direct file reads, not by trusting the doc: qmoe.rs:115,133,378,614
(qmoe_route kernel signature, write site, both GEMV consumer sites),
qmoe.rs:1671,1723 (routes = rows*k, scratch slot 0 allocation),
qmoe.rs:~2802-2809 (ScratchPool::ensure returns Err under capturing,
confirmed verbatim), qmoe.rs:1874,1999-2004,2035-2064 (debug route dump: off
by default, !capturing-guarded, blocking dtoh, steady-state path issues no
readback), block_quantized_moe.rs:79,183 (bqmoe_route signature, GEMV
consumer). All citations check out precisely against current source — this
design is grounded in the actual codebase, not invented.

Technical contract verification

  • Fixed buffer/header ABI: u32[HEADER_LEN=6] header
    (epoch,request,device,overflow,poison,count), bitmap u32[ceil(E/32)],
    dedup queue i32[capacity] + seen[ceil(E/32)] — all fixed-capacity,
    allocated once, matches design §2.3 exactly.
  • Epoch/request/device isolation: cuda_identity_isolation_fails_closed
    and cuda_capture_replay_reaccumulates_real_routes both non-vacuously
    exercise this — foreign request (101 vs owner 100) and foreign device (7 vs
    owner 0) each independently assert WholeBank, and the owning
    request+device+epoch combination is asserted to accept (HotSet). Stale
    epoch (record epoch=3, boundary=4) is asserted to fail closed; the same
    record at its own epoch (3) is asserted to accept — a real, not tautological,
    freshness check.
  • Overflow/poison fail-closed: cuda_dedup_overflow_fails_closed
    explicitly asserts its own precondition (distinct.len() > capacity) before
    checking the device sets overflow=1, so the test cannot pass by accident
    under capacity. cuda_poison_out_of_range_fails_closed injects real
    out-of-range ids (num_experts, -1) and both kernels' poison branch
    (if (e<0 || e>=num_experts) atomicOr(&header[4],1u)) is confirmed to set the
    bit and fail the validator closed.
  • Capture/replay address stability: cuda_capture_replay_reaccumulates_real_routes
    captures stamp_and_reset+route_bitmap via raw cuStreamBeginCapture_v2/
    cuStreamEndCapture/cuGraphInstantiateWithFlags, replays 3× with
    different synthetic routes each time (outside the graph, via htod), and
    asserts (a) the bitmap byte-equals the CPU oracle for that replay's own
    routes each time, (b) the buffer's device pointer is unchanged across all 3
    replays, (c) the epoch counter advances exactly once per replay (1→2→3). This
    is a genuine, non-trivial capture-replay proof, not a single-shot capture
    check. No dtoh/synchronize occurs between cuStreamBeginCapture_v2 and
    cuStreamEndCapture — confirmed by reading the capture region, which
    contains only two kernel launches.
  • No hidden host sync in steady state: the kernels themselves contain only
    atomics; the design correctly separates PRODUCE (device-only) from CONSUME
    (one dtoh at a boundary, after existing stream-completion authority) — the
    test's own capture region proves the PRODUCE side needs no host
    intervention.
  • PMM/VMM remains sole authority: the harness never calls
    PhysicalHandlePool/CudaVirtualBacking/any VMM primitive; it allocates
    purely through ExecutionProvider::allocate (a standard device buffer, not a
    raw VMM binding — even more conservative than spike(#1810): composable device+host-NUMA VMM feasibility slice (EXPERIMENTAL) #1813's approach). No
    cuMemMap/cuMemSetAccess anywhere in the file.

CPU oracle / non-vacuity

Every GPU assertion diffs against cpu_bitmap/cpu_dedup, pure-Rust
ground-truth functions with their own dedicated CPU-only test
(cpu_oracle_and_validator_self_consistent, verifies bitmap/dedup mutual
agreement and all 6 validator-reject paths) — not hardcoded expected values.
This is real cross-checking, matching this team's established bar from prior
harness reviews (#1813/#1823/#1829).

Measurement discipline (independently re-run, not trusted)

Re-ran on idle GPU 5 (A100, confirmed 4 MiB/210 MHz before start via
nvidia-smi) after Cycle 19's tests fully completed:

  • All 8/8 tests pass (1 CPU-only + 7 GPU --ignored), independently
    rebuilt (cargo build -p onnx-runtime-ep-cuda --features cuda,gpu-tests --tests --release).
  • Microbench re-run gave decode GPU/launch median 3.00 µs ≈ host-enqueue
    2.99 µs; prefill-512 GPU 3.06 µs ≈ host-enqueue 3.04 µs — close
    to, not identical to, the PR's reported 2.48/2.47 µs and 2.99/2.61 µs (run-to-
    run clock/driver variance is expected and disclosed as inherent in this
    methodology). The qualitative finding the PR rests its conclusion on —
    GPU-event time ≈ host-enqueue time, i.e. launch-bound, not compute-bound
    — reproduces independently. This directly answers the task's challenge: the
    reported number is not one-launch host timing (measured via a 512-launch
    batch enclosed by CUDA events, Trap 4) and not a cold-clock artifact (an
    explicit ~8s ramp precedes each shape's measurement, and idle is re-verified
    via nvidia-smi before/after each shape).
  • cargo fmt --check -p onnx-runtime-ep-cuda: clean.
  • cargo clippy --features cuda,gpu-tests --tests --release -- (JSON-filtered to this file): zero warnings in expert_route_telemetry_probe_gpu.rs.

GO/NO-GO gate framing (no overclaim)

The design's own G1 threshold (§6.4: "≤2 µs at decode") is nominally exceeded
by both the PR's reported number and my independent re-run (~2.5-3 µs) — but
the design/PR body correctly does not claim G1 passes: §7.1 explicitly
states "G1 is bounded but must be re-measured fused in Slice 7," reasoning
(supported by the GPU≈host-enqueue equality) that the standalone number is
dominated by launch overhead a fused kernel would not pay. This is honest,
falsifiable framing, not a hidden failure. No speedup is claimed anywhere, and
this number is never compared against the incompatible Slice 2-4 remap
quantities (150-430 µs/layer, 6.05 ms/token) — the exact defect class found and
fixed in the #1829 cycle is correctly avoided here from the start.

Non-blocking observations (do not require revision)

  1. Skill citation mislabel: the design doc and PR body cite "Traps 5/6" for
    "ramped and re-checked idle." Only Trap 5
    (.github/skills/cuda-perf-measurement/SKILL.md) covers clock-ramp
    discipline; Trap 6 in that skill is about a fully unrelated topic (a
    production fast-path that never fires). This is a citation nit, not a
    methodology defect — the actual ramp/idle-check behavior in the test is
    correct regardless of the mislabel.
  2. Trap 5's "re-measure the first shape at the end of the sweep" step is not
    implemented
    — the microbench measures decode then prefill-512 once
    each, never re-measuring decode afterward to report drift. Given neither
    number is load-bearing for any gate conclusion in this PR (G1 is explicitly
    deferred, not decided, from this data), this is a minor completeness gap
    worth closing in Slice 7's fused re-measurement, not a blocker here.

Neither observation touches correctness, scope, or any claim this PR actually
relies on.

Conclusion

All items in the review checklist check out: strict scope (zero #1854 overlap,
zero production wiring), accurate source citations, a sound and non-vacuous
fixed-capacity/fail-closed telemetry contract with real capture/replay proof,
disciplined and independently-reproduced measurement, and no overclaim of
production readiness (G5/G6 explicitly open, whole-bank stays default). This
harness durably documents a real, falsifiable design for Slice 7 and is worth
merging as-is.

APPROVE HARNESS FOR MERGE.

@justinchuby
justinchuby merged commit 31d0380 into main Aug 23, 2026
9 of 17 checks passed
@justinchuby
justinchuby deleted the squad/1810-slice6-route-observation-harness branch August 23, 2026 21:27
justinchuby added a commit that referenced this pull request Aug 23, 2026
…eration step (#1902)

## Context

This is a follow-up investigation, not a fix for a currently-broken
state. On fresh `main` (post #1884/#1893), both
`tests/fixtures/tiny-deepseek-v4-qmoe` and
`tests/fixtures/tiny-glm52-full-attention` already load and decode
correctly — the missing-`policies/*.onnx` problem I'd flagged in a prior
cycle was already fixed by **#1883** (re-emitting
`inference_metadata.yaml` with `migrate_model_io --reemit`, dropping ten
hand-authored, never-real auxiliary policy components) and **#1888**
(test-matrix follow-up). That fix is correct, already reviewed, and this
PR does not revisit it.

## What this PR actually fixes

Regeneration **determinism**. Re-running `generate.py` (and its
`--full-attention` variant) against each fixture's pinned Mobius commit
reproduces `model.onnx.textproto`/`tokenizer.json` byte-for-byte, but
**not** `inference_metadata.yaml` — it reproduces the
pre-canonicalization document `write_onnx_genai_config` emits, not the
canonical single-`decoder`-component shape actually committed. This is
true for **all three** tiny fixtures (not just the two #1883 touched),
meaning a naive future regeneration would silently reintroduce the exact
bug #1883 fixed.

Fix: document the required two-step recipe (`generate.py`, then
`migrate_model_io --reemit <output-dir>`) directly in each generator's
docstring. **Docstring-only change** — no functional/behavior change, no
fixture artifacts touched.

## Verification (2026-08-23, idle A100, `CUDA_VISIBLE_DEVICES` pinned,
`hostlock`-gated)

- `generate.py` alone vs committed: textproto/tokenizer byte-identical,
metadata differs — confirmed for all 3 fixtures.
- `generate.py` + `migrate_model_io --reemit` vs committed: metadata
byte-identical — confirmed for all 3 fixtures.
- `cargo test -p onnx-genai-metadata --test
decoder_recognizer_agreement`: 14/14 pass (existing matrix from #1888
already guards the committed shape; no new test needed for a
docstring-only change).
- `cargo test -p onnx-genai-engine --features native-cuda --test
deepseek_v4_tiny_qmoe_e2e --test glm_tiny_full_attention_e2e --test
glm_tiny_qmoe_native_cuda_e2e`: **12/12 pass**, CPU/CUDA tokens
byte-identical to `ANCHOR_IDS` for all 3 fixtures.
- DeepSeek-V4 CUDA: `captures=1 replays=10 fallbacks=0` (native e2e
test) / `captures=5 replays=310 fallbacks=0` (profile_native, n=3
steady). GLM full-attention CUDA: `captures=1 replays=10 fallbacks=0` in
a small run / `captures=3 replays=186 fallbacks=0` (n=3 steady) —
capture state recorded, not forced. GLM IndexShare CUDA: `captures=0
replays=0 fallbacks=0`, growing-KV-prefix decline (unchanged from prior
cycle).

Independent review (code-review agent) caught and this PR fixes: a `\`
vs `\\` line-continuation bug that silently collapsed a documented shell
command onto one line in the deepseek docstring, and confirmed the
default GLM-5.2 IndexShare variant also needs the reemit step (verified
directly via regenerate+reemit-vs-committed diff, since #1883 never
touched that fixture so there's no "pre-1883 blob" to diff against for
it).

## Scope

No fixture regeneration performed or needed. No Mobius changes. No
coarse-residency files touched. Independent of #1854.

Closes the "regenerate missing policies" instruction as **superseded** —
see full task report for the baseline-measurement portion of this cycle.

Signed-off-by: Sebastian <sebastian@squad.local>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…f which was reported

The `CUDA compile (Linux x86_64)` lane has been red since #1836 (`6e4b0ebb3`);
last green was `cb81745b0`. The reported compile error was masking two further
failures, so fixing it alone would have moved the lane from "red at build" to
"red at the honesty script". Two more arrived on `main` while this PR was open,
in #1884 and #1895. All five are fixed here; the lane cannot go green on any
proper subset.

1. Unresolved import (the reported error).

`content_preserving_transition_gpu.rs` imported
`transition_granule_range_with_phase8_faults` unconditionally, but that item is
gated `#[cfg(any(test, feature = "gpu-tests"))]`. The `test` arm does not cover
an integration test: `tests/*.rs` are separate crates linking the library built
*without* `cfg(test)`, so the item is genuinely absent in the
`without-gpu-tests` configuration. Invisible to any local run that passes
`--features gpu-tests`.

Fixed with a two-arm local shim, mirroring `with_faults` in
`crates/onnx-runtime-cuda-memory/tests/virtual_memory_gpu.rs`, which already
solves this exact problem for `CudaVirtualBacking::with_driver_faults`. The
gate itself is left alone: widening it would ship a driver fault injector that
can force `cuMemUnmap`/`cuMemMap`/`cuMemSetAccess` to fail into production, and
`required-features` is precisely what the honesty script exists to forbid
("exists only with gpu-tests enabled; CUDA tests must not hide from CPU
inventory").

#1895 (`1be9f2cc2`) reached this file first, with `#![cfg(feature =
"gpu-tests")]` on the whole target -- the option rejected above. It fixed the
compile and traded it for an inventory failure on the same lane: `main` at
`1be9f2cc2` reports "content_preserving_transition_gpu: Cargo reported no
integration tests" plus 19 x "test exists only with gpu-tests enabled" (job
`97266415620`). The target was not repaired, it was hidden. That line is
removed here and the module doc records why, so the option is not re-tried a
third time.

2. Five non-ignored tests in a `_gpu` target (4 honesty violations).

`verify_cuda_test_honesty.py` requires every test in a `_gpu` target to be
ignored, not passed, on a CPU-only runner -- that is how the suite is stopped
from reporting green for a GPU it never touched.

Three were genuinely CPU-only predicate tests over `verify_safe_point`. Moved
to a new non-`_gpu` target rather than `#[ignore]`d: silencing them would have
greened the lane by deleting coverage, and target naming is the escape hatch
the script's own comment names for CPU-only tests in a CUDA crate.

Moving them alone would *also* have stopped them running. A non-`_gpu` target
is skipped by the honesty script, and `workspace_test_packages.py` deny-lists
this crate from every offline lane, so the target would have been compiled and
never executed. The CUDA lane therefore gains an explicit `--test
content_preserving_transition` step, alongside the two CPU-only targets in this
crate that already have one for the same reason.

The other two asserted nothing -- `fault_injection_safe_point_recheck_rejects`
is comments plus a `println!` and self-describes as "a documentation test";
`zero_len_is_committed_noop` `println!`s that the behaviour is "verified by
implementation". Both passed unconditionally. Removed. The first has real
sibling coverage in `fault_injection_recheck_safe_point_rejected_gpu`; the
second does not -- the `len == 0` early return has no test anywhere. Deleting a
test that asserts nothing loses no coverage, but the gap is pre-existing and
real, and a genuine test needs a device (the early return still takes
`&CudaRuntime`/`&mut CudaReservation`), so it is left for a GPU-capable change
rather than papered over.

Added `safe_point_accepts_a_clean_state`, because the three moved tests only
ever assert `is_err()` and so all three survive a mutant that makes
`verify_safe_point` reject unconditionally. Verified: under that mutant the
inherited three pass and only the new test fails.

3. The manifest guard checked a proxy, not the property it claimed.

The check was the literal substring `"gpu-tests = []"`, so #1860 turned it red
by changing the value to `["onnx-runtime-cuda-memory/gpu-tests"]` -- a
legitimate and necessary forwarding. Now matches the feature *key* whatever its
value, extracted into `declares_gpu_tests_feature` so it is covered by the
script's own `--self-test` fixtures, which it previously was not.

4. A GPU test in a `_gpu` target with no `#[ignore]` (#1895).

`causal_conv_with_state_gpu::the_standard_domain_spelling_reaches_the_same_kernel`
calls `require_cuda()` exactly like its two siblings but is missing their
`#[cfg_attr(not(feature = "gpu-tests"), ignore = ...)]`, so it ran and failed on
a CPU runner. Fixed by adding the sibling attribute verbatim; no design choice
involved.

5. A CPU-only test inside a `_gpu` target (#1884).

`expert_route_telemetry_probe_gpu::cpu_oracle_and_validator_self_consistent`
passes in both configurations, which the script reports as "executed without
gpu-tests". Its doc comment states the intent plainly: it "runs without a GPU so
the reference cannot silently rot".

`#[ignore]` would clear the checker by destroying exactly that property -- an
ignored test runs nowhere -- so this takes the same route as defect 2. The
pure-CPU oracle (`cpu_bitmap`, `cpu_dedup`, `consume_and_validate`,
`synth_routes` and the header indices) moves verbatim to a shared
`tests/expert_route_oracle/mod.rs`, which is a module and not a target: Cargo
auto-discovers `tests/*.rs` and `tests/*/main.rs` only. Both the `_gpu` target
and a new `expert_route_telemetry_probe` target declare it, so there is one
copy of the oracle and the seven GPU tests keep diffing against the same code
the CPU test checks. The new target gets its own `ci.yml` step, for the reason
in defect 2.

Verified live in its new home by mutation rather than by its own green: perturb
the shift in `cpu_bitmap` -> FAILED, revert -> ok. "It compiles in the new file"
is not evidence that it executes there, which is the defect an earlier review
caught in defect 2's target.

Each fix independently falsified by re-introducing it: import -> E0432;
un-ignored test -> "must be ignored, not pass" (+3 inventory errors); manifest
value -> "must define a gpu-tests feature"; missing `cfg_attr` -> "1 tests
failed while checking ignored status"; relocated CPU test -> "executed without
gpu-tests; CUDA tests must be ignored, not pass".

Honesty script exits 0: 530 tests/79 targets identical in both configurations,
530 ignored without gpu-tests, 0 passed on this no-CUDA host.
`cargo clippy -p onnx-runtime-ep-cuda --features cuda -- -D warnings` clean (the
lane's exact command, both invocations of it); both explicit `ci.yml` test steps
pass 4/4 and 1/1; fmt clean.

Not fixed here: `cargo clippy -p onnx-runtime-ep-cuda --features cuda
--all-targets -- -D warnings` fails on ~8 pre-existing lints in `_gpu` targets
this PR does not touch. No lane runs that command -- both ep-cuda clippy steps
omit `--all-targets` and `workspace_test_packages.py` deny-lists the crate --
so it is a real gap but a separate one, and folding it in would put unrelated
files in a lane-restoration PR.

Closes #1875

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…rtWeightGroup (Cycle-19)

Roy's Cycle-19 review (#1854) found that ExpertWeightGroup members
(fc1/fc2/fc3/scales tensors of one logical MoE expert bank) could each
individually pass capability/alignment checks yet be assigned
DIFFERENT hot/cold expert partitions, silently tiering only part of a
logical expert across Device/Host.

apply_residency_plan_at_boundary now adds a group-level preflight pass
(4a) after the existing per-value precheck: for every multi-member
group it derives one canonical (hot-expert-set, expert_count) from the
first member that clears validation, and requires every other member
to match it exactly (HashSet equality, after each member's own
expert-count/index domain has already been validated). A group that
disagrees on expert_count/index domain, hot-expert set, or contains an
unsupported/missing-metadata member is rejected as a whole -- before
any mutation -- exactly like the existing (approved) group capability
fallback.

For groups that agree, every member's byte ranges are then RE-DERIVED
from the single group-authoritative hot set (via cold_ranges_for),
never from that member's own independently-computed (already-verified
equal) copy. This is a "cannot diverge by construction" fix: even if
the equality check above were ever weakened, no two members of an
agreeing group could still end up with different byte ranges, because
there is only one hot set in play by the time ranges are computed.

Scope, per the reviewer's narrowly-described gap:
- No tensor-name heuristics; graph-derived expert_groups is unchanged.
- No public API signature changes; ExpertWeightGroup, plan_residency,
  and validate_decision (weight.rs) are untouched -- the fix lives
  entirely in coarse_residency.rs's boundary-application logic.
- Per-member byte geometry/offsets may still legitimately differ
  across fc1/fc2/fc3/scales; only the logical hot-expert-index set is
  required to be identical.
- Default-off gate, same-device checks, Fatal/quarantine/progress
  reporting, and rollback fault behavior are unchanged; all previously
  approved tests (1-8) pass unmodified.

New tests (coarse_residency_plan_gpu.rs, tests 9-13), one logical group
with fc1/fc2-equivalent members per case:
- equal hot sets in different order/with duplicates normalize and the
  whole group transitions together;
- disagreeing hot sets between two supported members reject the whole
  group with zero side effects;
- expert-count/index-domain mismatch rejects the whole group with a
  distinct reason, zero side effects;
- an unsupported/missing-metadata member rejects the whole group
  before any mutation;
- a successful group transition followed by a later member's Fatal
  (forced via the existing DriverFaultPlan phase-8 fault-injection
  machinery) rolls back the entire group, including the earlier
  member's already-committed range, with bit-identical content
  preserved throughout.

Rebased onto latest main (through #1884/#1860); no conflicts. Local
validation: build, clippy (-D warnings, filtered to touched files),
fmt, ep-api lib (91/91), ep-cuda lib (543/543, 31 ignored), non-GPU
coarse_residency_plan (4/4), and on an idle A100
(CUDA_VISIBLE_DEVICES=4): coarse_residency_plan_gpu (12/12, including
all 5 new + all 7 previously-approved GPU tests), Slice2-4 GPU
regressions composable_vmm_production_gpu/content_preserving_transition_gpu/
expert_bank_remap_cost_gpu (29/29), and qmoe_gpu correctness (46/46).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…f which was reported

The `CUDA compile (Linux x86_64)` lane has been red since #1836 (`6e4b0ebb3`);
last green was `cb81745b0`. The reported compile error was masking two further
failures, so fixing it alone would have moved the lane from "red at build" to
"red at the honesty script". Two more arrived on `main` while this PR was open,
in #1884 and #1895. All five are fixed here; the lane cannot go green on any
proper subset.

1. Unresolved import (the reported error).

`content_preserving_transition_gpu.rs` imported
`transition_granule_range_with_phase8_faults` unconditionally, but that item is
gated `#[cfg(any(test, feature = "gpu-tests"))]`. The `test` arm does not cover
an integration test: `tests/*.rs` are separate crates linking the library built
*without* `cfg(test)`, so the item is genuinely absent in the
`without-gpu-tests` configuration. Invisible to any local run that passes
`--features gpu-tests`.

Fixed with a two-arm local shim, mirroring `with_faults` in
`crates/onnx-runtime-cuda-memory/tests/virtual_memory_gpu.rs`, which already
solves this exact problem for `CudaVirtualBacking::with_driver_faults`. The
gate itself is left alone: widening it would ship a driver fault injector that
can force `cuMemUnmap`/`cuMemMap`/`cuMemSetAccess` to fail into production, and
`required-features` is precisely what the honesty script exists to forbid
("exists only with gpu-tests enabled; CUDA tests must not hide from CPU
inventory").

#1895 (`1be9f2cc2`) reached this file first, with `#![cfg(feature =
"gpu-tests")]` on the whole target -- the option rejected above. It fixed the
compile and traded it for an inventory failure on the same lane: `main` at
`1be9f2cc2` reports "content_preserving_transition_gpu: Cargo reported no
integration tests" plus 19 x "test exists only with gpu-tests enabled" (job
`97266415620`). The target was not repaired, it was hidden. That line is
removed here and the module doc records why, so the option is not re-tried a
third time.

2. Five non-ignored tests in a `_gpu` target (4 honesty violations).

`verify_cuda_test_honesty.py` requires every test in a `_gpu` target to be
ignored, not passed, on a CPU-only runner -- that is how the suite is stopped
from reporting green for a GPU it never touched.

Three were genuinely CPU-only predicate tests over `verify_safe_point`. Moved
to a new non-`_gpu` target rather than `#[ignore]`d: silencing them would have
greened the lane by deleting coverage, and target naming is the escape hatch
the script's own comment names for CPU-only tests in a CUDA crate.

Moving them alone would *also* have stopped them running. A non-`_gpu` target
is skipped by the honesty script, and `workspace_test_packages.py` deny-lists
this crate from every offline lane, so the target would have been compiled and
never executed. The CUDA lane therefore gains an explicit `--test
content_preserving_transition` step, alongside the two CPU-only targets in this
crate that already have one for the same reason.

The other two asserted nothing -- `fault_injection_safe_point_recheck_rejects`
is comments plus a `println!` and self-describes as "a documentation test";
`zero_len_is_committed_noop` `println!`s that the behaviour is "verified by
implementation". Both passed unconditionally. Removed. The first has real
sibling coverage in `fault_injection_recheck_safe_point_rejected_gpu`; the
second does not -- the `len == 0` early return has no test anywhere. Deleting a
test that asserts nothing loses no coverage, but the gap is pre-existing and
real, and a genuine test needs a device (the early return still takes
`&CudaRuntime`/`&mut CudaReservation`), so it is left for a GPU-capable change
rather than papered over.

Added `safe_point_accepts_a_clean_state`, because the three moved tests only
ever assert `is_err()` and so all three survive a mutant that makes
`verify_safe_point` reject unconditionally. Verified: under that mutant the
inherited three pass and only the new test fails.

3. The manifest guard checked a proxy, not the property it claimed.

The check was the literal substring `"gpu-tests = []"`, so #1860 turned it red
by changing the value to `["onnx-runtime-cuda-memory/gpu-tests"]` -- a
legitimate and necessary forwarding. Now matches the feature *key* whatever its
value, extracted into `declares_gpu_tests_feature` so it is covered by the
script's own `--self-test` fixtures, which it previously was not.

The key is looked for only inside the `[features]` table. Review falsified the
first version of this fix: a file-wide regex also accepts a *dependency* named
`gpu-tests`, or one under `[target.'cfg(...)'.dependencies]` -- the same proxy
mistake in a new costume, inside the fix for that mistake. Twelve fixtures now,
including both false-positive shapes.

4. A GPU test in a `_gpu` target with no `#[ignore]` (#1895).

`causal_conv_with_state_gpu::the_standard_domain_spelling_reaches_the_same_kernel`
calls `require_cuda()` exactly like its two siblings but is missing their
`#[cfg_attr(not(feature = "gpu-tests"), ignore = ...)]`, so it ran and failed on
a CPU runner. Fixed by adding the sibling attribute verbatim; no design choice
involved.

5. A CPU-only test inside a `_gpu` target (#1884).

`expert_route_telemetry_probe_gpu::cpu_oracle_and_validator_self_consistent`
passes in both configurations, which the script reports as "executed without
gpu-tests". Its doc comment states the intent plainly: it "runs without a GPU so
the reference cannot silently rot".

`#[ignore]` would clear the checker by destroying exactly that property -- an
ignored test runs nowhere -- so this takes the same route as defect 2. The
pure-CPU oracle (`cpu_bitmap`, `cpu_dedup`, `consume_and_validate`,
`synth_routes` and the header indices) moves verbatim to a shared
`tests/expert_route_oracle/mod.rs`, which is a module and not a target: Cargo
auto-discovers `tests/*.rs` and `tests/*/main.rs` only. Both the `_gpu` target
and a new `expert_route_telemetry_probe` target declare it, so there is one
copy of the oracle and the seven GPU tests keep diffing against the same code
the CPU test checks. The new target gets its own `ci.yml` step, for the reason
in defect 2.

Verified live in its new home by mutation rather than by its own green: perturb
the shift in `cpu_bitmap` -> FAILED, revert -> ok. "It compiles in the new file"
is not evidence that it executes there, which is the defect an earlier review
caught in defect 2's target.

Each fix independently falsified by re-introducing it: import -> E0432;
un-ignored test -> "must be ignored, not pass" (+3 inventory errors); manifest
value -> "must define a gpu-tests feature"; missing `cfg_attr` -> "1 tests
failed while checking ignored status"; relocated CPU test -> "executed without
gpu-tests; CUDA tests must be ignored, not pass".

Honesty script exits 0: 530 tests/79 targets identical in both configurations,
530 ignored without gpu-tests, 0 passed on this no-CUDA host.
`cargo clippy -p onnx-runtime-ep-cuda --features cuda -- -D warnings` clean (the
lane's exact command, both invocations of it); both explicit `ci.yml` test steps
pass 4/4 and 1/1; fmt clean.

Not fixed here: `cargo clippy -p onnx-runtime-ep-cuda --features cuda
--all-targets -- -D warnings` reports 23 pre-existing errors and dies on
`index_share_gpu` and `qmoe_zero_copy_cold_expert_spike_gpu` before reaching the
rest, so 23 is a floor, not a total. None are in a file this PR touches -- no
diagnostic location matches any of them. No lane runs that command: both
ep-cuda clippy steps omit `--all-targets` and `workspace_test_packages.py`
deny-lists the crate. A real gap, but a separate one; folding it in would put
unrelated files in a lane-restoration PR.

Closes #1875

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
…f which was reported (#1881)

Closes #1875.

`CUDA compile (Linux x86_64)` has been red since **#1836
(`6e4b0ebb3`)**; last green was `cb81745b0`. Resch handed this off
explicitly rather than guessing at another domain's intended API
visibility — that was the right call, because **the reported compile
error was masking two further failures on the same lane.** Fixing only
the error would have moved it from "red at build" to "red at the honesty
script".

**Two more arrived on `main` while this PR was open** — #1884 and #1895,
both today. Five independent defects from five different PRs, reviewable
separately. The lane cannot go green on any proper subset of them.

---

### 1. Unresolved import — the reported error

`content_preserving_transition_gpu.rs` imports
`transition_granule_range_with_phase8_faults`, gated `#[cfg(any(test,
feature = "gpu-tests"))]`. The `test` arm does not cover an integration
test: `tests/*.rs` are separate crates linking the library built
*without* `cfg(test)`, so in the `without-gpu-tests` configuration the
item is genuinely absent. Invisible to any local run passing `--features
gpu-tests`.

**I rejected both options in the issue, with evidence:**

- *Widen the gate* contradicts the function's own doc — "Not reachable
from production" — and would ship an API that forces `cuMemUnmap` /
`cuMemMap` / `cuMemSetAccess` to fail to every consumer.
- *`required-features`* is exactly what the honesty lane exists to
forbid: `compare_inventories` emits *"exists only with gpu-tests
enabled; CUDA tests must not hide from CPU inventory"*. It trades a
compile error for a lane failure and deletes the CPU-side inventory of
14 tests.

**Fix:** a two-arm local shim, mirroring `with_faults` in
`crates/onnx-runtime-cuda-memory/tests/virtual_memory_gpu.rs`, which
already solves this identical problem for
`CudaVirtualBacking::with_driver_faults`. The gate stays intact; the
target compiles and lists its tests in both configurations. The
`unreachable!` arm is genuinely unreachable — after change 2 every test
in that file is `#[ignore]`d.

> **#1895 reached this file first, and took the option rejected above.**
> `1be9f2cc2` applied `#![cfg(feature = "gpu-tests")]` to the whole
target. It fixed the compile and **traded it for an inventory failure on
the same lane.** From the job log on `main@1be9f2c` (job
`97266415620`), not inferred:
> ```
> - content_preserving_transition_gpu: Cargo reported no integration
tests
> + 19 × "test exists only with gpu-tests enabled"
> ```
> The target was not repaired, it was **hidden** — the symptom stopped
being reported without the property becoming true. Its module doc argued
that a compile failure "is not a useful signal about a suite that cannot
run on a machine with no GPU anyway", which is the exact proposition
`verify_cuda_test_honesty.py` exists to reject. This PR removes that
line and records the reasoning in the module doc so the option isn't
tried a third time.

### 2. Five non-ignored tests in a `_gpu` target — 4 honesty violations

The script requires every test in a `_gpu` target to be **ignored, not
passed** on a CPU-only runner. That rule is how the suite is stopped
from reporting green for a GPU it never touched.

**Three were genuinely CPU-only** predicate tests over
`verify_safe_point` — pure logic over `ResizeSafePoint`, no driver.
**Moved** to a new non-`_gpu` target, not `#[ignore]`d: silencing them
would green the lane *by deleting coverage*, and target naming is the
escape hatch the script's own comment names, so that "a genuinely
CPU-only target is not policed as a device test merely because it lives
in a CUDA crate".

> **Correction after review — moving them was not sufficient, and my
first commit shipped the same defect this PR is about.**
> A non-`_gpu` target is skipped by the honesty script, and
`.github/scripts/workspace_test_packages.py:26` deny-lists this crate
from every offline lane. So the new target was **compiled and never
executed** — the moved tests, and the mutation guard I added
specifically to catch a bad `verify_safe_point`, would have run on
**zero** CI runs. My claim that moving "keeps them executing on every
run" was exactly backwards: before the move the honesty script *did*
execute them (that was the violation); after it, nothing did.
> This is the same failure mode as the defect being fixed — *a test that
isn't on the leg you're citing* — and it would have produced a green
check proving nothing. Fixed by adding an explicit `--test
content_preserving_transition` step to the CUDA lane, mirroring the two
CPU-only targets in this crate that already have one for precisely this
reason (`ci.yml:989-998`). Verified by running the added command
verbatim: 4 passed.

**Two asserted nothing** and are removed:
- `fault_injection_safe_point_recheck_rejects` — comments plus a
`println!`; self-describes as "this test documents the contract".
- `zero_len_is_committed_noop` — `println!`s that the behaviour is
"verified by implementation".

Both passed unconditionally. This touches another author's tests, so I
flag it explicitly rather than burying it in the diff. Deleting a test
that asserts nothing loses no coverage — but to be precise about what is
and isn't covered: the first *does* have real sibling coverage in
`fault_injection_recheck_safe_point_rejected_gpu`, and **the second does
not** — the `len == 0` early return is untested anywhere. That gap is
pre-existing, and a genuine test needs a device (the early return still
takes `&CudaRuntime`/`&mut CudaReservation`), so I have left it for a
GPU-capable change rather than claim coverage that doesn't exist.

**Added `safe_point_accepts_a_clean_state`.** The three moved tests only
ever assert `is_err()`, so they cannot distinguish "rejects the unsafe
field" from "rejects everything". Empirically confirmed — with
`verify_safe_point` mutated to reject unconditionally:

```
test result: FAILED. 3 passed; 1 failed
failures:
    safe_point_accepts_a_clean_state
```

All three inherited tests survive the mutant. Only the new one kills it.

### 3. The manifest guard checked a proxy, not the property it claimed

```python
if "gpu-tests = []" not in manifest:
    errors.append(f"crates/{crate.name}/Cargo.toml must define a gpu-tests feature")
```

A literal substring match. **#1860 turned this red** by changing the
value to `["onnx-runtime-cuda-memory/gpu-tests"]` — a legitimate and
necessary forwarding. The guard's message claims to check that the
feature *is defined*; it actually checked that it was defined *with one
particular empty value*.

Now matches the feature **key** whatever its value, extracted into
`declares_gpu_tests_feature()` so it is covered by the script's own
`--self-test` fixtures — which it previously was not. Fixtures include
the forwarding form, a multi-line list, a commented-out line, and a
`features = ["gpu-tests"]` dependency mention that must *not* count.

> Note for @resch: #1860 broke this lane as well as the shape-inference
pin you fixed in #1870 — same commit, two unrelated pins that both
encoded a value rather than the property.

### 4. A GPU test in a `_gpu` target with no `#[ignore]` — #1895


`causal_conv_with_state_gpu::the_standard_domain_spelling_reaches_the_same_kernel`
calls `require_cuda()` exactly like its two siblings in the same file,
but is missing their `#[cfg_attr(not(feature = "gpu-tests"), ignore =
…)]`. So it ran on a CPU runner and failed:

```
- causal_conv_with_state_gpu: 1 tests failed while checking ignored status
- causal_conv_with_state_gpu: Cargo inventory has 3 tests but libtest reported 2 ignored
```

**Fix:** add the sibling attribute verbatim. No design choice involved —
a four-line omission.

### 5. A CPU-only test inside a `_gpu` target — #1884


`expert_route_telemetry_probe_gpu::cpu_oracle_and_validator_self_consistent`
passes in **both** configurations, which the script reports as `executed
without gpu-tests`. Its doc comment states the intent plainly:

> *CPU-only: the oracle and the boundary validator are self-consistent.
Runs without a GPU so the reference cannot silently rot.*

That intent is correct and worth keeping. **`#[ignore]` would clear the
checker by destroying exactly the property the test was written for** —
an ignored test runs nowhere, so the reference could then rot precisely
as its author feared, with a green check over it. Same reasoning as
defect 2, so the same route.

**Fix:** the pure-CPU oracle — `cpu_bitmap`, `cpu_dedup`,
`consume_and_validate`, `synth_routes`, the `Decision` enum and the
header indices — moves **verbatim** to a shared
`tests/expert_route_oracle/mod.rs`, and the test moves to a new
`expert_route_telemetry_probe` target that declares it.
`tests/expert_route_oracle/mod.rs` is a module, not a target: Cargo
auto-discovers `tests/*.rs` and `tests/*/main.rs` only. Both targets
declare it, so there is **one** copy of the oracle and the seven GPU
tests keep diffing against the same code the CPU test checks — the
alternative, duplicating it, would let the two copies drift and quietly
defeat the point.

The new target gets its own `ci.yml` step, for the reason in defect 2's
correction.

**Verified live in its new home by mutation, not by its own green.**
Perturbing the shift in `cpu_bitmap`:

```
test result: FAILED. 0 passed; 1 failed
failures:
    cpu_oracle_and_validator_self_consistent
```

and `ok` again on revert. "It compiles in the new file" is not evidence
that it *executes* there — which is the defect the review caught in
defect 2, so this time it is checked rather than asserted.

> Defects 4 and 5 are in other people's very recent files. No open PR
touches either (checked across all open PRs), so there is no
concurrent-edit hazard, and neither fix requires domain knowledge of the
CUDA kernels or the telemetry design — only of where a test must live to
be run honestly.

---

## Validation

Every fix independently falsified by re-introducing it — each re-reds
with a **distinct** error:

| reverted | resulting failure |
|---|---|
| the shim | `error[E0432]: unresolved import ...with_phase8_faults` |
| one un-ignored test | `1 tests executed without gpu-tests; CUDA tests
must be ignored, not pass` + 3 inventory errors |
| the manifest check | `must define a gpu-tests feature` |
| the `cfg_attr` (4) | `1 tests failed while checking ignored status` |
| the relocation (5) | `1 tests executed without gpu-tests; CUDA tests
must be ignored, not pass` |

```
CUDA test honesty check passed: 530 tests/79 targets without gpu-tests (530 ignored),
530 tests/79 targets with gpu-tests (474 fail-loud, 56 ignored, 0 passed on this no-CUDA host)
```

- `cargo clippy -p onnx-runtime-ep-cuda --features cuda -- -D warnings`
— clean (the lane's exact command, and both of the two places `ci.yml`
invokes it)
- `cargo test -p onnx-runtime-ep-cuda --features cuda --lib` — **538
passed, 0 failed**
- both explicit CI test steps, run verbatim —
`content_preserving_transition` 4 passed, `expert_route_telemetry_probe`
1 passed
- `cargo fmt --all --check` — clean
- Rebased onto `496597764`; honesty script re-run green on the exact
tree being shipped (not on an earlier one — main moved four times during
this work).
- Independent Opus review — found the never-executed-target defect
above; all other categories (shim signature/arg-order vs the real
function across all 11 params and 7 call sites, error-string assertion,
regex, target classification, and every honesty-script rule) checked and
ruled out.

**One observation I am not fixing here.** The lane's clippy step omits
`--all-targets`, so this crate's test code is linted by **no** lane —
`workspace_test_packages.py` also deny-lists it. Running it now yields
23 errors and dies on `index_share_gpu` and
`qmoe_zero_copy_cold_expert_spike_gpu` before reaching the rest, all
pre-existing and none in a file this PR touches (verified: no diagnostic
location matches any of them). Unrelated and out of scope — folding it
in would put a pile of unrelated files into a lane-restoration PR — but
it is the same absent-coverage shape that let defects 2, 4 and 5 sit
un-noticed.

### Credit

@Gaff established that this is the **only** instance of the class in the
workspace — 24 `cfg(any(test, feature = …))` sites, the other 23 either
unreferenced from `tests/` or referenced only from gated positions — and
stated the root cause more crisply than I had: an item so gated is
broken **only when a `tests/` file names it from a position not itself
gated**, in practice a top-level `use`, because a `use` resolves
unconditionally regardless of how the functions below it are gated. A
gated call site is safe. He also retracted his own grep-based falsifier
for the class on #1817 after running the negative control on it.

One correction to his note, since it changes what the precedent
demonstrates: `onnx-runtime-cuda-memory` does **not** gate the
consumers. `vmm_release_quarantine_gpu.rs:59-67` is a **two-arm shim** —
a `#[cfg(feature = "gpu-tests")]` helper *and* a `#[cfg(not(feature =
"gpu-tests"))]` `unreachable!` twin — with its callers ungated. That
distinction is load-bearing here: genuinely gating the consumers would
delete the tests from the base inventory and red this lane under
`compare_inventories`. So the crate is still the right worked example;
it is an example of the shim, which is what this PR adopts (making it
the third in the repo, after `virtual_memory_gpu.rs::with_faults` and
`vmm_release_quarantine_gpu.rs::install_faults`).

🤖 Generated with [Copilot
CLI](https://githubnext.com/projects/copilot-cli)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
… test it means (#1911)

`CUDA compile (Linux x86_64)` is **red on `main` again**, one commit
after #1881 restored it.

```
CUDA test honesty check failed:
  - activations_gpu: 1 tests failed while checking ignored status
  - activations_gpu: Cargo inventory has 4 tests but libtest reported 3 ignored
```

#1905 added
`activations_gpu::mish_matches_cpu_including_the_saturating_tail`
without the `#[cfg_attr(not(feature = "gpu-tests"), ignore = …)]` that
its **three siblings in the same file** carry, so it runs and fails on a
CPU runner. Fix is the sibling attribute, verbatim.

No criticism of @-the-author intended, and I want to be explicit about
why. **This is the fourth instance of the same omission in three days**
— #1884, #1895, and now #1905. That is a property of the *signal*, not
of the authors: while the lane is red for an unrelated reason, adding a
`_gpu` test without an ignore is free, because the only check that would
object is a job nobody can distinguish from already-broken. #1881
restored the lane; this keeps it restored.

---

## The second half: the checker didn't say which test

The message above names a **count**, in a four-test target, and the run
that produced it was a CI job on a merge commit. Finding out which test
needed a rebase and a grep.

A check whose entire job is to notice that a `_gpu` test was added
without an ignore should **say which one**. `run_libtest` now parses
libtest's per-test outcome lines, so the same failure reads:

```
activations_gpu: 1 tests failed while checking ignored status
  (mish_matches_cpu_including_the_saturating_tail)
```

**Verified on the real path, not only in fixtures.** With the `cfg_attr`
reverted, the full script emits exactly that line; restored, it exits 0.
Fixtures alone would only have proved the formatter works.

Naming also applies to `executed without gpu-tests` and `passed with
gpu-tests on a no-CUDA host`, which have the same problem — those are
the two messages that fired for #1884 and #1895. `IgnoredResult` and
`ActiveResult` now share a `LibtestResult` base for it.

It **degrades to the bare count** if the per-test lines can't be parsed,
rather than rendering an empty `()`. The count is still true, so a
future change in libtest's output format must not turn a real failure
into a confusing one.

Four new `--self-test` fixtures cover the naming, the inventory-mismatch
message, the empty-name degradation, and the outcome regex itself —
because a guard whose own correctness is unverified is precisely the
defect this checker exists to catch, and I'd rather not add one to it
while fixing it.

## Validation

```
CUDA test honesty check passed: 531 tests/79 targets without gpu-tests (531 ignored),
531 tests/79 targets with gpu-tests (475 fail-loud, 56 ignored, 0 passed on this no-CUDA host)
```

- run on this branch, based on current `main` (`7e274a4e2`) — not on an
older tree
- `--self-test` passes
- `cargo fmt --all --check` clean
- mutation-effective both ways: revert the `cfg_attr` → the named
failure above; revert the naming → the bare count

Follow-up to #1881. Refs #1875.

🤖 Generated with [Copilot
CLI](https://githubnext.com/projects/copilot-cli)

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
…ndary window (#1810 Slice 7A)

Add default-disabled, producer-only expert-route telemetry (a device-side
bitmap + 6-word header [epoch,request,device,overflow,poison,count]) fused into
the existing QMoE and BlockQuantizedMoE route kernels as inert observability.
Off/on route outputs are byte-identical; PRODUCE performs zero host
readback/sync and zero VMM/map calls in steady-state replay.

Cycle-22 revision (Roy NO-GO on the rejected HEAD 8ec9d82): the prior design
launched a separate reset/epoch kernel on every execute/replay, which
contradicts the approved Slice-6 *coarse-boundary window* contract (design
EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md §2.3/§3: the epoch is bumped and the
record consumed at a safe boundary, NOT per replay). This revision makes the
fused route kernels *accumulate only* and moves reset/epoch-advance to an
explicit, host-ordered coarse boundary.

Invariant / API:
- `arm(request,device,experts)` validates identity, allocates the stable-VA
  record through the existing runtime allocator, and opens window 1 (stamps
  identity, epoch=1, zeroes bitmap/counters). No kernel is compiled or launched.
- On execute/replay the fused `route_telemetry_mark_row` inside
  `qmoe_route`/`bqmoe_route` ONLY accumulates into the stable record: atomicOr
  the routed bit into the bitmap, atomicAdd the in-range count, atomicOr the
  sticky poison bit on an out-of-range id, atomicOr the sticky overflow bit on
  count saturation. The epoch is fixed for the whole window; every eager call
  and captured replay in the window accumulates the routed-expert union + count.
  There is no reset/epoch kernel in the captured graph.
- `reset_route_telemetry_boundary()` is the ONLY place the window advances: it
  is rejected while the EP stream is capturing/replaying (fail closed), else it
  drains prior stream work via the existing `drain_for_unmap` authority, bumps
  the host-side epoch, and re-stamps identity + zeroes the record so the next
  window starts empty with no stale carryover. It allocates nothing, moves no
  pointer, and touches no PMM/VMM/cache/global coordinator.
- Overflow saturates and fail-closes into the sticky overflow bit (never wraps
  into a smaller "success" value); out-of-range routes poison. Snapshot/validate
  run on the host at a boundary against an already-copied record.
- Multi-request/multi-device/kernel instances are isolated (distinct records);
  the bitmap VA is stable across windows and captures. Epoch is a host counter,
  so the record footprint drops the former 4-byte device epoch buffer.

Scope (unchanged from Slice 7A): crate-internal / test-only / default-OFF. No
lifecycle/policy/consume wiring. Ordinary inference is byte-identical and has
zero overhead when disarmed.

Tests (rewritten to the new semantics, same commit):
- off/on output byte-identity;
- eager calls accumulate the CPU-oracle union + count within a window at a fixed
  epoch, then a boundary reset opens an empty epoch-2 window;
- >=3 graph replays with different routes accumulate the union at a fixed epoch
  and a stable VA (no per-replay reset);
- boundary reset increments the epoch and starts an empty window;
- boundary reset rejected during an active capture (epoch unchanged);
- capacity/device mismatch stays inert and never fails inference;
- multi-instance request/device isolation, footprint, teardown/accounting;
- QMoE and BQMoE route match the oracle for M in {1,2,4,8}.
Host unit tests cover poison/overflow fail-closed on the validator; the #1884
probe harness continues to cover device-level poison/overflow.

Measurement (G1 gate, idle pinned A100, DeepSeek-V2-Lite decode layer, CUDA
graph replay, BATCH=512, RUNS=5, ~8s ramp + first-shape recheck, n=3):
- GPU route+layer overhead 0.948-1.054us (0.82-0.91% of the 116us layer);
- host-enqueue delta ~0us; epoch fixed at 1 across the whole window
  (count accumulated to 30726); output byte-identical; clock drift < 0.5%.
- GATE (<=2us AND <=2%): GO. Telemetry remains default-OFF regardless.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
…ndary window (#1810 Slice 7A)

Add default-disabled, producer-only expert-route telemetry (a device-side
bitmap + 6-word header [epoch,request,device,overflow,poison,count]) fused into
the existing QMoE and BlockQuantizedMoE route kernels as inert observability.
Off/on route outputs are byte-identical; PRODUCE performs zero host
readback/sync and zero VMM/map calls in steady-state replay.

Cycle-22 revision (Roy NO-GO on the rejected HEAD 8ec9d82): the prior design
launched a separate reset/epoch kernel on every execute/replay, which
contradicts the approved Slice-6 *coarse-boundary window* contract (design
EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md §2.3/§3: the epoch is bumped and the
record consumed at a safe boundary, NOT per replay). This revision makes the
fused route kernels *accumulate only* and moves reset/epoch-advance to an
explicit, host-ordered coarse boundary.

Invariant / API:
- `arm(request,device,experts)` validates identity, allocates the stable-VA
  record through the existing runtime allocator, and opens window 1 (stamps
  identity, epoch=1, zeroes bitmap/counters). No kernel is compiled or launched.
- On execute/replay the fused `route_telemetry_mark_row` inside
  `qmoe_route`/`bqmoe_route` ONLY accumulates into the stable record: atomicOr
  the routed bit into the bitmap, atomicAdd the in-range count, atomicOr the
  sticky poison bit on an out-of-range id, atomicOr the sticky overflow bit on
  count saturation. The epoch is fixed for the whole window; every eager call
  and captured replay in the window accumulates the routed-expert union + count.
  There is no reset/epoch kernel in the captured graph.
- `reset_route_telemetry_boundary()` is the ONLY place the window advances: it
  is rejected while the EP stream is capturing/replaying (fail closed), else it
  drains prior stream work via the existing `drain_for_unmap` authority, bumps
  the host-side epoch, and re-stamps identity + zeroes the record so the next
  window starts empty with no stale carryover. It allocates nothing, moves no
  pointer, and touches no PMM/VMM/cache/global coordinator.
- Overflow saturates and fail-closes into the sticky overflow bit (never wraps
  into a smaller "success" value); out-of-range routes poison. Snapshot/validate
  run on the host at a boundary against an already-copied record.
- Multi-request/multi-device/kernel instances are isolated (distinct records);
  the bitmap VA is stable across windows and captures. Epoch is a host counter,
  so the record footprint drops the former 4-byte device epoch buffer.

Scope (unchanged from Slice 7A): crate-internal / test-only / default-OFF. No
lifecycle/policy/consume wiring. Ordinary inference is byte-identical and has
zero overhead when disarmed.

Tests (rewritten to the new semantics, same commit):
- off/on output byte-identity;
- eager calls accumulate the CPU-oracle union + count within a window at a fixed
  epoch, then a boundary reset opens an empty epoch-2 window;
- >=3 graph replays with different routes accumulate the union at a fixed epoch
  and a stable VA (no per-replay reset);
- boundary reset increments the epoch and starts an empty window;
- boundary reset rejected during an active capture (epoch unchanged);
- capacity/device mismatch stays inert and never fails inference;
- multi-instance request/device isolation, footprint, teardown/accounting;
- QMoE and BQMoE route match the oracle for M in {1,2,4,8}.
Host unit tests cover poison/overflow fail-closed on the validator; the #1884
probe harness continues to cover device-level poison/overflow.

Measurement (G1 gate, idle pinned A100, DeepSeek-V2-Lite decode layer, CUDA
graph replay, BATCH=512, RUNS=5, ~8s ramp + first-shape recheck, n=3):
- GPU route+layer overhead 0.948-1.054us (0.82-0.91% of the 116us layer);
- host-enqueue delta ~0us; epoch fixed at 1 across the whole window
  (count accumulated to 30726); output byte-identical; clock drift < 0.5%.
- GATE (<=2us AND <=2%): GO. Telemetry remains default-OFF regardless.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
…ndary window (#1810 Slice 7A) (#1922)

## Slice 7A — inert, default-OFF, producer-only expert-route telemetry
(QMoE/BQMoE)

Closes #1810 (Slice 7A). **Draft. Do not merge.** Crate-internal /
test-only / default-OFF.

> **Cycle-22 revision by Sebastian (Performance Engineer), revision
owner.**
> The original author is under reviewer-protocol lockout and did not
advise, pair, or contribute to this revision. Reworked independently
from the PR diff, the approved Slice-6 design doc, and Roy's Cycle-22
review.

### What changed vs the rejected HEAD (`8ec9d8230`)

Roy's Cycle-22 NO-GO was correct and is **addressed at the design
level**: the rejected design launched a **separate reset/epoch kernel on
every execute/replay**, which contradicts the approved Slice-6
**coarse-boundary window** contract
(`docs/memory/EXPERT_ROUTE_TELEMETRY_SLICE6_DESIGN.md` §2.3/§3 — the
epoch is bumped and the record consumed at a *coarse safe boundary*,
**not per replay**). That per-call reset is why the fused path measured
**2.526 us / 2.18 % -> NO-GO**.

This revision makes the fused route kernels **accumulate-only** and
moves reset/epoch-advance to an explicit, host-ordered **coarse
boundary**. The earlier NO-GO was scoped to the wrong (per-call reset)
design point; the boundary-window producer now measures **GO** (see
below), independently reproduced.

### Invariant / API (producer-only boundary-window contract)

- **`arm(request, device, experts)`** - validates identity, allocates
the stable-VA record via the existing runtime allocator, opens **window
1** (stamps identity, `epoch = 1`, zeroes bitmap/counters). **No kernel
is compiled or launched.**
- **execute / replay** - the fused `route_telemetry_mark_row` inside
`qmoe_route` / `bqmoe_route` **only accumulates** into the stable
record: `atomicOr` the routed bit into the bitmap, `atomicAdd` the
in-range count, `atomicOr` the sticky **poison** bit on an out-of-range
id, `atomicOr` the sticky **overflow** bit on count saturation. The
**epoch is fixed for the whole window**; every eager call and captured
replay in the window accumulates the routed-expert **union + count**.
**No reset/epoch kernel in the captured graph; no host
sync/alloc/drain/VMM on this path.**
- **`reset_route_telemetry_boundary()`** - the **only** place the window
advances. **Rejected while the EP stream is capturing/replaying** (fail
closed, checked *before* any drain); otherwise drains prior stream work
via the existing `drain_for_unmap` authority, bumps the **host-side**
epoch, and re-stamps identity + zeroes the record so the next window
starts empty with **no stale carryover**. Allocates nothing, moves no
pointer, touches no PMM/VMM/cache/global coordinator.
- **Fail-closed** - overflow **saturates** into the sticky overflow bit
(never wraps into a smaller "success"); out-of-range routes **poison**
without touching the bitmap; snapshot/validate run on the host at a
boundary against an already-copied record.
- **Isolation / stability** - multi-request / multi-device / multiple
kernel instances use distinct records; the bitmap **VA is stable**
across windows and captures. Epoch is a host counter, so the record
footprint **drops the former 4-byte device epoch buffer** (`footprint =
4*ceil(experts/32) + 24`).
- **Reuses only existing authorities** - `alloc_raw`/`free_raw`,
`htod`/`dtoh`, `is_capturing`, `drain_for_unmap`, `ordinal`. No new
coordinator/allocator/cache/PMM/VMM and **no host sync in steady-state
replay**.

Scope unchanged: **default-OFF, crate-internal, test-only**. Ordinary
inference is **byte-identical** and has **zero overhead when disarmed**.
No lifecycle/policy/consume wiring in this slice.

### Tests (rewritten to the new semantics, same commit)

QMoE (`tests/qmoe_gpu.rs`) and BQMoE
(`tests/block_quantized_moe_gpu.rs`) `mod route_telemetry`:

- off/on output **byte-identity**;
- eager calls **accumulate the CPU-oracle union + count** within a
window at a **fixed epoch**, then a boundary reset opens an empty
**epoch-2** window;
- **>=3 graph replays** with different routes accumulate the union at a
**fixed epoch** and a **stable VA** (QMoE - proves no per-replay reset);
- boundary reset **increments the epoch and starts an empty window**;
- boundary reset **rejected during an active capture** (epoch
unchanged);
- capacity / device mismatch stays **inert** and never fails inference;
- multi-instance **request/device isolation**, footprint,
teardown/accounting;
- **QMoE and BQMoE routes for M in {1, 2, 4, 8}** match the oracle.

Host unit tests cover **poison/overflow fail-closed** on the validator;
the #1884 probe harness continues to cover device-level poison/overflow.

**Results** (idle A100, GPU 5, `--test-threads=1`): QMoE
`route_telemetry` 10 passed; BQMoE 8 passed; telemetry host unit tests 7
passed; #1884 probe functional regressions 6 passed; **full QMoE suite
56 passed**, **full BQMoE suite 13 passed** (no regressions to
#1788/#1800/#1884/#1854).

### Measurement - G1 gate (`microbench_fused_route_telemetry_g1_gate`)

Idle **pinned A100**, serialized off/on back-to-back, whole QMoE layer
**CUDA-graph captured** and measured via `replay_graph` (host-enqueue
gaps excluded), **BATCH = 512, RUNS = 5**, ~8 s clock ramp + first-shape
recheck, GPU-event and host-enqueue reported separately. Gate
denominator is a **realistic** DeepSeek-V2-Lite decode layer (64
experts, top_k 6, hidden 2048); tiny synthetic is informational only.
**n = 3 independent runs:**

| Run | route+layer overhead (GPU) | % of layer | host-enqueue delta |
epoch | count | output |

|-----|---------------------------|-----------|----------------|-------|-------|--------|
| 1 | 0.976 us | 0.84 % | -0.016 us | fixed = 1 | 30726 | identical |
| 2 | 0.948 us | 0.82 % | -0.006 us | fixed = 1 | 30726 | identical |
| 3 | 1.054 us | 0.91 % | -0.008 us | fixed = 1 | 30726 | identical |

**Range: 0.948-1.054 us, 0.82-0.91 %. Clock drift < 0.5 %.** The epoch
stays **fixed at 1** across the whole BATCH x RUNS replay window (count
accumulates to 30726) - a moving epoch would have proven a forbidden
per-replay reset survived.

**G1 GATE (<= 2 us AND <= 2 %): GO.** Independently reproduced vs Roy's
boundary-only experiment (0.884 us / 0.76 %); same order, same verdict -
not a reused number. Telemetry remains **default-OFF** regardless of the
GO.

### Review

Independent review (excluding the locked-out author and the revision
owner): **APPROVE** - no blocking or substantive findings; confirmed
accumulate-only execute path, reject-before-drain boundary reset,
fail-closed overflow/poison, stable VA, isolation, footprint, and that
the tests assert the *new* semantics (they would fail under the rejected
per-replay-reset design). One MINOR note: the device-side poison branch
of the fused mark is defensive and unreachable-by-construction from the
production route kernel (Roy's prior approved framing) - covered by the
host validator unit tests and the #1884 probe harness.

**Roy's explicit final re-review is requested** before this leaves
draft.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
… coarse residency (#1810 Slice 7B) (#1971)

## #1810 Slice 7B — boundary-time route-telemetry **consumer**

**Status: DRAFT — awaiting independent review. Do not merge.**

Closes the loop opened by the merged Slice-6/7A **producer**
(expert-route telemetry, PR #1922 `e1ec495ee`) and the merged Slice-4/5
**coarse-boundary residency lifecycle** (PR #1854). This is the smallest
production seam the Slice-6 design §8 specifies:

> expose a boundary-time consumer that produces a per-expert
desired-set, and feed that set to the **existing** Slice 4/5
coarse-boundary plan application (`coarse_residency.rs`) as its policy
input — with **no** new allocator, **no** id→slot rewrite, and every
mapping change still owned by PMM/VMM.

### What it does

New module `crates/onnx-runtime-ep-cuda/src/route_residency.rs`:

`consume_route_window_at_boundary(...)`:
1. **Gate** — default-off via the existing `COARSE_RESIDENCY_ENABLE_ENV`
(`coarse_residency_profile_enabled()`). When off (shipped default) it
returns `Disabled` *before* reading the snapshot or touching any
allocator.
2. **Safe boundary** — re-reads the existing
`CudaWeightResidency::resize_safe_point` and fails closed with
`RejectedNotSafeBoundary { reason }` if a graph is capturing/replaying,
an admission is in flight, a deferred release has not settled, execution
is multi-device, or a routed-residency guard is live. Coarse-boundary
only — never a per-token remap.
3. **Validate** — runs the producer's own `consume_and_validate` on the
already-completed window snapshot; fail-closed (→ `WholeBank { reason
}`) on poison / overflow / stale epoch / foreign request / foreign
device, or when no in-range experts were recorded.
4. **Plan** — turns the routed-expert union into a desired **hot set**
and asks the already-validated `StaticProfileResidencyPolicy` to shape a
`ResidencyPlan` (the design's "record → desired set"
`RouteObserverPolicy` role — **reused, not duplicated**, so exactly one
validated policy emits `PerExpertCandidate`).
5. **Apply** — hands the plan to the existing
`CudaWeightResidency::apply_coarse_residency_plan`, which remains the
**sole** authority that maps / unmaps / accounts / quarantines / rolls
back through PMM/VMM.

`RouteWindowConsumeOutcome` carries the exact reason on every
non-applied path — **no silent fallback**.

### Invariants held by construction

- **Allocates nothing**, opens no stream, owns no VA, copies no device
bytes — pure host glue between two existing authorities.
- **No remap during capture/replay**; coarse-boundary only, never
per-token.
- **No new host sync** in steady state (only the producer's
already-taken snapshot and the existing transition primitive's drain).
- **Default-off & byte-identical** when disarmed (two independent
default-off switches: telemetry disarmed *and* this gate off).
- **Same-device fail-closed** and **PMM/VMM remains the sole
mapping/accounting/quarantine/rollback authority**.

### Tests — `tests/route_residency_consume_gpu.rs` (6 GPU tests, all
passing on an idle A100)

- `disabled_gate_is_structural_no_op` — off path is a structural no-op
(byte-identical).
- `route_window_hot_set_transitions_cold_experts` — routed hot-set stays
resident, cold set tiers to host, bytes identical.
- `expert_group_transitions_atomically_from_window` — atomic
expert-group transition driven from a window.
- `active_capture_and_multi_device_reject_consume` — active-capture and
multi-device boundaries reject the consume (fail-closed).
- `foreign_identity_and_defective_windows_fail_closed` — foreign
request/device and poison/overflow/stale windows fail closed to
whole-bank.
- `injected_fault_rolls_back_consumer_transition` — injected driver
fault (`fail_nth(Unmap, 3)` over an isolated + merged cold range) rolls
back range-precisely and quarantines; `rollback_count == 1`,
`values_touched == 0`, `committed_values` empty, content bit-identical
(mirrors the proven `coarse_residency_plan_gpu.rs` rollback fixture).

Run:
```
env -u ONNX_GENAI_WEIGHT_OFFLOAD_COARSE_RESIDENCY_ENABLE CUDA_VISIBLE_DEVICES=<idle> ONNX_GENAI_CUDA_DEVICE=0 \
  cargo test -p onnx-runtime-ep-cuda --features cuda,cuda-13000,gpu-tests --release \
  --test route_residency_consume_gpu -- --ignored --test-threads=1
```

### Honest scope

Like `coarse_residency::apply_residency_plan_at_boundary` when it
shipped (Slice 5), this consumer has **no live decode-loop call site
yet** — wiring it into a running session's request boundary is the next
slice. It ships here as the production seam: reachable, default-off, and
proven by the GPU tests. `cargo fmt` applied; `cargo clippy` on the
crate lib is clean for the new file (pre-existing unrelated
warnings/errors in `optimizer.rs` / `standard_attention.rs` are out of
scope).

### Constraints honored

Did not touch PagedAttention or IQ1 fusion files. Branch
`squad/1810-slice7b-telemetry-residency-consume` based on latest `main`
(`011fbb284`, includes #1922 `e1ec495ee`). #1788/#1800/#1884/#1854
lifecycle regressions preserved (reused, not forked).

Refs #1810.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant