Repository navigation
bench(cpu): print the realized decode width next to every t=N row - #1770
Conversation
`int4_decode_loop_ab` sweeps `ONNX_GENAI_CPU_DECODE_THREADS` and prints a row per width, but nothing in the row reported the width the run actually built. At least four paths silently reduce it below the request -- the pre-clamp in `resolve_persistent_decode_threads_with_override`, both headroom reservations, and the single-CPU-cpuset branch that drops decode to the flat path entirely -- and only one is even `NXRT_CALIB_DEBUG`-visible. A `t=N` row was a label, not a measurement of width N. That is not hypothetical. Before #1766 the benches never called `initialize()`, so the process was unbounded and the global Rayon pool ran full width at every budget. Measured at 9747b49 (current main minus #1766), qwen/acc0/block32/s=1, same binary otherwise: t=1 40.298 ms/token 100% CPU t=2 37.166 ms/token 267% CPU <- 1.08x for 2.67 cores That flat bottom end is what "ONNX_GENAI_CPU_DECODE_THREADS=2 does nothing" was: not a dispatch pathology, an unbounded-topology artifact. On current main the same sweep scales 1.98x / 3.91x / 7.55x / 15.0x at t=2/4/8/16, reproducible within 2% sweeping in both directions. Reports rather than asserts. A reduced width is legitimate when the host genuinely cannot honour the request (an 8-lane budget in a 2-CPU cpuset), and aborting there would leave a constrained container unable to benchmark at all. The `WIDTH-MISMATCH` token is for the caller -- human or matrix script -- to discard or mark the row, as a contended cell is already marked UNTRUSTED. Called after the phases, never before: the pool is built lazily at first decode and `decode_width()` is deliberately non-forcing, so reading it earlier reports `path=unresolved` -- and if it were forcing it would build the pool and change the topology being measured. Scoped to `int4_decode_loop_ab`. `gqa_decode` also reads a decode width but pins the ambient Rayon pool by design, so a width line there would report `path=flat` and mislead rather than verify. Refs #1763, #1766 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1770 +/- ##
==========================================
- Coverage 80.67% 80.58% -0.10%
==========================================
Files 411 411
Lines 199008 199008
Branches 199008 199008
==========================================
- Hits 160551 160366 -185
- Misses 33016 33201 +185
Partials 5441 5441
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
…d verdict Adversarial review caught a claim in the new doc that my own measurement in this branch had already refuted: `reserve_single_group_headroom` does not reduce the realized width. It reduces the *spawned thread* count, but it only runs in the single-group case, which is exactly where `dispatcher_owns_a_shard` (`shards.len() == 1 && shards[0].workers < requested`) is true, so `total_workers` adds the lane straight back. Measured: a 2-lane budget on a 2-CPU cpuset spawns one thread and realizes two lanes, `as_requested`. The genuine net reducers are three, not four: the pre-clamp to `available_parallelism`, `reserve_split_headroom` (NUMA-split only, and uncompensated because the dispatcher shard is single-group-only), and the single-CPU-cpuset fallback. Also corrects the same list in the shipped `DecodeWidth` doc from #1764, which had the mirror-image error: it named both headroom reservations, omitted the pre-clamp, and claimed all three report through `report_spmd_fallback`. Only the cpuset fallback does -- neither headroom function logs anything at all (call sites are :569, :1619, :1642, and neither :1749 nor :1763 is among them). Splits `WIDTH-MISMATCH` into `WIDTH-MISMATCH` (both widths known, unequal -- the row's label is wrong) and `WIDTH-UNRESOLVED` (never resolved -- no decode reached the pool). A matrix script wants to treat those differently. Verified by execution, all three verdicts: t=4 plain -> requested=4 realized=4 spmd-pool as_requested t=8 under taskset -c 0,2 -> requested=8 realized=2 spmd-pool WIDTH-MISMATCH t=1 plain -> requested=1 realized=1 flat as_requested The mismatch arm is the pre-clamp reducer caught in the act. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Adversarial review (Opus) — verdict and fixesNo BLOCKING findings. One NIT was a genuine defect and is fixed, plus a mirror-image error it exposed in already-shipped code. The real finding: my doc contradicted my own measurementThe doc claimed I had already measured this in this branch before writing the sentence: One spawned thread, two realized lanes, Genuine net reducers are three: the pre-clamp to Mirror-image error in shipped #1764The review's cross-check exposed the same list wrong in the opposite direction in Also fixedSplit the failure token, since a matrix script wants to treat these differently:
Verified by execution, all three verdictsThe mismatch arm is the pre-clamp reducer caught in the act — precisely the silent path that made a Reviewer questions answered
Gates
|
What
int4_decode_loop_absweepsONNX_GENAI_CPU_DECODE_THREADSand prints a row per width, but nothing in the row reported the width the run actually built. Addscommon::report_decode_width()and calls it after the measured phases.Why this is not hypothetical
At least four paths silently reduce realized width below the request — the pre-clamp in
resolve_persistent_decode_threads_with_override, both headroom reservations, and the single-CPU-cpuset branch that drops decode to the flat path. Only one is evenNXRT_CALIB_DEBUG-visible.ONNX_GENAI_CPU_DECODE_THREADS=2was recorded indocs/benchmarks/2026-08-21-int4-acc4-execution-regime.mdas producing timings identical to=1. The row was dropped rather than explained.It was an unbounded-topology artifact, and it is already fixed. Measured at
9747b4971(current main minus #1766), same binary otherwise, qwen/acc0/block32/s=1:Before #1766 the benches never called
initialize(), so the process was unbounded and global Rayon ran full width at every budget. The "t=1" arm was not 1-wide.On current main, same sweep, both directions (<2% between passes):
A/A null control at t=4: 10.179 vs 10.077 (1.0%).
Design notes
WIDTH-MISMATCHis for the caller to discard or mark the row, as a contended cell is already marked UNTRUSTED.decode_width()is deliberately non-forcing; reading it earlier reportspath=unresolved, and a forcing read would build the pool and change the topology being measured.int4_decode_loop_ab.gqa_decodealso reads a decode width but pins the ambient Rayon pool by design, so a width line there would reportpath=flatand mislead rather than verify.Validation
cargo fmt --checkclean;clippy --benches --all-targets -D warningscleanrealizedtracks the request at every width,pathflipsflat→spmd-poolat t=2Bench-only; no shipped code changes.
Refs #1763, #1766