diff --git a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py index f2a8f2733a..1d9f704fd0 100755 --- a/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py +++ b/crates/onnx-runtime-ep-cpu/benches/acc0_gap_matrix.py @@ -166,7 +166,7 @@ def native(binary, model, block, acc, threads, sessions, tokens, reps, extra=Non if extra: env.update(extra) r = sh(f"taskset -c {PIN} {binary}", env) - steady, width_line, cpu = None, None, None + steady, width_line, cpu, dispatcher = None, None, None, None for line in r.stdout.splitlines() + r.stderr.splitlines(): if line.strip().startswith("steady"): f = line.split() @@ -174,6 +174,12 @@ def native(binary, model, block, acc, threads, sessions, tokens, reps, extra=Non "tps": float(f[3]), "spread": float(f[4])} if line.strip().startswith("decode_width"): width_line = line.strip() + # Optional, same contract as the `cpu` row: only binaries built after + # the dispatcher-placement change emit it. Carries the last token, the + # verdict, so a caller can reject an intervention arm whose pin did not + # take instead of scoring it as a control. + if line.strip().startswith("dispatcher reserved_cpu="): + dispatcher = line.split()[-1] # Optional: only binaries built after the CPU-accounting change emit # it, so its absence is not an error -- older harnesses and older # binaries keep working and simply carry no `cpu_*` keys. @@ -208,6 +214,8 @@ def native(binary, model, block, acc, threads, sessions, tokens, reps, extra=Non raise RuntimeError(f"native width vacuous: {width_line}") if cpu: steady["cpu"] = cpu + if dispatcher: + steady["dispatcher"] = dispatcher return steady diff --git a/crates/onnx-runtime-ep-cpu/benches/acc0_w16_dispersion.py b/crates/onnx-runtime-ep-cpu/benches/acc0_w16_dispersion.py new file mode 100644 index 0000000000..56f5d07523 --- /dev/null +++ b/crates/onnx-runtime-ep-cpu/benches/acc0_w16_dispersion.py @@ -0,0 +1,243 @@ +#!/usr/bin/env python3 +"""Score an A/B run for *dispersion*, not for its median. + +Written and committed BEFORE the run it scores. The rule below is the whole +point of the file; reading it after seeing the numbers and adjusting it would +make it worthless. + +Why a separate script +--------------------- +`acc0_w16_blocktime_ab.py` is a validated instrument -- its thresholds, A/A +arm, arm rotation and width-8 regression guard have been replayed against +archived JSON and reproduce published numbers exactly. A dispersion claim needs +a different statistic, so it gets a different file rather than another edit to +that one. This script *runs nothing*: it reads the JSON the A/B already wrote, +so the two claims come from the same launches and cannot disagree about which +host state they describe. + +The question +------------ +Width-16 A/B work on this host is blocked by its own noise floor. Two +*identical* arms measured in the same launch differ by a median of ~11-22% and +by as much as 58%, which is large enough that no width-16 intervention of +realistic size can clear a pre-registered bar -- a +23% steal-tiles candidate +was rejected on exactly this ground, with the A/A half-width at 21.5% in the +same run. + +That dispersion is not symmetric noise. Pooling 90 archived arms: + +* At width 8, arms slower than 1.3x the best are 5/45, sit at execution + positions 1 and 2 (never 3), and every one of them has an intra-arm rep + spread of >=23% against a 6.9% median. They are self-detectable. +* At width 16 they are 15/45 -- three times as common -- spread evenly over + all three positions, and NOT self-detectable: their median intra-arm spread + is 22.8% against 18.9% for the normal arms, and one slow arm was internally + consistent to 0.5% while running 1.3x slow for every one of its reps. + +An arm that is uniformly slow for its whole life, with a tight internal spread, +is a per-process state, not a disturbance. The hypothesis this scores is that +the state is thread placement, and specifically the placement of the *inline +dispatcher*: the pool reserves a CPU for it (`DISPATCHER_RESERVED_CPUS`, +justified by a measured 1.57x) but nothing binds it there, so where it settles +is decided per launch by the scheduler. + +The rule (pre-registered) +------------------------- +Per arm, over trusted launches, using each launch's `tps_rep` at +`PRIMARY_WIDTH`: + + D(arm) = (p90(tps) - p10(tps)) / median(tps) + +`p90 - p10` rather than max-min so a single launch cannot define the result, +and normalised by the median so the two arms are comparable in scale. + + PIN-STABILISES iff D(test) <= DISPERSION_RATIO * D(control) + and n_trusted >= MIN_TRUSTED + +Self-test, evaluated FIRST and able to veto: + + The A/A arm is the same configuration as its own reference arm, so its + dispersion estimates the same quantity. If + + |D(aa) - D(ref)| > SELFTEST_TOLERANCE * mean(D(aa), D(ref)) + + the estimator is too noisy at this n to support any dispersion claim, and + the verdict is REPORT NOTHING no matter what the arms did. + + `ref` is the arm the A/A repeats: `aa_arm` names it per launch, and the + harness rotates it, so this is resolved per launch rather than assumed. + +Note on what the A/A arm is here. In `acc0_w16_blocktime_ab.py` the A/A is a +*second run of one arm within the same launch*, so it measures within-launch +repeatability. That is the right null for the median rule. For dispersion it is +also the right self-test, because a dispersion estimator that cannot reproduce +itself across two runs of one configuration cannot distinguish two others. + +This is a dispersion rule only. It says nothing about which arm is faster; the +median rule in the A/B script answers that, and the two are reported side by +side deliberately, because "faster" and "more reproducible" are different +claims and an intervention can win one and lose the other. + +What it returned +---------------- +On the dispatcher-pin run this was written for (`7e274a4e2`, 16 launches, 15 +trusted) the **self-test vetoed the run**: |D(aa) - D(ref)| = 0.1432 against an +allowed 0.5 x 0.2591 = 0.1296. Two arms of identical configuration disagreed +about dispersion by more than the estimator's own tolerance, so no dispersion +claim was made -- D(control) = 0.3610 and D(test) = 0.0780 are unscored +observations and are not a result. An earlier 6-launch run scored +PIN-STABILISES (0.2781 -> 0.0416); it did not survive the larger n. Recorded +here because a rule that has only ever fired positively is a rule nobody has +tested. See `docs/benchmarks/2026-08-24-acc0-dispatcher-placement.md`. +""" + +import argparse +import json +import statistics +import sys + +PRIMARY_WIDTH = 16 +MIN_TRUSTED = 6 +DISPERSION_RATIO = 0.5 +SELFTEST_TOLERANCE = 0.5 + + +def pctl(values, q): + """Linear-interpolated quantile; `statistics.quantiles` needs n >= 2.""" + ordered = sorted(values) + if len(ordered) == 1: + return ordered[0] + pos = q * (len(ordered) - 1) + low = int(pos) + high = min(low + 1, len(ordered) - 1) + return ordered[low] + (ordered[high] - ordered[low]) * (pos - low) + + +def dispersion(values): + if len(values) < 2: + return float("nan") + median = statistics.median(values) + if median <= 0: + return float("nan") + return (pctl(values, 0.90) - pctl(values, 0.10)) / median + + +def arm_series(cells, arm): + """Per-launch `tps_rep` for one arm at the primary width.""" + out = [] + for cell in cells: + width = cell["widths"].get(str(PRIMARY_WIDTH)) + if not width: + continue + entry = width.get(arm) + if not entry or "cpu" not in entry or "tps_rep" not in entry["cpu"]: + continue + out.append(entry["cpu"]["tps_rep"]) + return out + + +def aa_reference_series(cells, control_value): + """The arm each launch's A/A repeats, resolved per launch via `aa_arm`. + + The harness rotates which arm is doubled, so a fixed choice here would + compare the A/A against the *other* configuration in half the launches -- + which would not be a self-test at all. + """ + out = [] + for cell in cells: + width = cell["widths"].get(str(PRIMARY_WIDTH)) + if not width: + continue + arm = "control" if width.get("aa_arm") == control_value else "test" + entry = width.get(arm) + if not entry or "cpu" not in entry or "tps_rep" not in entry["cpu"]: + continue + out.append(entry["cpu"]["tps_rep"]) + return out + + +def pin_state(cells): + """Every distinct dispatcher verdict seen, per arm. + + Non-vacuity, checked rather than assumed: an intervention arm whose pin did + not take is not a test arm, and scoring it as one would silently report a + null result as a negative one. `PIN-MISSED` or a missing row anywhere in + the test arm invalidates the run. + """ + seen = {} + for cell in cells: + width = cell["widths"].get(str(PRIMARY_WIDTH)) + if not width: + continue + for arm in ("control", "test", "aa"): + entry = width.get(arm) + if not entry: + continue + seen.setdefault(arm, set()).add(entry.get("dispatcher", "MISSING")) + return {arm: sorted(v) for arm, v in seen.items()} + + +def report(cells, control_value): + lines = [] + trusted = [c for c in cells if c.get("trusted")] + n = len(trusted) + lines.append(f"n_trusted={n} of {len(cells)} launches") + if n < MIN_TRUSTED: + lines.append(f"REPORT NOTHING (n_trusted={n} < {MIN_TRUSTED})") + return lines + + control = arm_series(trusted, "control") + test = arm_series(trusted, "test") + aa = arm_series(trusted, "aa") + ref = aa_reference_series(trusted, control_value) + + d_control, d_test, d_aa, d_ref = ( + dispersion(control), dispersion(test), dispersion(aa), dispersion(ref)) + lines.append(f"D(control)={d_control:.4f} D(test)={d_test:.4f}") + lines.append(f"self-test: D(aa)={d_aa:.4f} D(aa's own arm)={d_ref:.4f}") + + mean_pair = (d_aa + d_ref) / 2.0 + gap = abs(d_aa - d_ref) + if mean_pair <= 0 or gap > SELFTEST_TOLERANCE * mean_pair: + lines.append( + f"SELF-TEST FAILED: |D(aa)-D(ref)|={gap:.4f} > " + f"{SELFTEST_TOLERANCE} x {mean_pair:.4f}") + lines.append("REPORT NOTHING (the dispersion estimator cannot " + "reproduce itself at this n)") + return lines + lines.append(f"self-test passed: |D(aa)-D(ref)|={gap:.4f} <= " + f"{SELFTEST_TOLERANCE} x {mean_pair:.4f}") + + states = pin_state(trusted) + lines.append(f"dispatcher verdicts: {states}") + bad = [s for s in states.get("test", []) if "PIN-TOOK" not in s] + if bad: + lines.append(f"VACUOUS: test arm did not pin ({bad})") + return lines + + lines.append(f"medians: control={statistics.median(control):.2f} tps " + f"test={statistics.median(test):.2f} tps " + f"ratio={statistics.median(test) / statistics.median(control):.4f}") + if d_test <= DISPERSION_RATIO * d_control: + lines.append(f"PIN-STABILISES: D(test)={d_test:.4f} <= " + f"{DISPERSION_RATIO} x D(control)={d_control:.4f}") + else: + lines.append(f"NOT STABILISED: D(test)={d_test:.4f} > " + f"{DISPERSION_RATIO} x D(control)={d_control:.4f}") + return lines + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("json", help="output of acc0_w16_blocktime_ab.py") + ap.add_argument("--control-value", default="0", + help="the --control value the A/B ran with; used to " + "resolve which arm each launch's A/A repeats") + args = ap.parse_args() + cells = json.load(open(args.json)) + for line in report(cells, args.control_value): + print(line) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/crates/onnx-runtime-ep-cpu/benches/common/mod.rs b/crates/onnx-runtime-ep-cpu/benches/common/mod.rs index a68d49b3ba..ef5d29078e 100644 --- a/crates/onnx-runtime-ep-cpu/benches/common/mod.rs +++ b/crates/onnx-runtime-ep-cpu/benches/common/mod.rs @@ -370,3 +370,52 @@ pub fn report_decode_width() { width.path, ); } + +/// Report the dispatcher's reserved CPU against the CPU it is actually on. +/// +/// The non-vacuity check for `ONNX_GENAI_CPU_DECODE_DISPATCHER_PIN`, and the +/// measurement that motivated it. Two independent facts, never inferred from +/// each other: which CPU the headroom reserve kept clear, and which CPU the +/// dispatching thread is running on now. +/// +/// Both come from the pool, which samples the dispatcher from *inside* the +/// dispatch path. Neither can be obtained here. The first version of this read +/// `sched_getcpu()` on the reporting thread and was exactly inverted: the +/// reporter is idle while the pool works, so with the dispatcher unpinned the +/// scheduler parks the reporter on the one free core, and pinning the +/// dispatcher *evicts* it -- so "unpinned" read as on-the-reserved-CPU and +/// "pinned" read as off it. Reading the wrong thread does not merely add noise, +/// it can invert the sign. The second version read the dispatcher's own +/// `/proc/self/task//stat`, which parses correctly but almost always +/// returns nothing: the dispatcher is a transient thread and has usually exited +/// by the time a bench reports. +/// +/// `PIN-TOOK` only when the knob was asked for and the two agree; `PIN-MISSED` +/// when it was asked for and they do not, which is a failed intervention and +/// must not be scored as a control. With the knob off this is pure +/// observation -- `observed` is where the scheduler left the dispatcher, which +/// is the quantity the experiment is about. +pub fn report_dispatcher_cpu() { + let pools = onnx_runtime_ep_cpu::decode_spmd::pools(); + let reserved = pools.and_then(|p| p.dispatcher_cpu()); + let observed = pools.and_then(|p| p.dispatcher_observed_cpu()); + let tid = pools.and_then(|p| p.dispatcher_thread_id()); + let moves = pools.map(|p| p.dispatcher_cpu_changes()); + let requested = onnx_runtime_ep_cpu::decode_spmd::dispatcher_pin_requested(); + let verdict = match (requested, reserved, observed) { + (false, _, _) => "PIN-OFF", + (true, None, _) => "PIN-UNRESERVED", + (true, Some(_), None) => "PIN-UNOBSERVABLE", + (true, Some(r), Some(o)) if r == o => "PIN-TOOK", + (true, Some(_), Some(_)) => "PIN-MISSED", + }; + let show = |v: Option| v.map_or("none".to_string(), |v| v.to_string()); + println!( + "dispatcher reserved_cpu={} requested={} tid={} observed_cpu={} moves={} {verdict}", + show(reserved), + u8::from(requested), + tid.map_or("none".to_string(), |v| v.to_string()), + show(observed), + moves.map_or("none".to_string(), |v| v.to_string()), + ); +} diff --git a/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs b/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs index 0e2dab25e6..55701ed7ca 100644 --- a/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs +++ b/crates/onnx-runtime-ep-cpu/benches/int4_decode_loop_ab.rs @@ -484,4 +484,7 @@ fn main() { // After the phases, never before: the pool is built at first decode, so this // is the earliest point the realized width exists to be read. common::report_decode_width(); + // Same reason, same place: the reservation only exists once the pool does, + // and `sched_getcpu` must be read on the thread that dispatched. + common::report_dispatcher_cpu(); } diff --git a/crates/onnx-runtime-ep-cpu/src/decode_spmd.rs b/crates/onnx-runtime-ep-cpu/src/decode_spmd.rs index 44847beb10..e128308d76 100644 --- a/crates/onnx-runtime-ep-cpu/src/decode_spmd.rs +++ b/crates/onnx-runtime-ep-cpu/src/decode_spmd.rs @@ -91,7 +91,9 @@ #[cfg(test)] use std::cell::Cell; use std::cell::UnsafeCell; -use std::sync::atomic::{AtomicBool, AtomicU32, AtomicU64, AtomicUsize, Ordering, fence}; +use std::sync::atomic::{ + AtomicBool, AtomicI64, AtomicU32, AtomicU64, AtomicUsize, Ordering, fence, +}; use std::sync::{Arc, Mutex, OnceLock}; use std::thread::{self, JoinHandle}; use std::time::{Duration, Instant}; @@ -1125,6 +1127,61 @@ pub struct SpmdDecodePools { /// underneath it was not -- the count was never the thing that was wrong. /// A label nothing can check is a label that drifts. worker_cpus: Vec>, + /// The allowed CPU that [`DISPATCHER_RESERVED_CPUS`] freed for the inline + /// dispatcher, or `None` when the pool has no dispatcher shard or the + /// reservation could not be located (unpinned workers, empty CPU list). + /// + /// Reserving the CPU and *using* it are two different things. The + /// reservation is made in [`reserve_single_group_headroom`] / + /// [`reserve_split_headroom`] and only guarantees that no worker is pinned + /// there; the dispatcher itself is an ordinary unpinned thread that the + /// scheduler is free to leave on a worker's core while the reserved one + /// sits idle. Recording the reserved CPU here is what lets + /// [`Self::dispatcher_observed_cpu`] check whether that actually happens, + /// and [`DISPATCHER_PIN_ENV`] test whether closing the gap is worth + /// anything. Measured answer so far: it is worth 9.5%, which is below the + /// bar that was set for it. + dispatcher_cpu: Option, + /// OS thread id of the first thread to dispatch on this pool, or `0` before + /// any dispatch has happened. + /// + /// Recorded because the dispatcher is *not* the thread that owns the pool, + /// nor in general the process's main thread, so "which CPU is the + /// dispatcher on" cannot be answered by asking whoever is reporting. The + /// first attempt at this measurement read `sched_getcpu()` on the reporting + /// thread and produced an exactly inverted result -- with the dispatcher + /// unpinned the reporter sat on the reserved CPU (it is idle while the pool + /// works, so the scheduler parks it on the one free core), and pinning the + /// dispatcher *evicted* the reporter from that CPU. Both readings were of + /// the wrong thread. + /// + /// Written once per pool via `compare_exchange` from the dispatch path's + /// existing per-thread one-shot, so it costs one syscall per process and + /// nothing on the steady path. + dispatcher_tid: AtomicI64, + /// The CPU the dispatcher was last seen on, sampled every + /// [`DISPATCHER_CPU_SAMPLE_MASK`] + 1 dispatches, or `-1` before the first + /// sample. + /// + /// Sampled inside the dispatch path because the dispatcher is a transient + /// thread: by the time a harness reports, `/proc/self/task/` is + /// usually already gone, so its placement has to be recorded while it is + /// still running. + dispatcher_observed_cpu: AtomicI64, + /// How many consecutive dispatcher CPU samples differed from the one + /// before. + /// + /// A lower bound on migrations, not a count of them: sampling every + /// [`DISPATCHER_CPU_SAMPLE_MASK`] + 1 dispatches sees a thread that left + /// and came back as no change at all. That is the right direction for the + /// question it answers -- "does the unpinned dispatcher stay put?" -- since + /// any non-zero reading is a migration that definitely happened, while zero + /// is only evidence of stillness at this sampling rate. + /// + /// Sampled from one thread only (see `DISPATCHER_IS_RECORDED`), so a + /// successfully pinned dispatcher reads exactly zero and this doubles as a + /// check on the pin. + dispatcher_cpu_changes: AtomicU64, } #[derive(Clone, Copy, Debug, Eq, PartialEq)] @@ -1200,6 +1257,15 @@ impl SpmdDecodePools { *last += 1; } let total_workers = total_threads + usize::from(dispatcher_shard.is_some()); + // The CPU the headroom reservation freed. Workers take `shard.cpus` in + // order (`worker % len`), so index `shard.workers` is the first CPU of + // that node that no worker was pinned to -- exactly the CPU the reserve + // exists to keep clear. The dispatcher's shard lives on the last node, + // so that node's spare is the one it should sit on. + let dispatcher_cpu = dispatcher_shard.and_then(|_| { + let last = shards.last()?; + last.cpus.get(last.workers).copied() + }); #[cfg(feature = "mlas")] if schedule == DecodeSchedule::Steal { @@ -1224,6 +1290,11 @@ impl SpmdDecodePools { worker_cpus: Vec::new(), total_workers: total_threads, schedule, + // No inline dispatcher on this path, so nothing reserved one. + dispatcher_cpu: None, + dispatcher_tid: AtomicI64::new(0), + dispatcher_observed_cpu: AtomicI64::new(-1), + dispatcher_cpu_changes: AtomicU64::new(0), }; } @@ -1423,6 +1494,10 @@ impl SpmdDecodePools { total_workers, schedule, worker_cpus, + dispatcher_cpu, + dispatcher_tid: AtomicI64::new(0), + dispatcher_observed_cpu: AtomicI64::new(-1), + dispatcher_cpu_changes: AtomicU64::new(0), } } @@ -1481,6 +1556,68 @@ impl SpmdDecodePools { &self.worker_cpus } + /// The allowed CPU that the dispatcher reservation freed, if any. + /// + /// This is a *reservation*, not a placement: no worker is pinned here, but + /// nothing pins the dispatcher here either unless [`DISPATCHER_PIN_ENV`] is + /// on. Exposed so a harness can check the two independently -- which CPU + /// was kept clear, and which CPU the dispatching thread actually ran on -- + /// rather than inferring one from the other. + pub fn dispatcher_cpu(&self) -> Option { + self.dispatcher_cpu + } + + /// OS thread id of the thread that dispatched on this pool, once one has. + /// + /// `None` before the first dispatch, and on platforms with no stable + /// per-thread id. The dispatcher is neither the pool's builder nor + /// necessarily the process's main thread, so this is the only reliable way + /// to ask where the dispatcher is actually running. + pub fn dispatcher_thread_id(&self) -> Option { + let tid = self.dispatcher_tid.load(Ordering::Relaxed); + (tid != 0).then_some(tid) + } + + /// The CPU the dispatcher was last sampled on, or `None` before any + /// dispatch (and on platforms without `sched_getcpu`). + /// + /// A *sample*, not a residence: an unpinned dispatcher migrates, so equality + /// with [`Self::dispatcher_cpu`] means "it was there when last looked", and + /// only a pinned dispatcher can be said to stay. + pub fn dispatcher_observed_cpu(&self) -> Option { + let cpu = self.dispatcher_observed_cpu.load(Ordering::Relaxed); + (cpu >= 0).then_some(cpu as usize) + } + + /// Observed dispatcher CPU changes between consecutive samples. + /// + /// See [`Self::dispatcher_cpu_changes`]: a lower bound on migrations, and + /// necessarily zero once the dispatcher is pinned. + pub fn dispatcher_cpu_changes(&self) -> u64 { + self.dispatcher_cpu_changes.load(Ordering::Relaxed) + } + + /// Record the calling thread's current CPU as the dispatcher's placement, + /// and count the sample as a change if it moved since the last one. + fn sample_dispatcher_cpu(&self) { + #[cfg(target_os = "linux")] + { + // SAFETY: `sched_getcpu` takes no arguments and only reads the + // calling thread's current CPU. + let cpu = unsafe { libc::sched_getcpu() }; + if cpu < 0 { + return; + } + let cpu = i64::from(cpu); + let previous = self.dispatcher_observed_cpu.swap(cpu, Ordering::Relaxed); + // `previous < 0` is the first sample, which has nothing to differ + // from and must not be counted as a move. + if previous >= 0 && previous != cpu { + self.dispatcher_cpu_changes.fetch_add(1, Ordering::Relaxed); + } + } + } + /// Whether the pinned workers occupy distinct physical cores. /// /// `None` only when the question is unanswerable: an unpinned pool (no @@ -1653,6 +1790,77 @@ impl SpmdDecodePools { /// dispatcher at a time. When another thread already owns them this runs /// the shards inline instead of waiting or racing -- see /// [`SharedState::dispatching`]. + /// Record this thread as the dispatcher and, if [`DISPATCHER_PIN_ENV`] is + /// on, bind it to the CPU the headroom reserve freed. At most once per + /// thread. + /// + /// The tid is recorded whether or not the pin is requested: with the knob + /// off, "which CPU did the scheduler leave the dispatcher on" is the + /// measurement the knob exists to answer, and it needs the same identity. + /// + /// The pin attempt is recorded before it is made, so a host that refuses it + /// costs one failed syscall for the life of the thread rather than one per + /// op. + /// + /// Called from [`Self::dispatch`] rather than at build time because the + /// pool is a process-wide static built on whichever thread decodes first, + /// which need not be the thread that goes on to dispatch. + fn bind_dispatcher_to_reserved_cpu(&self) { + let tick = DISPATCHER_TICK.with(|t| { + let seen = t.get(); + t.set(seen.wrapping_add(1)); + seen + }); + if tick != 0 { + if tick & DISPATCHER_CPU_SAMPLE_MASK == 0 + && DISPATCHER_IS_RECORDED.with(std::cell::Cell::get) + { + self.sample_dispatcher_cpu(); + } + return; + } + // First dispatcher wins the identity slot: it is the one whose + // placement is sampled from here on. Later dispatching threads still + // take the reserved CPU below -- the point of the reserve is that + // whoever is dispatching sits there -- they just do not report. + let recorded = match current_thread_os_id() { + Some(tid) => self + .dispatcher_tid + .compare_exchange(0, tid, Ordering::Relaxed, Ordering::Relaxed) + .is_ok(), + None => false, + }; + DISPATCHER_IS_RECORDED.with(|flag| flag.set(recorded)); + let Some(cpu) = self.dispatcher_cpu else { + if recorded { + self.sample_dispatcher_cpu(); + } + return; + }; + if !dispatcher_pin_requested() { + if recorded { + self.sample_dispatcher_cpu(); + } + return; + } + match crate::decode_affinity::pin_current_thread_to_cpu(cpu) { + Ok(()) => report_dispatcher_pin(&format!( + "{DISPATCHER_PIN_ENV} on: dispatcher pinned to reserved cpu {cpu}" + )), + Err(message) => report_dispatcher_pin(&format!( + "{DISPATCHER_PIN_ENV} on, but pinning the dispatcher to reserved cpu \ + {cpu} failed: {message}; dispatcher left unpinned" + )), + } + // After the pin, never before: the first sample is the baseline every + // later one is compared against, so taking it pre-pin would score the + // pin itself as a migration and a successfully pinned dispatcher could + // never read zero. + if recorded { + self.sample_dispatcher_cpu(); + } + } + fn dispatch(&self, job: &F) where F: Fn(usize) + Sync, @@ -1671,6 +1879,7 @@ impl SpmdDecodePools { .as_ref() .expect("fixed SPMD dispatch requires shared worker state"); shared.panic_if_poisoned(); + self.bind_dispatcher_to_reserved_cpu(); unsafe fn call(data: *const (), global_index: usize) where F: Fn(usize) + Sync, @@ -3057,6 +3266,99 @@ fn node_shards_with( /// subscribed. const DISPATCHER_RESERVED_CPUS: usize = 1; +/// Opt-in: bind the inline dispatcher to the CPU [`DISPATCHER_RESERVED_CPUS`] +/// reserved for it (`1`/`on`/`true`/`yes`). Default off. +/// +/// The reservation already keeps one allowed CPU clear of workers, because a +/// dispatcher sharing a core with a worker makes that worker a straggler the +/// whole barrier waits on -- 1.57x on qwen int4 at 16 cores, which is why +/// [`reserve_single_group_headroom`] exists. But reserving a CPU does not put +/// anything on it: the dispatcher is an ordinary unpinned thread, and nothing +/// stops the scheduler leaving it on a worker's core while the reserved CPU +/// sits idle. Whether that happens is now measurable rather than assumed -- +/// [`SpmdDecodePools::dispatcher_observed_cpu`] samples the dispatcher's actual +/// CPU during the run, and a harness can compare it against +/// [`SpmdDecodePools::dispatcher_cpu`]. If the two diverge, the collision the +/// reserve was built to prevent is happening anyway, non-deterministically per +/// launch, which would also make it a candidate explanation for the +/// launch-to-launch dispersion that swamps width-16 A/B work. +/// +/// Off by default deliberately, and for two independent reasons. +/// +/// The first is that it has not earned a default. Measured against the +/// pre-registered single-knob rule on `7e274a4e2` -- 16 launches, 15 trusted -- +/// it is faster in **15 of 15** launches and the median gain is **1.0953**, +/// under the 1.10 the rule requires: **REJECT**. A companion rule about +/// launch-to-launch dispersion failed its own self-test and certified nothing. +/// An earlier 6-launch run scored 1.1910 and ACCEPT; it did not replicate. The +/// mechanism is also unproven -- the migration counter says the unpinned +/// dispatcher moves at most once per launch, which is far too little to explain +/// anything, so whatever this does is not "it stops migrating". See +/// `docs/benchmarks/2026-08-24-acc0-dispatcher-placement.md`. +/// +/// The second is that binding the dispatching thread is not free of +/// consequence: it is the session thread, so it keeps that affinity after the +/// decode loop ends, and a subsequent prefill on the same thread would run +/// one-CPU-wide. Turning this into a default needs evidence that covers prefill +/// as well as decode, and that evidence does not exist yet. The knob is here so +/// the decode half can be measured at all. +pub const DISPATCHER_PIN_ENV: &str = "ONNX_GENAI_CPU_DECODE_DISPATCHER_PIN"; + +/// Whether [`DISPATCHER_PIN_ENV`] asks for the dispatcher to be pinned. +/// Read once per process, like every other decode knob, so a mid-run +/// environment change cannot make two ops disagree. +pub fn dispatcher_pin_requested() -> bool { + static REQUESTED: OnceLock = OnceLock::new(); + *REQUESTED.get_or_init(|| { + std::env::var(DISPATCHER_PIN_ENV) + .ok() + .map(|raw| dispatcher_pin_from_raw(Some(raw.as_str()))) + .unwrap_or(false) + }) +} + +/// Parse of [`DISPATCHER_PIN_ENV`], split out so the accepted spellings are +/// directly testable without touching process environment. +fn dispatcher_pin_from_raw(raw: Option<&str>) -> bool { + matches!( + raw.map(|value| value.trim().to_ascii_lowercase()) + .as_deref(), + Some("1" | "on" | "true" | "yes") + ) +} + +/// Sample the dispatcher's CPU once every this many dispatches (minus one). +/// +/// The dispatcher's *placement* is the quantity under study, and it is not a +/// constant: an unpinned thread migrates, so a single reading taken at the +/// first dispatch describes startup rather than steady state. Sampling costs +/// one `sched_getcpu` -- a vDSO read, no syscall -- per 1024 dispatches, which +/// at ~400 barriers per token is under three tokens' spacing and immaterial +/// beside the barrier itself. +const DISPATCHER_CPU_SAMPLE_MASK: u32 = 1023; + +thread_local! { + /// Per-dispatching-thread tick. Zero means "this thread has not dispatched + /// before", which drives the one-shot identity record and pin attempt; the + /// low bits then drive periodic CPU sampling. + /// + /// The pin is attempted exactly once per thread, success or failure. A + /// retry loop would put a failing syscall on the hot path forever on any + /// host that refuses the pin. + static DISPATCHER_TICK: std::cell::Cell = const { std::cell::Cell::new(0) }; + /// Whether this thread is the one whose id the pool recorded. + /// + /// A process can dispatch from more than one thread over its life -- a + /// session per phase is enough -- and every one of them takes the reserved + /// CPU, which is the intended behaviour: the point is that whoever is + /// dispatching sits there. But *placement sampling* has to follow a single + /// thread or it reports thread changes as movement. Measured before this + /// distinction existed, a pinned dispatcher read 2 to 7 "moves" when it + /// could only ever have made one -- the samples were coming from different + /// threads. + static DISPATCHER_IS_RECORDED: std::cell::Cell = const { std::cell::Cell::new(false) }; +} + /// Cap a single pinned worker group so at least [`DISPATCHER_RESERVED_CPUS`] /// allowed CPU stays free for the inline dispatcher. /// @@ -3111,8 +3413,40 @@ fn reserve_split_headroom(shards: &mut [NodeShard]) { } } -/// Log the first persistent-pool fallback/pinning problem once so a restricted -/// or unsupported host surfaces the reason without spamming every worker. +/// The calling thread's OS thread id, or `None` where the platform exposes no +/// stable per-thread id. Linux only in practice: this exists so a harness can +/// find the dispatcher's `/proc/self/task/` entry, which has no +/// counterpart elsewhere. +fn current_thread_os_id() -> Option { + #[cfg(target_os = "linux")] + { + // SAFETY: `gettid` takes no arguments, cannot fail, and returns the + // calling thread's kernel id. + let tid = unsafe { libc::gettid() }; + (tid > 0).then_some(i64::from(tid)) + } + #[cfg(not(target_os = "linux"))] + { + None + } +} + +/// Report the dispatcher-pin outcome once. Separate static from +/// [`report_spmd_fallback`]'s: this is not a fallback, and folding the two +/// would let whichever fired first silence the other. +fn report_dispatcher_pin(message: &str) { + static REPORTED: OnceLock<()> = OnceLock::new(); + if REPORTED.set(()).is_ok() { + #[cfg(feature = "tracing")] + tracing_crate::debug!(dispatcher_pin = %message, "cpu decode dispatcher pin"); + #[cfg(not(feature = "tracing"))] + if std::env::var("NXRT_CALIB_DEBUG").is_ok() { + eprintln!("onnx-genai: {message}"); + } + } +} + +/// Log the first persistent-pool fallback/pinning problem once so a restricted/// or unsupported host surfaces the reason without spamming every worker. /// Emitted as `tracing::debug!` when the `tracing` feature is enabled, or /// gated behind `NXRT_CALIB_DEBUG` otherwise. fn report_spmd_fallback(message: &str) { @@ -3627,6 +3961,92 @@ mod tests { SpmdDecodePools::build_with_schedule(&shards, DecodeSchedule::Fixed, true) } + #[test] + fn dispatcher_pin_env_accepts_only_affirmative_spellings() { + for on in ["1", "on", "true", "yes", " ON ", "True"] { + assert!( + dispatcher_pin_from_raw(Some(on)), + "{on:?} should enable the dispatcher pin" + ); + } + for off in ["0", "off", "false", "no", "", " ", "maybe"] { + assert!( + !dispatcher_pin_from_raw(Some(off)), + "{off:?} should not enable the dispatcher pin" + ); + } + // Unset is off: this knob never opts a host in by accident. + assert!(!dispatcher_pin_from_raw(None)); + } + + #[test] + fn dispatcher_cpu_is_the_cpu_the_headroom_reserve_freed() { + // A node with more CPUs than workers is exactly what the reserve + // produces, and the first unused CPU is the one it kept clear. + let shards = vec![NodeShard { + index: 0, + cpus: vec![0, 2, 4, 6], + workers: 3, + }]; + let pools = SpmdDecodePools::build_with_schedule(&shards, DecodeSchedule::Fixed, true); + assert_eq!(pools.dispatcher_cpu(), Some(6)); + // The reserved CPU is precisely the one no worker claimed. + let taken: Vec> = pools.worker_cpus().to_vec(); + assert!( + !taken.contains(&Some(6)), + "reserved cpu must not also be a worker's pin target: {taken:?}" + ); + pools.shutdown(); + } + + #[test] + fn dispatcher_cpu_is_none_without_a_reservation_or_a_dispatcher_shard() { + // Fully subscribed: every CPU has a worker, so nothing was reserved and + // there is no free CPU to name. Reporting one here would be the + // unverified-label failure the placement accessors exist to avoid. + let full = vec![NodeShard { + index: 0, + cpus: vec![0, 2], + workers: 2, + }]; + let pools = SpmdDecodePools::build_with_schedule(&full, DecodeSchedule::Fixed, true); + assert_eq!(pools.dispatcher_cpu(), None); + pools.shutdown(); + + // No dispatcher shard: the dispatcher computes nothing, so the pool + // makes no claim on a CPU for it even when one is spare. + let spare = vec![NodeShard { + index: 0, + cpus: vec![0, 2, 4], + workers: 2, + }]; + let pools = SpmdDecodePools::build_with_schedule(&spare, DecodeSchedule::Fixed, false); + assert_eq!(pools.dispatcher_cpu(), None); + pools.shutdown(); + } + + #[test] + fn dispatcher_cpu_comes_from_the_node_that_owns_the_dispatcher_shard() { + // `node_worker_counts` adds the dispatcher's shard to the *last* node, + // so the reserved CPU must come from that node -- taking node 0's spare + // would pin the dispatcher on the far side of the barrier it serves. + let shards = vec![ + NodeShard { + index: 0, + cpus: vec![0, 2, 4], + workers: 2, + }, + NodeShard { + index: 1, + cpus: vec![16, 18, 20], + workers: 2, + }, + ]; + let pools = SpmdDecodePools::build_with_schedule(&shards, DecodeSchedule::Fixed, true); + assert_eq!(pools.dispatcher_cpu(), Some(20)); + pools.shutdown(); + } + #[test] fn decode_schedule_parses_env_values() { assert_eq!(decode_schedule_from_raw(None), DecodeSchedule::Fixed); diff --git a/docs/benchmarks/2026-08-24-acc0-dispatcher-placement.md b/docs/benchmarks/2026-08-24-acc0-dispatcher-placement.md new file mode 100644 index 0000000000..3098c60eff --- /dev/null +++ b/docs/benchmarks/2026-08-24-acc0-dispatcher-placement.md @@ -0,0 +1,272 @@ +# The decode pool reserves a CPU for its dispatcher and never binds it — an opt-in pin, REJECTED by its own bar + +**Date:** 2026-08-24 +**Harness:** `crates/onnx-runtime-ep-cpu/benches/acc0_w16_blocktime_ab.py` (reused +unchanged), `crates/onnx-runtime-ep-cpu/benches/acc0_w16_dispersion.py` (new) +**Instrument:** `dispatcher …` row emitted by `benches/int4_decode_loop_ab.rs` +**Unblocks:** nothing yet — see the verdict. +**Follows:** [2026-08-23-acc0-width-16-worker-attribution.md](2026-08-23-acc0-width-16-worker-attribution.md) + +## Verdict first + +The decode pool leaves one CPU empty for its dispatcher and then never puts the +dispatcher on it. Binding it there (`ONNX_GENAI_CPU_DECODE_DISPATCHER_PIN=1`) is +faster in **15 of 15** trusted launches on current main — and the median gain is +**9.5%**, under the **10%** its pre-registered rule requires. **REJECT.** A +second, separately pre-registered rule about measurement dispersion **failed its +own self-test** and returned REPORT NOTHING. An earlier 6-launch run of the same +knob returned ACCEPT at 19.1%; it **does not replicate** at n=15 on the current +tree, and this record supersedes it. + +The knob merges **off**, as apparatus and as a recorded negative. What is worth +keeping is not the number — it is that the reservation exists without a binding, +that the knob demonstrably takes the reserved CPU, and that three separate ways +of measuring "where is the dispatcher" gave confident wrong answers first. + +## Why this study exists + +The worker-attribution study ended by naming the **measurement**, not the +kernel, as the binding constraint at width 16. A steal-tiles change measured +**+23%** with exactly the mechanism it predicted (`sys_frac` 0.280 → 0.192) and +no width-8 regression, and was nonetheless **REJECTED**, because the A/A null in +the same run — two arms that differ in nothing at all — was **±21.5%**. No +improvement of a realistic size can clear a pre-registered bar against a null +that wide. Until the null is understood, width-16 work cannot ship. + +This study is about the null. + +## The null is not width-16-specific. The *kind* of slow arm is. + +First, from archived data only, at zero host cost — pooling +`bb/steal_ab.json` and `bb/blocktime_ab.json`, 20 launches, 120 arms: + +| | width 8 | width 16 | +|---|---:|---:| +| mean \|aa − 1\| (steal run) | 18.7% | 21.6% | +| mean \|aa − 1\| (blocktime run) | 16.0% | 10.9% | +| worst single A/A deviation | **+90%, +105%** | +49% | + +**This corrected my own earlier localisation.** I had been calling this a +width-16 instability; it is as bad at width 8, and the two worst single +deviations in the whole archive are at width 8. What actually differs is the +*shape* of a slow arm. Splitting arms slower than 1.3× the launch's best by +execution position and by intra-arm rep spread: + +| | width 8 | width 16 | +|---|---|---| +| slow arms | 5 / 45 | **15 / 45** | +| position | 1 and 2 only, never 3 | evenly 5 / 5 / 5 | +| intra-arm spread | **≥23%** every one (median 6.9%) | tight — one ran 1.3× slow with 0.5% internal spread | + +Width 8's slow arms are warmup-shaped, position-dependent, and **self-detectable** +— the rep spread announces them. Width 16's are none of those things. An arm +that is uniformly slow for every rep, with a tight internal spread, is not being +disturbed: it is in a **different state for its whole life**. That is a +per-process property, and it points at what the process got at startup rather +than at what happened to it during the run. + +## What the process gets at startup + +- **NUMA is not it.** `numactl -H` reports **one** node, CPUs 0–31. First-touch + placement cannot vary between launches when there is nowhere else to place. + One command, hypothesis closed. +- **But the cache topology is not flat.** L3 is 64 MiB in **two** instances: + CPUs 0–15 and CPUs 16–31. The pool's `node_worker_counts = [8, 8]` maps onto + these two L3 complexes — not onto NUMA nodes, of which there is one. +- **The workers are pinned, and it is not `decode_affinity`.** On a single-node + host `decide_affinity` returns `Off`, yet every worker's + `/proc//status:Cpus_allowed_list` is a single CPU. The pinning comes from + the CPU-decode-budget path, which also confines the whole process to `w` CPUs. +- **At width 16 one core is deliberately left empty.** Workers 0–14 take even + CPUs 0–28, one each. **CPU 30 is free.** `DISPATCHER_RESERVED_CPUS = 1`, and + `reserve_single_group_headroom` caps workers at `core_count − 1` inside the + physical-core budget. The in-tree justification is a measured **1.57×** + (16 workers at 4.41 ms/token vs 15 workers at 2.81 ms/token): a dispatcher + sharing a core turns that core's worker into the straggler everyone waits for. + +**And the dispatcher was never put on it.** The reservation frees the core and +then nothing binds anything to it. The dispatcher is left to the scheduler, +which may place it on the free core, or on a worker's core, and may move it. +Where it lands is decided once per process, early, by the scheduler — which is +exactly the signature the archived data pointed at. + +## Three instrument failures, all in the same 20 lines + +I record these because two of them produced *confident wrong answers*, not +noise, and the class is general. + +**1. Reading the wrong thread inverted the sign.** The first reporter called +`sched_getcpu()` on the *reporting* thread. Pin off, it read CPU 30; pin on, it +read 18. That is exactly backwards, and it is not a coincidence: the reporter is +idle, so with the pin off the scheduler parks it on the one core nobody is +using — CPU 30 — and with the pin on the dispatcher has **evicted** it from +there. A wrong-thread reading does not add variance to a placement measurement; +it can report the negative of the truth with a straight face. + +**2. The dispatcher is transient.** The second version read the dispatcher's own +`/proc/self/task//stat` (field index verified) and returned `none`. The +dispatcher is neither the process main thread nor the pool's builder, and it has +usually **exited by the time a bench can report**. Any question of the form +"where is the dispatcher" has to be answered from *inside* the dispatch path. + +**3. A process dispatches from more than one thread over its life.** A +migration counter built from consecutive samples reported 2–7 moves on **pinned** +runs, which is impossible. The samples were coming from different dispatching +threads, each of which had correctly taken the reserved CPU. Fixed by recording +the first dispatching thread's tid and sampling **only** that thread — with the +baseline sample taken *after* the bind, so a pinned dispatcher reads exactly +zero. The *pin* behaviour was deliberately left unchanged (every dispatching +thread still takes the reserved CPU) so that the A/B already run was not +invalidated by a diagnostics fix. + +The earlier "dispatcher/worker CPU collision was tested and excluded (one +partial match in four launches)" claim in the ledger came from a probe that +sampled `/proc//stat` — the **main thread**. That claim rested on failure +mode 1 and every statement derived from it was removed from the code before +commit. It is not re-asserted here in either direction. + +## The intervention + +`ONNX_GENAI_CPU_DECODE_DISPATCHER_PIN=1` binds the dispatching thread to the CPU +the reservation already freed (`shards.last().cpus[shards.last().workers]` — the +reserved core belongs to the **last** node, because that is the shard +`node_worker_counts` adds the dispatcher to). Default **off**. One `sched_setaffinity` +per dispatching thread, from a thread-local one-shot inside the single `fn dispatch` +funnel. + +## Pre-registered rules + +Two, both written down before the first measurement. + +1. **Throughput** — the existing validated single-knob A/B, reused byte-identical + and pointed at the new knob via `--env-name/--control/--test`: + median ratio ≥ **1.10**, sign consistency ≥ **80%**, effect > **3×** the A/A + half-width, no width-8 regression below 0.95, ≥ 6 trusted launches. +2. **Dispersion** — a *new* file with its own rule, rather than an edit to the + validated instrument: D = (p90 − p10) / median over per-launch throughput must + fall by ≥ **2×**. Replay-only, self-tested, and it refuses to score a run whose + test arm does not report `PIN-TOOK`. + +## Result: both rules say no + +Current main `7e274a4e2`, 16 launches, **15 trusted**, qwen / acc0 / block 32 / +384 tokens / 3 reps, arms interleaved with the order rotated per launch and the +A/A taken in the same launch as the effect it has to clear. + +| launch | peak | ratio w16 | A/A w16 | sys_frac control | sys_frac test | ratio w8 | +|---:|---:|---:|---:|---:|---:|---:| +| 0 | 19 | 1.0117 | 0.9454 | 0.165 | 0.170 | 1.0685 | +| 1 | 23 | 1.6715 | 1.0164 | 0.333 | 0.192 | 0.9895 | +| 2 | 22 | 1.3442 | 0.9954 | 0.232 | 0.148 | 0.9993 | +| 3 | 22 | 1.0174 | 1.0157 | 0.193 | 0.188 | 0.9931 | +| 4 | 22 | 1.1980 | 1.1731 | 0.201 | 0.164 | 1.0510 | +| 5 | 22 | 1.1572 | 1.0181 | 0.214 | 0.144 | 1.1607 | +| 6 | 19 | 1.0401 | 0.6822 | 0.140 | 0.154 | 1.0248 | +| 7 | 23 | 1.0008 | 1.0028 | 0.150 | 0.156 | 0.6670 | +| 8 | 24 | 1.0953 | 0.9299 | 0.192 | 0.170 | 1.0048 | +| 9 | 23 | 1.0704 | 1.0172 | 0.198 | 0.168 | 1.0440 | +| 10 | 22 | 1.0299 | 1.0245 | 0.150 | 0.145 | 1.6965 | +| 11 | 22 | 1.5696 | 1.0255 | 0.276 | 0.170 | 1.5189 | +| 12 | 23 | 1.0863 | 0.8645 | 0.191 | 0.214 | 1.9997 | +| 13 | 23 | 1.1811 | 1.0354 | 0.222 | 0.173 | 1.0865 | +| 14 | 25 | 1.2158 | 0.7295 | 0.223 | 0.170 | 1.0258 | +| 15 | 75 | *(discarded — runnable peak 75)* | | | | | + +**THROUGHPUT: REJECT.** Median ratio **1.0953**, below the pre-registered +**1.10**. The other two conditions passed — sign consistency is **100%** (the +pinned arm was faster in **15 of 15** trusted launches, against a bar of 80%), +and the effect **+0.0953** clears 3× the A/A half-width, **0.0765**. The +composite rule is nonetheless REJECT, and REJECT is the verdict. The bar was +written down before the first measurement and is not being moved now that a +result has landed 0.005 underneath it. + +**MECHANISM: UNPROVEN.** `sys_frac` falls 0.198 → 0.170, but the shift holds in +only 73% of launches, under the 80% the same rule requires. + +**REGRESSION at width 8: none.** Ratio 1.0440, down-sign 27%. + +**DISPERSION: REPORT NOTHING.** The scorer's own self-test failed: +|D(A/A) − D(A/A's reference arm)| = **0.1432**, over the allowed 0.5 × 0.2591 = +0.1296. Two arms of identical configuration produced dispersion estimates that +differ by more than the estimator's tolerance, so the estimator cannot reproduce +itself at n=15 and is not entitled to compare anything. D(control) = 0.3610 and +D(test) = 0.0780 are recorded here as **unscored observations only** — they are +what the rule refused to certify, not a result. + +### The earlier ACCEPT does not replicate, and this supersedes it + +An earlier run of the same knob against the same throughput rule returned +**ACCEPT** — ratio 1.1910, effect 0.1910 vs a 3× A/A half-width of 0.1477 — and +the dispersion rule returned **PIN-STABILISES**, 0.2781 → 0.0416. That run had +**6** trusted launches and was taken on `d5e585d2a`, before #1868 landed. The +run above has 15 and is on the tree this branch actually merges into. Two things +moved: + +- **n.** Six launches put the median 19% up; fifteen put it 9.5% up with the + same sign in every launch. The six-launch median was optimistic, which is the + ordinary behaviour of a median over a heavy-tailed sample, not a defect in + either run. +- **The baseline.** #1868 fixed the spin deadline at two yield sites, and + control `sys_frac` at width 16 fell from 0.257 to 0.198 between the two runs. + Some of what the pin was recovering has already been recovered upstream. + +The larger, on-tree run wins. **The pin does not clear its bar on current main.** + +## What is established, and what is not + +**Established, and not by timing:** + +- The pool reserves a CPU for the dispatcher (`DISPATCHER_RESERVED_CPUS = 1`, + justified in-tree by a measured 1.57×) and **binds nothing to it**. At width + 16 on this host that is CPU 30, and with the knob off the dispatcher is an + ordinary unpinned thread. +- The knob does what it says. Direct measurement, 4 launches per arm, + interleaved: pinned reports `observed_cpu=30` and **0 migrations** in every + launch; unpinned reports 1, 1, 1 and 0, and in one launch was last seen on + **CPU 2** — a worker's core — rather than the reserved one. +- Native's width-16 A/A instability is **not** width-16-specific in magnitude, + and the width-16 slow arms are internally consistent, which makes them a + per-process state rather than a disturbance. + +**Not established:** + +- **That the pin helps.** Directionally consistent 15/15 and below its bar. +- **Why it would.** The migration counter samples once per 1024 dispatches — + roughly 150 samples per launch — and sees at most one change per unpinned + launch. That rate is far too low to make steady-state migration the + explanation, so the counter's honest reading is that **migration is not the + mechanism**, not that it is. The leading remaining candidate is *wakeup* + placement rather than residence: the dispatcher parks and is woken hundreds of + times per token, and a single-CPU affinity mask lets the kernel skip the + idle-sibling search on each wake. `sys_frac` falling is consistent with that + and is not evidence for it at 73% sign. **Nothing here should be cited as a + mechanism.** +- **That the A/A null is fixed.** It is the thing that blocked the +23% + steal-tiles candidate, and the dispersion rule declined to certify any change + in it. + +## Why the knob ships off, and what flipping it would need + +Beyond it not having cleared its bar: the dispatcher is the **session thread**, +and `sched_setaffinity` is not scoped to a decode. A thread pinned during decode +keeps that mask afterwards, so a subsequent **prefill** on the same thread would +run one CPU wide. This harness measures decode only and cannot see that. Any +proposal to flip the default needs prefill in the matrix, not more decode +launches. + + +## Reproduce + +```bash +cargo build --release -p onnx-runtime-ep-cpu --benches +BIN=$(ls target/release/deps/int4_decode_loop_ab-* | grep -v '\.d$' | head -1) +./scripts/hostlock.sh run --wait --gate 8 -- \ + python3 crates/onnx-runtime-ep-cpu/benches/acc0_w16_blocktime_ab.py \ + --binary "$BIN" --env-name ONNX_GENAI_CPU_DECODE_DISPATCHER_PIN \ + --control 0 --test 1 --launches 16 --out pin_ab.json +python3 crates/onnx-runtime-ep-cpu/benches/acc0_w16_dispersion.py --replay pin_ab.json +``` + +Non-vacuity is checked by the harness itself: the `dispatcher …` row must read +`PIN-OFF` on the control arm and `PIN-TOOK` on the test arm, and the dispersion +scorer aborts if it does not. diff --git a/docs/performance/CPU_MATMUL_ASSIGNMENT.md b/docs/performance/CPU_MATMUL_ASSIGNMENT.md index 45998d688e..c2c7abc22e 100644 --- a/docs/performance/CPU_MATMUL_ASSIGNMENT.md +++ b/docs/performance/CPU_MATMUL_ASSIGNMENT.md @@ -2004,10 +2004,39 @@ time. That figure is now stale; the re-measurement replaces it. > and not proposed, because the t=16 A/A null in the same run is **+-21.5%**. > **The binding constraint at t=16 is now the measurement, not the kernel:** > until the A/A instability is understood no improvement of realistic size can -> clear a pre-registered bar there. Full records: +> clear a pre-registered bar there. +> +> **The dispatcher/worker collision line above is withdrawn as unevidenced.** +> It rested on a probe that sampled `/proc//stat` -- the process **main +> thread**, which is not the dispatcher. The dispatcher is a transient thread +> that is usually gone before a bench can report, so its placement can only be +> read from inside the dispatch path; a first attempt from the *reporting* +> thread returned the exactly **inverted** answer, because the reporter is idle +> and the scheduler parks it on the very core the reserve freed. Collision is +> now neither asserted nor excluded. +> +> **What is established is structural: the pool reserves a CPU for the +> dispatcher and binds nothing to it.** `DISPATCHER_RESERVED_CPUS = 1` and +> `reserve_single_group_headroom` keep one allowed CPU clear of workers -- CPU +> 30 at t=16 on this host, justified in-tree by a measured 1.57x -- and the +> dispatcher is then left to the scheduler. Direct measurement confirms the gap +> is real: unpinned, the dispatcher was last seen on a **worker's** core in one +> launch of four. `ONNX_GENAI_CPU_DECODE_DISPATCHER_PIN=1` closes it and +> **fails its own bar**: 15 of 15 trusted launches faster, median **1.0953** +> against a pre-registered 1.10, with no t=8 regression; the companion +> dispersion rule **failed its self-test** and certified nothing. An earlier +> 6-launch run scored 1.1910/ACCEPT and **did not replicate** -- partly small-n, +> partly because #1868's spin-deadline fix already took control `sys_frac` at +> t=16 from 0.257 to 0.198. The mechanism is **unproven**: the unpinned +> dispatcher migrates at most once per launch, so migration is not it. The knob +> ships **off**, and would need prefill in the matrix before it could ship on -- +> the dispatcher is the session thread and keeps its affinity after decode ends. +> **The A/A null therefore remains open, and the +23% steal-tiles candidate +> remains blocked behind it.** Full records: > [`docs/benchmarks/2026-08-23-acc0-gap-at-width-16.md`](../benchmarks/2026-08-23-acc0-gap-at-width-16.md), > [`docs/benchmarks/2026-08-23-acc0-width-16-cpu-attribution.md`](../benchmarks/2026-08-23-acc0-width-16-cpu-attribution.md), -> [`docs/benchmarks/2026-08-23-acc0-width-16-worker-attribution.md`](../benchmarks/2026-08-23-acc0-width-16-worker-attribution.md). +> [`docs/benchmarks/2026-08-23-acc0-width-16-worker-attribution.md`](../benchmarks/2026-08-23-acc0-width-16-worker-attribution.md), +> [`docs/benchmarks/2026-08-24-acc0-dispatcher-placement.md`](../benchmarks/2026-08-24-acc0-dispatcher-placement.md). > > The old figure was **not mislabelled — it was a correct measurement of a tree > that no longer exists.** An earlier draft argued this from the ORT arm alone