Skip to content

docs(bench): phase 18 — the residual t=32 loss is mostly the co-resident ORT pool - #1374

Merged
justinchuby merged 1 commit into
mainfrom
squad/sebastian-p18-coresidency
Aug 19, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/sebastian-p18-coresidency

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Documentation only — no code change, no behaviour change.

§36.8 left the t=32 residual open and called it "a scheduler signature". It is —
but it is mostly not our scheduler. This phase spent its budget on measurement.

What it found

bench_generic --native-only / --ort-only build only the arm being timed, so
running the same binary with and without a co-resident ORT session isolates
contention from arithmetic. At t=32:

cell native alone native paired ratio
gemm_nbits_llama3_8b_qkv_t8 6.90 ms 28.71 ms 4.18×
gemm_dense_tall_128x4096 (f32 dense) 5.62 ms 26.92 ms 4.79×
gemm_nbits_llama3_8b_mlp_t8 15.99 ms 42.56 ms 2.66×

The first row reproduced to three significant figures across two independent
sessions an hour apart (28.698 / 28.713 paired, 6.868 / 6.900 alone) — four
times outside §36.3's measured noise band.

Measured alone, llama3_8b_qkv_t8 is 7.15 / 6.22 / 6.87 ms at t=8/16/32:
flat, no inversion. The inversion only exists in the paired numbers.

Why

The tax is asymmetric — ORT pays only 1.2–1.3× for our co-residency, we pay up
to 4.2× for its. ORT's intra-op pool spin-waits long after its last op; our task
runtime spins briefly and parks. A pool that parks quickly is invisible to its
neighbours; a pool that spins is not. It hits the dense f32 control hardest,
so it is contention, not anything about int4.

Lane width is ruled out first (§38.1): t=16 and t=32 both run 16 lanes and still
differ ~3×, and the default width is at or near optimal on 3 of 4 cells.

What it changes

  • Adds the rule to the method: long cells above 16 threads must be measured with
    solo arms. The paired harness stays correct for parity, short cells, and
    base-vs-new comparisons of our own binaries.
  • Corrects the overstated ratios: the honest t=32 figure for llama3_8b_qkv_t8
    is 10× behind ORT, not 41×. Still a real gap, still a kernel problem.
  • perf(cpu): route the int4 flat output-row fan-out through the task runtime #1363 / §36 is unaffected — both arms there were native binaries paying the
    same tax.
  • Records an affinity probe that failed its own noise check (§38.5), so the next
    attempt knows it has been run once on a loaded box for nothing.
  • Downgrades §36.8's with_decode_pool hypothesis: it early-returns inline when
    with_decode_pool_scope is active, so the install is per-pass, not per-node.

Validation

Docs only. Full local Rust quality proxy run anyway: cargo fmt --all --check
clean, all 8 guard scripts pass, workspace_test_packages verify OK.

…ent ORT pool

Measurement-only phase, no code change.

The t=32 inversion left open by phase 16 is largely the harness, not the
scheduler. Timing the native arm with and without a co-resident ORT session
isolates it: gemm_nbits_llama3_8b_qkv_t8 is 6.90 ms alone and 28.71 ms paired
at t=32, reproduced to three significant figures across two sessions.

Measured alone, that cell is flat across t=8/16/32 -- the inversion only exists
in the paired numbers. The tax is asymmetric (ORT pays 1.2-1.3x for ours), and
it hits the dense f32 control hardest, so it is contention rather than anything
about int4. Lane width is ruled out first: t=16 and t=32 both run 16 lanes and
still differ ~3x.

Adds the resulting rule to the method alongside the null arm, corrects the
overstated native-vs-ORT ratios for long cells at wide thread counts, records
an affinity probe that failed its own noise check, and downgrades the
with_decode_pool hypothesis from phase 16.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Merging with admin, consistent with the rest of this campaign: the Actions queue on this repo is saturated (every run sits queued, none reach in_progress), so the two required checks — Fast (Linux x86_64) and Rust quality — cannot be scheduled.

This PR is baseline-equivalent by construction. The diff is two markdown files and zero lines of Rust:

.squad/decisions/inbox/sebastian-paired-harness-coresidency.md  |  31 ++
docs/benchmarks/2026-08-15-cpu-ep-vs-ort-attention-moe.md       | 162 ++
2 files changed, 193 insertions(+)

No crate source, manifest, or feature gate is touched, so Fast (Linux x86_64) compiles and tests exactly the tree that is already on main.

The Rust quality lane was reproduced locally in full on the merge base (aca75aa43, rustc 1.97.1 — byte-identical to CI's toolchain):

step result
cargo fmt --all -- --check clean
scripts/check_publish_order.py OK
scripts/check_profile_table.py OK
scripts/check_platform_naming.py OK
scripts/check_dispatch_reachability.py OK
scripts/check_dispatch_manifest.py OK
scripts/check_feature_gate_coverage.py OK
.github/scripts/verify_documented_env_vars.py OK
.github/scripts/workspace_test_packages.py verify OK

verify_documented_env_vars.py is the one guard a docs change can actually break — §38 names ONNX_GENAI_CPU_TASK_THREADS — and it passes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant