Repository navigation
docs(bench): phase 18 — the residual t=32 loss is mostly the co-resident ORT pool - #1374
Conversation
…ent ORT pool Measurement-only phase, no code change. The t=32 inversion left open by phase 16 is largely the harness, not the scheduler. Timing the native arm with and without a co-resident ORT session isolates it: gemm_nbits_llama3_8b_qkv_t8 is 6.90 ms alone and 28.71 ms paired at t=32, reproduced to three significant figures across two sessions. Measured alone, that cell is flat across t=8/16/32 -- the inversion only exists in the paired numbers. The tax is asymmetric (ORT pays 1.2-1.3x for ours), and it hits the dense f32 control hardest, so it is contention rather than anything about int4. Lane width is ruled out first: t=16 and t=32 both run 16 lanes and still differ ~3x. Adds the resulting rule to the method alongside the null arm, corrects the overstated native-vs-ORT ratios for long cells at wide thread counts, records an affinity probe that failed its own noise check, and downgrades the with_decode_pool hypothesis from phase 16. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Merging with admin, consistent with the rest of this campaign: the Actions queue on this repo is saturated (every run sits This PR is baseline-equivalent by construction. The diff is two markdown files and zero lines of Rust: No crate source, manifest, or feature gate is touched, so The
|
Documentation only — no code change, no behaviour change.
§36.8 left the t=32 residual open and called it "a scheduler signature". It is —
but it is mostly not our scheduler. This phase spent its budget on measurement.
What it found
bench_generic --native-only/--ort-onlybuild only the arm being timed, sorunning the same binary with and without a co-resident ORT session isolates
contention from arithmetic. At t=32:
gemm_nbits_llama3_8b_qkv_t8gemm_dense_tall_128x4096(f32 dense)gemm_nbits_llama3_8b_mlp_t8The first row reproduced to three significant figures across two independent
sessions an hour apart (28.698 / 28.713 paired, 6.868 / 6.900 alone) — four
times outside §36.3's measured noise band.
Measured alone,
llama3_8b_qkv_t8is 7.15 / 6.22 / 6.87 ms at t=8/16/32:flat, no inversion. The inversion only exists in the paired numbers.
Why
The tax is asymmetric — ORT pays only 1.2–1.3× for our co-residency, we pay up
to 4.2× for its. ORT's intra-op pool spin-waits long after its last op; our task
runtime spins briefly and parks. A pool that parks quickly is invisible to its
neighbours; a pool that spins is not. It hits the dense f32 control hardest,
so it is contention, not anything about int4.
Lane width is ruled out first (§38.1): t=16 and t=32 both run 16 lanes and still
differ ~3×, and the default width is at or near optimal on 3 of 4 cells.
What it changes
solo arms. The paired harness stays correct for parity, short cells, and
base-vs-new comparisons of our own binaries.
llama3_8b_qkv_t8is 10× behind ORT, not 41×. Still a real gap, still a kernel problem.
same tax.
attempt knows it has been run once on a loaded box for nothing.
with_decode_poolhypothesis: it early-returns inline whenwith_decode_pool_scopeis active, so theinstallis per-pass, not per-node.Validation
Docs only. Full local
Rust qualityproxy run anyway:cargo fmt --all --checkclean, all 8 guard scripts pass,
workspace_test_packages verifyOK.