Repository navigation
docs(perf): record the acc0 route, the SMT placement finding and the f16 layout divergence - #1701
Conversation
…f16 layout divergence Three entries in the shared CPU matmul ledger for work that landed or was closed today, kept together because two of them are the same lesson. **§23 — the acc0 route (`99f105d52`).** #1104's register-blocked int4 kernel shipped default-off "until the win is measured" and the measurement never happened, so `accuracy_level = 0` — the production default — took the per-column path for the kernel's whole life. Route counters proved it (`nblock = 0`). Records the corrected attribution too: the first one was inflated ~3x because the probe's `fetch_add` was in the timed path. With a clean instrument the hreduce removal is worth 1.02x, i.e. nothing, and the whole 1.48x is four-column activation reuse. The numerics regress up to 3.70x relatively; that is disclosed, not buried. **§24 — the t=8 "wash" (#1680).** The premise does not reproduce: the win is flat through pool width 8 and collapses at 12+. Root cause is `decode_spmd.rs::node_shards` pinning worker *i* to `allowed_cpus()[i]` in logical order, so 16 workers land on 8 physical cores. Bandwidth and task grain were tried first and discarded. No kernel change — handed to the runtime owner with the measurement. **§25 — the f16/bf16 layout divergence (`2e1cfb67c`).** Same math, same bytes, three prices; `[K,N]` crosses a page every `p`. Software prefetch was tried first and is a negative result. Accuracy moves the same way (`nk` is 2.7-9.3x better), so there was no trade to weigh. Also records the memory-plan coupling that could have gone badly — the #1056 predictor was Apple-only for `MatMul`, so the transpose would have been invisible to the plan on x86. §23 and §25 both contain the same error, caught in different disguises: an instrument that changes what it measures, and a numerics test built from exactly-representable operands that could not see reassociation at all. Both are written up as such. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ll records Adversarial review found two numbers stated more strongly than the evidence beneath them and one convention break. All three are fixed. **§23's headline numerics multiplier was wrong.** 2.422e-5 / 6.739e-6 is **3.59x**, not 3.70x, and the overstatement was repeated in the summary sentence. This is the one number in that section that has to be exact — it is the cost side of "1.48x speed for X worse accuracy" — so it is corrected in both places and pinned to "the worst cell measured" rather than an unqualified "up to". **§24 quoted a bandwidth figure this file already refuted.** My all-thread sweep read 83 GB/s; §22 measured this host at 31-36 GB/s within a CCX and ~56.6 GB/s across both, and §22's numbers are the ones taken with the access pattern decode actually uses. Against those, a 41 GB/s draw is *above* the within-CCX ceiling — the opposite of "not bandwidth-bound". §24 cannot dismiss bandwidth on a number §22 refutes. Rewritten to state the conflict and rule bandwidth out on the placement A/B instead, which holds shapes, bytes, thread count and binary constant and varies only which CPUs the workers sit on. The conclusion is unchanged; its support is now sound. **Every earlier section links a `docs/benchmarks/` full record; these three did not.** That left the numbers a skeptical reader would want unbacked — the kernel A/B range in particular was not reconstructable from anything quoted. Adds the three records with the full matrices, the controls, the discarded hypotheses and the negatives, and links them. Also: dropped a "cf. §18" that claimed a parallel §18 does not support (§18's probe was real kernel overhead that slowed the shipping kernel; §23's was measurement-only and corrupted its own baseline) — the distinction is now spelled out instead; fixed an ordinal that said "fourth" while citing four priors; replaced a "model-shaped rows are 1.2-1.6x" generalisation that its own table contradicts at attn_out with the large-row claim the data supports; and made one "an Opus review" match the file's "adversarial review" voice. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Adversarial review (Opus) returned REQUEST CHANGES with three Medium and three Low findings. All six are fixed in Two of the Mediums were numbers stated more strongly than the evidence sitting directly beneath them, which in this file is the whole failure mode:
The third Medium was a convention break with real consequences: every earlier section links a
Lows, all taken: dropped a The reviewer explicitly checked and confirmed clean: every ratio in §23's table, the 1.45x group-1→group-4 decomposition, all nine §25 production speedups, §24's internal consistency against §22's SMT model, and that the four disclosures I was most worried about softening (§23's shipped numerics regression, §25's 272 MB / worst-at-1.21x Verification on the merged base: Follow-up split out rather than folded in: #1702 ( |
|
Merging under Justin's direct-merge directive, with disclosure. The one red check is inherited from Change scope: The failure: Attribution. Byte-for-byte identical on main's own run at Local verification on the merged base (
Adversarial review returned REQUEST CHANGES, all six findings fixed in No test, kernel or build input is touched by this PR, so there is no code failure to investigate before merge — only main's, which is now tracked. |
Three entries in
docs/performance/CPU_MATMUL_ASSIGNMENT.mdfor work that landed or was closed today. Documentation only — no code.99f105d52, perf(cpu): enable the register-blocked int4 decode kernel at accuracy_level=0 (1.17-1.69x) #1679). Absorb MLAS SQNBit decode locality into borrowed int4 path (no resident copy) #1104's register-blocked int4 kernel shipped default-off "until the win is measured" and the measurement never happened, soaccuracy_level = 0took the per-column path for the kernel's whole life; route counters readnblock = 0. Records the corrected attribution: the first one was inflated ~3x because the probe'sfetch_addsat in the timed path. Clean, the hreduce removal is worth 1.02x — nothing — and the entire 1.48x is four-column activation reuse. The numerics regress up to 3.70x relatively; disclosed rather than buried.decode_spmd.rs::node_shardspinning worker i toallowed_cpus()[i]in logical order, so 16 workers land on 8 physical cores. Bandwidth (83 GB/s available vs 41 GB/s drawn) and task grain were tried first and discarded. No kernel change — handed to the runtime owner with the measurement.2e1cfb67c, perf(cpu): route f16/bf16 decode through the [N,K] GEMV (1.64-2.89x), closing #1381 #1687, closing perf(cpu): give bf16 decode a GEMV, and split decode at the measured weight #1381). Same math, same bytes, three prices;[K,N]crosses a page everyp. Software prefetch was tried first and is a recorded negative. Accuracy moves the same way, so there was no trade to weigh. Also records the memory-plan coupling that could have gone badly: the Every resident weight side-buffer must be in the memory plan before it is allocated #1056 predictor was Apple-only forMatMul, so the transpose would have been invisible to the plan on x86.§23 and §25 contain the same error in two disguises — an instrument that changed what it measured, and a numerics test built from exactly-representable operands (
*0.125) that could not see reassociation at all and reported a confident zero. Both are written up as such, next to §18's version of it.All seven repo policy scripts pass (
check_publish_order,check_profile_table,check_platform_naming,check_dispatch_reachability,check_dispatch_manifest,check_feature_gate_coverage,verify_documented_env_vars).