Skip to content

docs(perf): record the k-split negative result and correct §16's proposed fix - #1618

Merged
justinchuby merged 1 commit into
mainfrom
squad/roy-ksplit-negative
Aug 21, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/roy-ksplit-negative

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 21, 2026 •

Copy link
Copy Markdown
Owner

Docs-only. Closes out the hypothesis §16 named, with the evidence, after an adversarial review corrected both tallies and the scope of the conclusion.

Result

36 cells (6 llama/qwen shapes × t=1,2,4,8,16,32), 3 reps, two prebuilt test binaries alternated, portable drift control 0.999, t=1 rows as an identical-code null control.

That control fails on 1x1024x3072 (−29% in one matrix, +14% in the other, on code that is identical at one thread), so all five of that shape's cells are dropped — its wins and its losses — leaving 25 counted cells:

  • v1 hybrid band × column, serial reduction: 8 wins, 5 losses, 12 neutral, geomean 0.997 — a wash. And it inverts its own prediction: the gain should grow with thread count as the stripe narrows; t=4 is the best column and t=32 the worst.
  • v2 pool-aligned bands + parallel reduction (the Amdahl fix, bands*threads/k ≈ 6%): 3 wins, 13 losses, 9 neutral, geomean 0.846 (3/12/9, 0.904 with one unexplained 0.17x cell also excluded).

A second parallel plan, a scratch allocation and a reduction have to earn their place. 0.997 does not buy them. Code #1617 is closed unmerged.

Kept on the record

Accumulation is wrapping i32, so summing bands in band order is bit-identical to one pass — a k split owes an equality, not a tolerance. Asserted at every thread count; both plan gates mutation-checked.

Scope, stated narrowly

The experiment moved two things at once — the access pattern and the addition of scratch — and the scratch is the best explanation of the v2 result. So the streaming layout was never measured in isolation, and this does not rule out prepacked/reordered B, huge pages, software prefetch, NUMA first-touch, or gating the split to only the sub-page-stripe case.

The "instruction-bound inside L3" reading rests on §16's aspect-ratio sweep, not on this — a confounded A/B corroborates, it does not independently confirm. The planning consequence is a priority call, not a proof that layout work is dead: the instruction budget is the measured term (vpmaddwd at ~0.31 total uops/byte of B, ~0.25 vector-only), so paths cutting bytes and uops per weight together — the packed-nibble int4 kernel at 0.5 B/weight — are the better next spend.

Docs-only: no code, no test or CI gate reads these files.

…'s proposed fix

Section 16 named a k split with private accumulators as the fix for the
m=1 parallel-scaling loss. It was built, it is bit-identical to the
column split, and it does not pay for itself: over the 25 cells whose
null control holds, the hybrid plan is a wash (geomean 0.997) that
inverts its own predicted trend, and removing its Amdahl term makes it
worse (geomean 0.846).

Scope is stated narrowly. The experiment moved the access pattern and
added scratch at the same time, and the scratch is the best explanation
of the second result, so the streaming layout was never measured in
isolation -- prepacked B, huge pages, prefetch and NUMA first-touch are
untouched. The instruction-bound reading rests on the aspect-ratio sweep
in section 16, not on this.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the squad/roy-ksplit-negative branch from 747af24 to 02cf10a Compare August 21, 2026 00:16
@justinchuby
justinchuby marked this pull request as ready for review August 21, 2026 00:17
@justinchuby
justinchuby merged commit e8ee477 into main Aug 21, 2026
10 checks passed
@justinchuby
justinchuby deleted the squad/roy-ksplit-negative branch August 21, 2026 00:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant