Repository navigation
docs(perf): record the k-split negative result and correct §16's proposed fix - #1618
Merged
Merged
Conversation
…'s proposed fix Section 16 named a k split with private accumulators as the fix for the m=1 parallel-scaling loss. It was built, it is bit-identical to the column split, and it does not pay for itself: over the 25 cells whose null control holds, the hybrid plan is a wash (geomean 0.997) that inverts its own predicted trend, and removing its Amdahl term makes it worse (geomean 0.846). Scope is stated narrowly. The experiment moved the access pattern and added scratch at the same time, and the scratch is the best explanation of the second result, so the streaming layout was never measured in isolation -- prepacked B, huge pages, prefetch and NUMA first-touch are untouched. The instruction-bound reading rests on the aspect-ratio sweep in section 16, not on this. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
force-pushed
the
squad/roy-ksplit-negative
branch
from
August 21, 2026 00:16
747af24 to
02cf10a
Compare
justinchuby
marked this pull request as ready for review
August 21, 2026 00:17
This was referenced Aug 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Docs-only. Closes out the hypothesis §16 named, with the evidence, after an adversarial review corrected both tallies and the scope of the conclusion.
Result
36 cells (6 llama/qwen shapes × t=1,2,4,8,16,32), 3 reps, two prebuilt test binaries alternated,
portabledrift control 0.999,t=1rows as an identical-code null control.That control fails on
1x1024x3072(−29% in one matrix, +14% in the other, on code that is identical at one thread), so all five of that shape's cells are dropped — its wins and its losses — leaving 25 counted cells:t=4is the best column andt=32the worst.bands*threads/k≈ 6%): 3 wins, 13 losses, 9 neutral, geomean 0.846 (3/12/9, 0.904 with one unexplained 0.17x cell also excluded).A second parallel plan, a scratch allocation and a reduction have to earn their place. 0.997 does not buy them. Code #1617 is closed unmerged.
Kept on the record
Accumulation is wrapping i32, so summing bands in band order is bit-identical to one pass — a
ksplit owes an equality, not a tolerance. Asserted at every thread count; both plan gates mutation-checked.Scope, stated narrowly
The experiment moved two things at once — the access pattern and the addition of scratch — and the scratch is the best explanation of the v2 result. So the streaming layout was never measured in isolation, and this does not rule out prepacked/reordered
B, huge pages, software prefetch, NUMA first-touch, or gating the split to only the sub-page-stripe case.The "instruction-bound inside L3" reading rests on §16's aspect-ratio sweep, not on this — a confounded A/B corroborates, it does not independently confirm. The planning consequence is a priority call, not a proof that layout work is dead: the instruction budget is the measured term (
vpmaddwdat ~0.31 total uops/byte ofB, ~0.25 vector-only), so paths cutting bytes and uops per weight together — the packed-nibble int4 kernel at 0.5 B/weight — are the better next spend.Docs-only: no code, no test or CI gate reads these files.