Skip to content

The diffstat is read first, so being wrong about a record's size is worse than being wrong about its content - #10075

Merged
gunbai-bot[bot] merged 1 commit into
mainfrom
session/stern-otter-633
Sep 2, 2026
Merged

gunbai-bot[bot] merged 1 commit into
mainfrom
session/stern-otter-633

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

What this files

One recurring_failure_mode row: salience_instrument_blind_to_the_record_it_sizes.

A change's size in the diff is read to decide how much attention it deserves. For one class of artifact that number is wrong by orders of magnitude, so a substantive edit is allocated a trivial edit's scrutiny.

The specimen, measured

$ git diff --stat -- dag/gunbc/rung_drop.dag
 1 file changed, 1 insertion(+), 1 deletion(-)

$ git diff -- dag/gunbc/rung_drop.dag | grep '^+' | wc -c
 15286

Every RungDrop in this corpus is one very long line, so a 15 kB addition renders exactly like a typo fix.

On #10044, three consecutive independent approvals (reviews 58738, 58749, and one prior) described that hunk as a small wording tweak. A fourth, earlier review (58702) attributed it to direct_call_arg_seam_v2_exemption — a row that diff does not touch — because git's @@ header names the declaration preceding the hunk, and a single-line record is therefore labelled with its neighbour. One representation, two distinct review failures.

Why it is worse than an inaccurate number

The diffstat is a salience instrument: it is read first, to decide where to look. Nobody re-checks a figure they have already used to conclude the thing is not worth checking. The error is self-concealing in a way a wrong content claim is not — a wrong claim gets disputed, a wrong salience signal gets obeyed silently.

The approvals were not wrong on what they read. Each verdict is defensible over the prose rows in the same diff, which render normally. What the count concealed is that #10044 had three approvals and no review coverage of the half carrying the receipt — the gap was closed only when the landing manager read it with --word-diff before merging. A review tally is a claim about attention, and this instrument redirects attention before any reviewer forms a judgment.

Rung and trigger

  • Rung found at: silent wrongness. The misallocation leaves no trace, produces a green, and is indistinguishable from a reviewer who looked and found nothing.
  • Ceiling: mechanically preventable — a record's diff size can be made to track its content size. Not structurally impossible, since nothing stops a future record being authored as one long line.
  • Trigger — a representation, deliberately not advice: the long-line record broken across lines so hunks are proportional to the change, or a review surface that reads the generated projection (docs/design-ledgers.md carries the same content as prose and diffs legibly). "Look harder" cannot be discharged and is what a class gets when nobody wants to pay for the fix.
  • Recognition rule: when a review's description of a hunk is smaller than the hunk, check whether the artifact's line structure and the change's content structure agree. Where one record is one line, every size signal derived from lines — diffstat, hunk header, line-count budget, review-effort heuristic — is answering about the file's shape rather than the change's.

The class was found by reviewing this session's own landed work, not reported from outside it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

…eing wrong about content

Three consecutive approvals on #10044 described its rung_drop.dag hunk as a small wording
tweak; an earlier one attributed it to a row the diff never touched. Measured:

  git diff --stat -- dag/gunbc/rung_drop.dag  ->  1 insertion(+), 1 deletion(-)
  bytes actually added                        ->  15,286

Every RungDrop is one very long line, so a 15 kB receipt renders exactly like a typo fix. The
same shape produced the misattribution: git's @@ header names the declaration PRECEDING the
hunk, so a single-line record is labelled with its neighbour.

WHAT MAKES THIS WORSE THAN AN INACCURATE NUMBER: the diffstat is a SALIENCE instrument. It is
read FIRST, to decide where to look. Nobody re-checks a figure they have already used to
conclude the thing is not worth checking, so the error is self-concealing in a way a wrong
content claim is not.

The approvals were not wrong on what they read — each verdict is defensible over the prose rows
in the same diff, which render normally. What the count concealed is that the PR had three
approvals and NO review coverage of the half carrying the receipt. A review tally is a claim
about attention, and this instrument redirects attention before any reviewer forms a judgment.

TRIGGER IS A REPRESENTATION, NOT ADVICE: a record whose diff size tracks its content size —
the long-line record broken across lines, or a review surface reading the generated projection
(docs/design-ledgers.md renders the same content as prose and diffs legibly). "Look harder"
cannot be discharged and is what a class gets when nobody wants to pay for the fix.

Specimen n=3 in one PR, with a second failure mode (misattribution) from one cause. The class
was found by review of this session's own work rather than reported from outside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants