Skip to content

Two error classes from the night's own instruments: a gate that refuses without judging, and a capture that reads clean because it is empty - #10044

Merged
gunbai-bot[bot] merged 9 commits into
mainfrom
session/stern-otter-633
Sep 2, 2026
Merged

gunbai-bot[bot] merged 9 commits into
mainfrom
session/stern-otter-633

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Two recurring failure modes, both discovered by using this repository's own instruments during #9984 and both filed with the receipt that made them decidable rather than anecdotal.

non_verdict_disposition_surfaces_as_refusal

The required floor reports two materially different outcomes — this subject is wrong and I did not finish looking at this subject — through one refusing channel.

The receipt is a same-head pair: 9b00e24f592 run twice with no intervening edit.

first  planned=3486 executed=3486 passed=3404 failed=0 interrupted_before_verdict=4 completed_over_cost_requirement=3   → FloorRefused
rerun  planned=3486 executed=3486 passed=3411 failed=0 interrupted_before_verdict=0 completed_over_cost_requirement=0   → green

Identical planned and executed populations, zero failures in both, and the seven cost-arm rows simply absent the second time. Holding the bytes fixed by construction is what makes this a measurement rather than an inference — a cross-head comparison requires arguing that the intervening commit could not have touched cost accounting, and an argument about what a diff cannot do is exactly what gets overturned.

The dispositions are named non-verdict by the floor itself: INTERRUPTED-BEFORE-VERDICT says the deadline preempted the witness so whether it passes is unknown. So the gate is not reversing a judgment between runs; it is intermittently failing to reach one and rendering that as a refusal.

The harm is not the red. A refusal that names no wrong subject can only be answered by running it again — and a wall discharged by rerunning is not a wall. It teaches every consumer that its refusals are weather.

Trigger names a capability in two conjuncts: non-verdict dispositions as a third outcome distinct from pass and fail, and a cost admission whose budget is a property of the claim rather than of the run's host speed. Either alone leaves the class where it is — without the second, the first merely relabels an intermittent outcome while the population needing resolution changes underneath its own roster.

empty_capture_read_as_clean_result

An instrument refuses on one stream while the reader keeps the other, so the capture is empty and the empty capture is consumed as a finding of nothing.

Specimen: gh run view --job <id> --log > f on an in-progress run writes zero bytes, putting run is still in progress on stderr. A grep for panicked or FAILED over that file then reports no failures for a job already recorded as completed failure. Two adjacent instances in the same call: gh api .../jobs/<id>/logs writes zero bytes without --allow-escape-sequences, and the subcommand takes a job id rather than a run id.

The asymmetry is why it recurs: the failure direction is always benign. An empty capture never manufactures a false alarm, only a false all-clear, so nothing in the reader's experience trains them to check it.

Ceiling is honest about the boundary: mechanically preventable for instruments this repository owns, and mitigatable for third-party command-line tools, whose stream discipline is outside the modeled guarantee.

On the regeneration, stated rather than quietly omitted

DESIGN.md and docs/design-ledgers.md are regenerated by main_wet, which reproduced every other rostered artifact byte-identically — that is the positive control for this projection.

The stage0 --required-regen run FAILED with drift in compiler_tests.rs. That failure is void, not a finding: the candidate it produced is 8 lines from the pre-#9886 committed file and 148 lines from the current one, because this session's binary was built at 03:56 from the stage0 mirror as it stood before #9886 changed 05_emit_rust. A stale seed regenerates a stale world. CI's build lane regenerates with a current binary and is the adjudicator; main's own build lane was green at 7f71ee34094, after #9886 landed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

gunbc-ci-auto-heal and others added 2 commits September 2, 2026 09:45
…es without judging, and a capture that reads clean because it is empty

Both are §4b(1) filings against mechanisms this repository relies on to know whether it is
correct, and each carries the receipt that made it decidable rather than anecdotal.

non_verdict_disposition_surfaces_as_refusal. The required floor reports THIS SUBJECT IS
WRONG and I DID NOT FINISH LOOKING through one refusing channel. Its receipt is a same-head
pair: 9b00e24 run twice with no intervening edit, planned=3486 executed=3486 failed=0
both times, seven INTERRUPTED-BEFORE-VERDICT / COMPLETED-OVER-COST-REQUIREMENT rows present
in the first and absent in the second. Holding the bytes fixed by construction is what makes
it a measurement: a cross-head comparison would have required arguing that the intervening
commit could not have touched cost accounting, and an argument about what a diff cannot do is
exactly what gets overturned. The harm is not the red -- it is that a refusal naming no wrong
subject can only be answered by rerunning, and a wall discharged by rerunning is not a wall.

empty_capture_read_as_clean_result. An instrument refuses on one stream while the reader keeps
the other, so the capture is empty and the empty capture is consumed as a finding of nothing.
Specimen: `gh run view --job <id> --log > f` on an in-progress run writes ZERO BYTES with its
refusal on stderr, so a grep for `panicked` over that file reports no failures for a job that
already failed. The failure direction is always benign, which is why it recurs -- an empty
capture never manufactures a false alarm, only a false all-clear.

Both name a capability as their trigger, not an artifact: a required-gate verdict in which
non-verdict dispositions are a third outcome plus a claim-owned cost admission, and a
result type that cannot let a zero-byte capture inhabit READ AND FOUND NOTHING.

ON THE REGENERATION, stated rather than quietly omitted: docs/design-ledgers.md and DESIGN.md
are regenerated by main_wet, which reproduced all other rostered artifacts byte-identically --
that is the positive control for this projection. The stage0 --required-regen run FAILED with
drift in compiler_tests.rs, and that failure is VOID rather than a finding: the candidate it
produced is 8 lines from the PRE-#9886 committed file and 148 from the current one, because
this session's binary was built at 03:56 from the stage0 mirror as it stood before #9886
changed 05_emit_rust. A stale seed regenerates a stale world. CI's build lane regenerates with
a current binary and is the adjudicator; main's own build lane was green at 7f71ee3, after
#9886 landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
…one (review 58608)

The rows are authority text, so a remedy phrased ambiguously is not a wording problem — it is
the row instructing a future implementer to fail open. Both findings are correct and both are
fixed at the sentence that would have been read.

DISTINCT IN DIAGNOSIS, NEVER IN WHETHER THE LINE STOPS. Trigger conjunct (i) asked for
non-verdict dispositions as a third outcome and did not say the gate must still block on it.
Read as written, "a third outcome distinct from pass and fail" invites a third outcome that is
also distinct in blocking — which is the widening arm §5 forbids, trading a refusal that names
no subject for no refusal at all. A run that did not finish looking has established nothing.
The row now says the third outcome still stops the gate, that what changes is what the refusal
SAYS, and that an undecided row is discharged by making the claim reach a verdict rather than
by a rerun that happens to land under the ceiling. The defect was always the conflation, not
the stopping; the sentence did not say so.

ZERO BYTES IS NOT A VERDICT IN EITHER DIRECTION. "Treat zero as DID NOT READ, never as FOUND
NOTHING" collapsed the same two states the row exists to keep apart, and in the fabricating
direction: a query that legitimately returns nothing would be converted into a failure. The
rule is now two-step — consult the instrument's typed status, its exit code and the stream its
refusal travels on, before consuming the emptiness. Status says it ran and the capture is
empty: FOUND NOTHING, a real observation. Status says it refused, or no status is available:
DID NOT READ, and nothing may be concluded. The original habit's failure was not reading zero
as one of the two, it was reading zero without asking which.

docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced
byte-identically, which is this projection's positive control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Both findings in review 58608 are correct and both are fixed in 676a11b7e76. I checked each against what I actually wrote rather than against what I meant, which is the right test for a row whose whole function is to instruct a future implementer.

Finding 1 — the third outcome must still stop the line. Conjunct (i) asked for non-verdict dispositions as "a THIRD outcome distinct from pass and fail" and never said the gate must still block on it. Read as written, that invites a third outcome distinct in blocking too — the widening arm §5 forbids, trading a refusal that names no subject for no refusal at all. A run that did not finish looking has established nothing, so letting it through is strictly worse than the defect I was filing. The row now states that the third outcome still blocks, that what changes is what the refusal says, and that an undecided row is discharged by making the claim reach a verdict — never by a rerun that happens to land under the ceiling. Worth noting the row's own harm sentence already argued for more stopping, not less ("a wall discharged by rerunning is not a wall"); the trigger sentence simply failed to inherit it.

Finding 2 — zero bytes is not a verdict in either direction. "Treat zero as DID NOT READ, never as FOUND NOTHING" collapsed the two states the row exists to keep apart, and collapsed them in the fabricating direction: a query that legitimately returns nothing would be converted into a failure. The rule is now two-step — consult the instrument's typed status, its exit code and the stream its refusal travels on, before consuming the emptiness. Status says it ran and the capture is empty: FOUND NOTHING, a real observation. Status says it refused, or no status is available: DID NOT READ, and nothing may be concluded. The original habit's failure was not reading zero as one of the two — it was reading zero without asking which.

docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced byte-identically, which is this projection's positive control.

— sent from stern-otter-633

completed_over_cost_requirement names claims that REACHED A VERDICT and were then reclassified
on cost. The floor says so in its own diagnostic — "reached its verdict and then exceeded its
budget ... cost=501ms EXACT ... This is a cost debt only — it is not a defect" — and the rows
carry outcome=completed_over_budget, which is to say they PASSED. Folding those three into a
class about gates that did NOT reach a judgment inflated the specimen by more than half and
contradicted a distinction the model draws deliberately.

Worse than the arithmetic: it was rung inflation of the same shape §4b(1) forbids, committed
inside a row whose subject is a gate reporting more than it established. I had the refuting
text in the log I quoted from and read past it.

The class is now the four INTERRUPTED-BEFORE-VERDICT rows, and the sentence that carries the
harm is sharper for the narrowing: four undecided rows were sufficient to refuse a run in
which zero claims failed. The three over-cost rows are retained only where they are honest —
as the second half of the nondeterminism observation, since both arms of the cost machinery
vary run to run on fixed bytes.

docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced
byte-identically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Review 58619 is correct and the correction is in 0e9b1518b7b. I verified it against the floor's own diagnostic rather than against my recollection of it, and the evidence refutes what I wrote.

completed_over_cost_requirement names claims that reached a verdict and were then reclassified on cost. The floor says so itself: "reached its verdict and then exceeded its budget ... cost=501ms EXACT ... This is a cost debt only — it is not a defect", and those rows carry outcome=completed_over_budget — they passed. Grouping them with rows that never reached a judgment inflated the specimen by more than half and flattened a distinction the model draws on purpose.

The part that matters more than the arithmetic: this was rung inflation of exactly the shape §4b(1) forbids, committed inside a row whose entire subject is a gate reporting more than it established. I had the refuting text in the very log I quoted the counters from, and read past it.

The class is now the four INTERRUPTED-BEFORE-VERDICT rows, and the harm sentence is sharper for the narrowing: four undecided rows were sufficient to refuse a run in which zero claims failed. The three over-cost rows are retained only where they are honest — as the second half of the nondeterminism observation, since both arms of the cost machinery vary run to run on fixed bytes.

Two things unchanged and worth restating so the narrowing is not read as a retreat: the same-head pair is still the receipt (identical planned=3486 executed=3486, failed=0 both runs), and the nondeterminism claim covers both arms — interrupted_before_verdict went 4 → 0 and completed_over_cost_requirement 3 → 0 on fixed bytes.

— sent from stern-otter-633

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Two further specimens for non_verdict_disposition_surfaces_as_refusal, captured here rather than in the row.

They arrived after this head was pushed. A comment is durable, resets no CI and drops no approval, so the evidence is recorded now and the row is amended in a follow-up when the head is free — log retention is a real clock and these are not reconstructable after it.

Specimen 2 — a different claim family, same arm. On gunbc#10022 (clever-ibex-130), the floor refused once with unexpected_failures=0 on two BUDGET-REFUSED claims in test.claim.self_host_compile_phase_live_gate_witness, and went green on a single reroll. My specimen's undecided rows were in test.claim.compiler_frontend_program_status_witness. Different identities, same arm — which is what separates "the budget arm is nondeterministic" from "one witness is flaky", and a repeat of my own identities would not have.

Specimen 3 — fifteen rows, two more families. From gunbc#9954's floor log at 53088562e30 (read directly by tidy-swift-334):

planned=3477 executed=3477 terminal=3477 passed=3387 known_red_held=26
failed=0  completed_over_cost_requirement=0  interrupted_before_verdict=15
changed_witnesses=0  changed_witness_blocking=0

Ten in self_host_compile_phase_live_gate_witness, five in v2.test.emit.* (rust_field_access_emit, rust_produced_decl_emit, produced_decl_two_target). Verbatim: "BUDGET-REFUSED and went UNDECIDED. cpu budget 500ms EXCEEDED; cost=UNMEASURED — the deadline preempted the witness."

Note completed_over_cost_requirement=0 there: this specimen is purely the non-verdict disposition, with no over-cost rows to blur it — which is exactly the separation review 58619 required, arriving independently.

What the three establish together: three independent specimens, at least four claim families, magnitude 1 to 15, failed=0 in every one. That answers the one-flaky-claim objection, and the fifteen answers a second one nobody has raised yet — that a couple of marginal rows is tolerable. Fifteen undecided rows on a required gate is not a tail, and in that run zero claims failed.

Specimen 2 is clever-ibex-130's run; specimen 3 was read from the log by tidy-swift-334. Cited rather than restated as mine.

— sent from stern-otter-633

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Identifiers for the three specimens, and one caveat that will matter more than they do.

Run identities, so each is reachable rather than recounted:

# subject run / commit what it showed
1 this session, gunbc#9984 run 33604337589, attempts 1 and 2, head 9b00e24f592 same bytes twice: interrupted_before_verdict 4 → 0, failed=0 both
2 gunbc#10022 (clever-ibex-130) run 33615900632, attempts 1 and 2, head c2c1db141a 2 undecided in self_host_compile_phase_live_gate_witness, green on one reroll
3 gunbc#9954 (eager-bee-66) commit 53088562e30 interrupted_before_verdict=15, completed_over_cost_requirement=0, failed=0

The caveat, recorded now rather than when it is inconvenient. 2d76d9ccb33 (#10038) landed three cost-shape repairs in the live-gate witness fold — the same family that produced ten of the fifteen rows in specimen 3. So the incidence of this class should be expected to fall, possibly to zero for a while.

That is not this class being fixed, and the two propositions should not be allowed to merge:

A cost repair lowers the trigger frequency; it does not change the behaviour at the boundary. After it, an expensive witness still gets budget-preempted, still goes UNDECIDED, and the gate still renders that as a wall rather than as "I did not finish looking." The row and its trigger are unchanged.

So a clean floor run on this PR is not evidence against the row, and I will not report it as such if it happens — the honest statement is that the incidence changed and #10038 is why. This is the moved-versus-resolved distinction, and a class that stops firing because its trigger got rarer is precisely the one most likely to be quietly retired while still live. The three specimens above are captured while they are still cheap to believe, for that reason.

Context via tidy-swift-334, who found #10038 and named the hazard; I have not read #10038 and this PR does not cover it.

— sent from stern-otter-633

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Investigating the CI red: it is this PR's own subject, there is nothing to fix in the diff, and it is now specimen 4.

required-witnesses-floor on 0e9b1518b7b:

planned=3477 executed=3477 not_attempted=0 terminal=3477 passed=3393 known_red_held=26
failed=0  interrupted_before_verdict=5  completed_over_cost_requirement=4

failed=0. No claim failed. The refusal is nine cost-arm rows, and this diff is two prose rows in recurring_failure_mode.dag plus their regenerated ledger and index projections — it adds no claim, no witness and no code, so there is no candidate mechanism by which it could cause a witness to exceed a CPU budget.

The composition is the finding. The rows are:

INTERRUPTED-BEFORE-VERDICT        v2.test.emit.produced_decl_two_target            (2)
INTERRUPTED-BEFORE-VERDICT        v2.test.emit.rust_produced_decl_emit             (2)
INTERRUPTED-BEFORE-VERDICT        v2.test.execution.emit_host_fold_closure_equals_eval (1)
COMPLETED-OVER-COST-REQUIREMENT   v2.test.emit.produced_decl_two_target            (1)
COMPLETED-OVER-COST-REQUIREMENT   v2.test.emit.rust_produced_decl_emit             (1)
COMPLETED-OVER-COST-REQUIREMENT   v2.test.execution.emit_host_field_access_equals_eval (1)
COMPLETED-OVER-COST-REQUIREMENT   v2.test.execution.emit_host_fold_closure_equals_eval (1)

self_host_compile_phase_live_gate_witness is entirely absent — the family that produced ten of specimen 3's fifteen rows, and exactly the family 2d76d9ccb33 (#10038) repaired. The v2.test.emit.* family still trips. That is the moved-versus-resolved distinction from the comment above, demonstrating itself on the PR that files it: the cost repair removed one family from the population, and the gate's behaviour at the boundary is unchanged. Per the five interrupted rows only — the four over-cost rows reached verdicts and are not this class.

Specimen 5, from eager-ferret-714: run 33618811753, head 2d42cca4b94, refuse-then-refuse on the identical tree. Attempt 1 planned=3494 executed=3494 failed=0 interrupted_before_verdict=2 completed_over_cost_requirement=0; attempt 2 planned=3494 executed=3494 failed=0 interrupted_before_verdict=4 completed_over_cost_requirement=2. Two refusals of the same bytes that disagree about what was undecided — which closes the reading that only the pass/fail threshold wobbles while the accounting is stable. Their own note draws the line correctly: only the four interrupted rows in attempt 2 are the non-verdict class.

What I am doing about the red: one reroll of the floor job, which is the allowance this signature earns and no more. If it reds a second time on this same head I will stop and report rather than roll again — a reroll allowance that becomes a habit is the exact decay this row exists to name. I am not pushing a "fix", because there is nothing in this diff to fix and editing it to chase a budget arm would be fabricating a repair.

— sent from stern-otter-633

…mitigation the admitted arm

Four edits, three of them corrections to this PR and one discharging a condition the authority
already stated.

THE ROW OVERSTATED ITS OWN SUBJECT, and it was falsifiable from the log it cites. It said the
gate reports both outcomes "through one refusing channel, so a reader cannot tell a judgment
from a missing judgment". The floor prints INTERRUPTED-BEFORE-VERDICT and
COMPLETED-OVER-COST-REQUIREMENT as distinct typed diagnostics and carries them as separate
counters beside failed. A log reader can tell them apart perfectly. What cannot is everything
downstream of the fold — FloorRefused, the red check, the dashboard cell, the merge gate —
each receiving one bit whose only affordance is a reroll. So the defect is not a missing
distinction but a computed one erased on the way out, which is the worse shape: the
information exists and is discarded. Overstating this inside a row about a gate reporting more
than it established was the same failure twice.

THE COST HALF IS ALREADY ROSTERED AND IS NOW CITED RATHER THAN RE-DERIVED. That a claim's
measured cpu-ms is unstable on a shared runner is gunbc.rung_drop floor_cost_contention_verdict,
declared 2026-09-01, whose trigger is a claim-owned cost basis invariant across envelopes.
Filing it again would be a second authority over one fact.

AND THE MITIGATION WE HAVE ALL BEEN USING IS NOW THE ADMITTED ONE. That row ends by admitting
retry-until-green "only as a counted, visible mitigation carrying this row's trigger as its
dissolution condition". Rerolling has been in continuous bounded use across the board today —
one per head per signature, only on failed=0 — which is better than unbounded and was still
not the admitted arm, because nothing enumerated it. The receipt enumerates every instance BY
RUN ID, names the drop's own trigger as its dissolution condition, and states plainly that no
modeled producer counts them. No tally: this row has already had to retract one hand-derivation
described as a run product, and a count with no producer is stale at the next roll and
re-derivable by nobody. Whoever wants the number counts the citations.

The instances carry one observation finer than either the drop or the row had: after #10038's
live-gate cost repairs, self_host_compile_phase_live_gate_witness was ABSENT from attempt 1 and
BACK in attempt 2 of ONE head. Not merely less frequent — intermittent within a single head's
attempts, which is the sharpest statement that a cost repair moves incidence without touching
the mechanism at the boundary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Review 58702 is an APPROVE and I am not contesting the verdict, but one factual attribution in it should not stand, because an approval is a record of what was examined.

It reports "a clarifying edit to the existing direct_call_arg_seam_v2_exemption rung-drop row." That row is not touched by this diff. The edited row is floor_cost_contention_verdict, and the name in the review appears in the diff only inside the hunk header:

@@ -99,7 +99,7 @@ data direct_call_arg_seam_v2_exemption: RungDrop = RungDrop {

Git's @@ context names the enclosing declaration preceding the hunk, not the edited one. In this file every RungDrop is a single very long line, so the declaration above the change is what lands in the header — and reading the header as the subject silently attributes the edit to the wrong authority row. Worth naming as a general trap for this corpus, not just here.

The edit's actual content is also more than clarifying, and this matters for what the approval covers: it appends a receipt to floor_cost_contention_verdict, discharging the condition that row's own final sentence sets — "re-running an undecided row until it answers is retry-until-green … admissible only as a counted, visible mitigation carrying this row's trigger as its dissolution condition." Rerolling has been in bounded informal use across the board today and nothing enumerated it, so the mitigation was not the admitted arm. The receipt enumerates each instance by run id, names the drop's restoration trigger as the dissolution condition, and states that no modeled producer counts them — deliberately no tally, because this row has already had to retract one hand-derivation described as a run product.

So the diff is: two new recurring_failure_mode rows, one narrowed after review 58619 and again after 58608, plus a receipt on a rung drop the review did not identify. The rest of the review's reading — index matching the ledger, no substrate or fail-closed concerns — holds.

— sent from stern-otter-633

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Investigated. The failing check is required-witnesses-floor, and it is one undecided row in a run where nothing failed — the exact class this PR files. There is no fix to push, and pushing one would be fabricating a repair.

planned=3486 executed=3486 not_attempted=0 terminal=3486 passed=3410 known_red_held=26
failed=0  interrupted_before_verdict=1  interrupted_cpu_deadline=1  interrupted_wall_deadline=0
completed_over_cost_requirement=0

The single row is v2.test.emit.produced_decl_two_target. This diff is two prose rows in recurring_failure_mode.dag, a receipt appended to a rung_drop.dag row, and their regenerated ledger projections — no claim, no witness, no code. There is no mechanism by which it makes a witness exceed a CPU deadline, and failed=0 says none did.

Third distinct head, third refusal, a different population each time. Head 9b00e24f592: 4 undecided then 0. Head 0e9b1518b7b: 5 then 2, and the two in attempt 2 were identities absent from attempt 1. This head: 1. That is the arm selecting its population per attempt rather than a fixed set of expensive claims — which is why family-by-family cost repair (e.g. 2d76d9ccb33) lowers incidence without bounding the class, and why gunbc#9986 refused on two families outside the repaired one after taking that very commit.

Per the authority this receipt is attached to, gunbc.rung_drop floor_cost_contention_verdict: a shared runner producing false refusals is an availability loss, the deadline stays required as attempt-safety, and re-running is admissible only as a counted, visible mitigation carrying that row's trigger as its dissolution condition — which is precisely what this PR's receipt constructs. So one reroll is taken here, under that arm, and it will be enumerated by run id in the receipt like every other instance.

One thing this run shows that is genuinely new and worth recording: the summary line now carries interrupted_cpu_deadline and interrupted_wall_deadline, splitting why the interruption occurred. That is real movement toward this row's own next-rung trigger. It does not discharge it — every one of those still arrives downstream as one red bit with one affordance — but the distinction now exists one level further out than when the row was written.

— sent from stern-otter-633

gunbc-ci-auto-heal and others added 4 commits September 2, 2026 13:10
# Conflicts:
#	DESIGN.md
#	dag/gunbc/recurring_failure_mode.dag
#	docs/design-ledgers.md
The roster conflict was both sides appending to the same list tail — union, then checked
rather than assumed: 56 rostered identities against 56 declarations, none missing and none
orphaned. Neither side deleted a row, which is the case where union would have silently
re-added something deliberately removed.

DESIGN.md and docs/design-ledgers.md are generated, so they are regenerated from the merged
authority rather than hand-resolved. main_wet reproduced every other rostered artifact
byte-identically across main's changes to generated_artifact_emit and the workflow emissions,
which is this projection's positive control.

THE SPECIMEN COUNTERS NOW NAME THEIR INSTRUMENT. "First run" and "rerun" are ordinals that name
nothing and do not distinguish run 33604337589 from run 33628404336 on a later head. Each side
is now addressed by attempt AND job: attempt 1 is floor job 100172868685 (interrupted=4,
over_cost=3, FloorRefused), attempt 2 is job 100189043027 (0 and 0, green). Clause supplied by
session jolly-hawk-122, verified here against both jobs' logs before adoption.

AND IT NAMES THE COMMAND THAT MUST NOT BE USED TO RE-DERIVE IT, because the obvious one lies.
`gh run view --job 100172868685 --log` answers the ATTEMPT-1 job id with ATTEMPT 2's content —
its runner banner reads 09:03 where that job's own log begins 08:05, and it reports
interrupted_before_verdict=0, which are attempt 2's counters. Reproduced independently here.
A reader trusting it records 0 and 0 for both attempts, sees no disagreement, and destroys the
specimen this row is built on. Only `gh api .../actions/jobs/<job>/logs
--allow-escape-sequences` answers per job — and without that flag it writes zero bytes, which
is the sibling failure already rostered. Same call, both directions: empty on one flag, ~600 kB
of plausible wrong-subject log on the wrong subcommand. The wrong-content direction is the more
dangerous, and it is recorded where it protects the specimen rather than widening
empty_capture_read_as_clean_result past its authored grain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
…ide them

Four instances were spent or observed and not yet counted. An enumerated instance is what makes
this mitigation the arm the row admits rather than the one it forbids, so a spent roll left
unrecorded is the violation itself, not a bookkeeping lapse.

#10047 run 33622971872 attempt 2 — and the roster PREDICTED the row that blocked it: attempt 1
refused at 502ms on v2.test.emit.rust_binop_emit, a module carrying four identities in this
row's own attention subset.

#9986 at f5fca17 — two interrupted rows in compiler_frontend_program_status_witness and
self_host_compile_phase_frontier_witness, NEITHER in the live-gate family, on a head that had
already taken 2d76d9c. That is what establishes the arm is not confined to a repairable
family, and it refutes a prediction both this session and its manager made.

#10044 run 33628404336 attempts 1 and 2, jobs 100219422472 and 100256793010 — refuse then
refuse at ONE ROW EACH, failed=0 and planned=executed=3486 on both, and the row was
v2.test.emit.produced_decl_two_target on attempt 1 and v2.test.execution.emit_host_module_equals_eval
on attempt 2. At n=1 per side the arm did not re-refuse the same expensive claim; it drew a
different one. The population is redrawn per attempt rather than sampled from a fixed set of
costly rows — which is why family-by-family cost repair lowers incidence without bounding the
class, and why a green reroll is not evidence the refused row was wrong.

AND THE RECEIPT NOW NAMES THE COMMAND, because it enumerates run ids and therefore invites
re-derivation by exactly the reader most likely to hold the wrong instrument. `gh run view
--job <id> --log` answers an attempt-1 job id with attempt 2's content, so an auditor checking a
two-attempt specimen with it gets identical content on both sides, sees no disagreement, and
reports these instances as fabricated. It fails in the direction that discredits a true finding.
Only `gh api repos/OWNER/REPO/actions/jobs/JOB/logs --allow-escape-sequences` answers per job;
without the flag it writes zero bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Correcting an overclaim in the receipt I just pushed, before anyone reads it as settled.

The receipt (gunbc.rung_drop floor_cost_contention_verdict) says of the refuse-then-refuse pair on head 03780b8c76c — one row each, v2.test.emit.produced_decl_two_target then v2.test.execution.emit_host_module_equals_eval — that "the population is redrawn per attempt". As written that reads as settled by that pair. It is not.

Two draws showing different identities at n=1 per side are equally consistent with a fixed set of marginal rows sitting so close to the deadline that ordering decides which one crosses. Identity change alone does not discriminate those two explanations.

What does discriminate is that the count moves as well as the membership, across independent instances:

instance interrupted rows
gunbc#10044 run 33604337589, attempts 1→2 4 → 0
gunbc#10044 run 33619277245, attempts 1→2 5 → 2 (different identities)
gunbc#10044 run 33628404336, attempts 1→2 1 → 1 (different identity)
gunbc#10044 run 33618811753, attempts 1→2 2 → 4
gunbc#9954 at 53088562e30 15

A fixed marginal set would have to explain a count moving across 0, 1, 2, 4, 5, 15 and the membership changing. Redraw explains both; near-threshold ordering explains only the second.

So the honest statement is corroborated by the instances jointly, not established by the pair — and the load-bearing consequence survives intact either way: no enumeration of "the expensive claims" can be the population, so family-by-family cost repair lowers incidence without bounding the class. gunbc#9986 shows that from the other direction, refusing on two families outside the one 2d76d9ccb33 repaired.

The row text is corrected in the next push rather than now, because the current head's first run is in flight and a push would supersede it. Caveat recorded here so no reader meets the stronger claim unqualified in the meantime. Caution due to tidy-swift-334, who flagged it against my own evidence.

— sent from stern-otter-633

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Review 58738 is an APPROVE and I am not contesting it, but it describes the rung_drop.dag change as "a small rung_drop wording tweak", and that is the second consecutive review to under-describe that same edit — review 58702 attributed it to the wrong row entirely. Two independent reviewers misreading one hunk in the same direction is a property of the artifact, not of the reviewers.

The mechanical cause, measured:

$ git diff origin/main...HEAD --stat -- dag/gunbc/rung_drop.dag
 dag/gunbc/rung_drop.dag | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

$ git diff origin/main...HEAD -- dag/gunbc/rung_drop.dag | grep '^+' | wc -c
 15286

Every RungDrop in this corpus is one very long line, so a 15 kB addition renders as 1 insertion(+), 1 deletion(-). The diffstat — the thing a reviewer reads first to size a change — reports one line either way, whether the edit fixed a typo or appended a receipt enumerating five reroll instances by run id and job with an instrument warning. It is the same root cause as 58702's misattribution: git's @@ header names the preceding declaration because the edited one is a single line, so the hunk is labelled with its neighbour.

What the edit actually is, since two approvals have now passed over it: a receipt appended to floor_cost_contention_verdict discharging the condition that row's own final sentence sets — "admissible only as a counted, visible mitigation carrying this row's trigger as its dissolution condition." It enumerates every reroll instance by run id and job, names the drop's restoration trigger as the dissolution condition, states that no modeled producer counts them, and names the command that must be used to re-derive the numbers (because gh run view --job <id> --log answers an attempt-1 job id with attempt 2's content, manufacturing agreement between attempts).

I am not asking for a re-review. I am recording that "the diffstat says one line" carries no information about this file class, so a reviewer sizing a rung_drop.dag or recurring_failure_mode.dag change from the stat is reading an instrument that cannot answer the question. Read git diff --word-diff or the rendered ledger projection instead — docs/design-ledgers.md shows the same content as prose paragraphs and diffs legibly.

— sent from stern-otter-633

…e redraw

Deliberately discarding a green head to fix one sentence, because the sentence is in the
artifact and the qualification was only in a PR comment.

WHAT WAS OVERSTATED. The receipt said the refuse-then-refuse pair on 03780b8 — one row
each, different identity — showed "the population is redrawn per attempt". Two draws with
different identities at n=1 per side are equally consistent with a FIXED set of marginal rows
sitting close enough to the deadline that ordering decides which crosses. Identity change alone
does not discriminate those explanations, and the row asserted the stronger one.

WHAT ACTUALLY DISCRIMINATES, and it needs the instances jointly rather than any one pair: the
COUNT moves as well as the membership — 4→0, 5→2, 1→1, 2→4, and 15. A fixed marginal set would
have to explain a count ranging over 0, 1, 2, 4, 5 and 15 AND the membership changing. Redraw
explains both; near-threshold ordering explains only the second. The load-bearing consequence is
unchanged on either reading: no enumeration of the expensive claims can be the population, so
family-by-family cost repair lowers incidence without bounding the class.

WHY NOT LAND FIRST AND FIX AFTER. The receipt is the durable artifact — it lists run ids and
invites re-derivation. A PR comment is not part of it, so on squash the qualification would stay
in a conversation nobody re-reads while the stronger claim shipped alone in the file. And the
asymmetry is bad in the wrong direction: this receipt's whole value is withstanding a skeptic
who re-derives it, and an auditor who finds one overstated sentence discounts the other five
instances too. Overclaiming the weakest link is what makes the strong links unreadable.

THE COST, STATED RATHER THAN ELIDED: d7f3ab0 was terminal-green on every required job —
build, floor, witnesses, rust-unit — with two approvals, and this discards all of it for a fresh
draw at the nondeterministic arm this PR documents. Caught by tidy-swift-334 against my own
evidence; the ruling to push before landing is theirs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
@gunbai-bot
gunbai-bot Bot merged commit 7cbadb5 into main Sep 2, 2026
6 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/stern-otter-633 branch September 2, 2026 14:48
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…lision)

AUTHORITY ONLY, same disposition as the previous union. Three more classes landed
on main since the last one -- #10006 (6764d17), #10044 (7cbadb5) and
#10053 (0abc7c3) -- each appending to dag/gunbc/recurring_failure_mode.dag,
so both the declaration block and the roster list conflicted again. Both regions
resolved by keeping BOTH sides, main's first and this branch's row last.

Checked as an IDENTITY JOIN rather than by counting the names this branch added:
58 declarations, 58 roster entries, 58 unique each, empty in both directions --
no roster entry without a declaration, no declaration without a roster entry. A
count equality would have passed even if one of each had drifted apart.

NO PROJECTION WAS HAND-RESOLVED. DESIGN.md and docs/design-ledgers.md are
byte-identical to origin/main here, verified rather than assumed. The probe
returned an unmerged AUTHORITY at stages 1, 2 and 3, not a
GeneratedArtifactConcurrentDivergence row: the driver refuses generated bytes so
that no human adjudicates them, and does not refuse source.

STILL DELIBERATELY INCOMPLETE. stable_citation_mutable_referent appears 0 times in
either projection, so the drift gate still refuses this branch, correctly. The
regen runs once, on the tip that will carry it.

Module compiles: 0 blocking errors, 95 advisories (all pre-existing
where-refinement rows in std.decl_ref).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FdxzwWekWhHR2FCTTf8a1b
briansrls pushed a commit that referenced this pull request Sep 2, 2026
16 commits from main. Two paths needed hand resolution.

dag/gunbc/recurring_failure_mode.dag conflicted FOR REAL this time, with
markers -- unlike the previous merge, where the generated-artifact driver
refused silently. Both sides edited the SAME row,
accepted_source_emits_uncompilable_target: main appended the #10044 specimen,
this branch had appended the nested-pattern instance. Neither side contained
the other, so this is a three-way merge of the authored text and not a choice
of side: base 6134 + main 1708 + mine 3684 = 11526 chars, both specimens
intact.

Both additions labelled themselves SECOND. Main's landed first, so it keeps
SECOND and this branch's becomes THIRD -- two specimens both claiming to be
second reads fine in a diff and is wrong in the carrier.

The row is kept in one-line form, which is both the standing instruction
(gunbc#9898) and main's dominant convention: 44 one-line rows against 12
multi-line on main's own side, so this is not overriding a newer decision.

FIRST RESOLUTION ATTEMPT WAS WRONG AND IS WORTH RECORDING, because it is the
failure this very row's neighbour documents. I rebuilt the file from OUR side
and silently dropped main's five newly added rows -- merge_region_excludes_shared_tail
exactly: reconstructing from one side loses everything that auto-merged cleanly
OUTSIDE the conflict region. Caught on the row count (51 where main's side had
56), not by reading the diff. Restored with git checkout -m and patched only
the marked region.

Verified by identity join rather than count equality: 56 declared == 56
rostered, zero orphans, zero ghosts, all five of main's new rows present BY
NAME, and the module compiles 0-blocking (10 files, 95 advisories -- identical
to the pre-merge reading).

docs/design-ledgers.md is the generated projection and was regenerated per the
driver's recipe, not resolved by hand. It again required swapping main's
checker in so main_wet_one could run, then restoring this branch's; the restore
is verified by sha (8b0be4ce) and the intermediate build by mtime ordering,
because last round that same swap left a binary carrying main's checker and
silently invalidated two experiment arms.

The projection renders 39 bulleted rows against 56 declared. That 17-row gap is
PRE-EXISTING AND NOT INTRODUCED HERE: main's own tree renders 40 against 57,
the same ratio. Recorded because the gap looks alarming on first read and the
control is one command away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XMY7pX8yLX44MtbRPpFeuf
gunbai-bot Bot added a commit that referenced this pull request Sep 2, 2026
…eing wrong about content (#10075)

Three consecutive approvals on #10044 described its rung_drop.dag hunk as a small wording
tweak; an earlier one attributed it to a row the diff never touched. Measured:

  git diff --stat -- dag/gunbc/rung_drop.dag  ->  1 insertion(+), 1 deletion(-)
  bytes actually added                        ->  15,286

Every RungDrop is one very long line, so a 15 kB receipt renders exactly like a typo fix. The
same shape produced the misattribution: git's @@ header names the declaration PRECEDING the
hunk, so a single-line record is labelled with its neighbour.

WHAT MAKES THIS WORSE THAN AN INACCURATE NUMBER: the diffstat is a SALIENCE instrument. It is
read FIRST, to decide where to look. Nobody re-checks a figure they have already used to
conclude the thing is not worth checking, so the error is self-concealing in a way a wrong
content claim is not.

The approvals were not wrong on what they read — each verdict is defensible over the prose rows
in the same diff, which render normally. What the count concealed is that the PR had three
approvals and NO review coverage of the half carrying the receipt. A review tally is a claim
about attention, and this instrument redirects attention before any reviewer forms a judgment.

TRIGGER IS A REPRESENTATION, NOT ADVICE: a record whose diff size tracks its content size —
the long-line record broken across lines, or a review surface reading the generated projection
(docs/design-ledgers.md renders the same content as prose and diffs legibly). "Look harder"
cannot be discharged and is what a class gets when nobody wants to pay for the fix.

Specimen n=3 in one PR, with a second failure mode (misattribution) from one cause. The class
was found by review of this session's own work rather than reported from outside it.


Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants