Skip to content

File two recurring failure classes found in this session, and give positional_citation its edit-form instance - #10052

Merged
gunbai-bot[bot] merged 9 commits into
mainfrom
session/sunny-gull-270-harm-axis
Sep 2, 2026
Merged

gunbai-bot[bot] merged 9 commits into
mainfrom
session/sunny-gull-270-harm-axis

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Draft deliberately — two roster rows are not urgent, the fleet is the active incident, and a draft costs less. This branch was pushed by the autocommit daemon and auto-opened; it goes ready when the fleet is healthy.

Both rows were found while doing other work in this session, and both specimens are this session's own mistakes.

hedged_benefit_leaves_the_harm_axis_unexamined

An actor takes an irreversible or side-effecting action, states a calibrated caution about whether it helps, and never asks whether it harms. Different axes; no amount of hedging on the first reaches the second. The hedge is what makes it dangerous rather than merely wrong — it reads as due diligence, so a reviewer sees the epistemic care and endorses, and the endorsement generalises a single act into practice.

Specimen. This session cancelled a superseded CI run on a stale head to return a fleet slot during saturation, reporting it as "I have no evidence it helps rather than merely stops waste." The manager verified the reasoning and endorsed it as "the only lever any of us actually has." Both wrong, and the authority was one grep away: gunbc.witness_floor_workflow emits cancel-in-progress: false with NeverSupersedeRunning, because a cancelled run SIGKILLs cargo mid-build and every borrowed jobserver permit is lost for the daemon's lifetime. The act is the mechanism that produces the exhaustion symptom it was taken to relieve.

Recognition rule, mechanical: read the hedge and ask which axis it is on. An actor who has examined harm names the mechanism by which harm would occur, then says why it does not apply. An actor who has only hedged benefit names no mechanism, because there is none to name.

Rung 1, ceiling 2 — whether side effects are harmful is not decidable from the action alone, but whether the governing policy was consulted is.

receipt_names_a_property_not_the_tree_it_holds_of

Evidence of the form "regenerating produces the committed bytes" is a claim about a base, recorded and cited as a claim about a change. When the base moves it keeps reading as valid while describing a tree that no longer exists.

It is the worst-signalling member of its family, which is why it earns a row: a stale citation still shows its text, an incompatible merge still conflicts, an absent observation still reports an absence — an expired green produces nothing. Worse, the author's own careful re-verification re-affirms it.

Specimen. #10036's entire acceptance test was byte-identity of an emitted workflow. It was re-verified after each of three commits, and all three measured against a base main had already left — #10024 had rewritten the same authority by 151 lines and the same artifact by 140. No signal at any point.

Recognition rule: before citing a zero-diff receipt as merge evidence, ask whether the base moved on the files the receipt is about — git diff --stat $(git merge-base HEAD origin/main) origin/main -- <those paths>.

Rung 1, ceiling 2. The trigger is a tracked stall rather than a wish, and that is the load-bearing part: the capability is already implemented one surface over. Review 58587 on #10036 failed with worktree freshness check failed: HEAD is ad388184ac but PR #10036 head is c80c9d88db — refusing to review a stale/wrong checkout. Reviews carry their base and refuse on mismatch; regeneration receipts do not. The gap is an unapplied mechanism, not a missing one.

Notes

Checked for existing rows before minting each, since two earlier findings this session turned out to belong to rows that already existed — one of those became an append to positional_citation (the edit-form instance) rather than a new row. execution_provenance_loss asks whether an observation ran; positional_citation is a decaying handle. Different invalid states, different repairs.

Appended as one-line units, per merge_region_excludes_shared_tail's own prescription for this carrier. Both projections (DESIGN.md index, docs/design-ledgers.md content) regenerated from the authority.

🤖 Generated with Claude Code

https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct

gunbc-ci-auto-heal and others added 5 commits September 2, 2026 09:46
…hile leaving HARM unexamined

DESIGN §4b: every newly discovered error class files one row. This one was
discovered with a live specimen, and the specimen is the author.

THE CLASS. An actor takes an irreversible or side-effecting action, states a
calibrated caution about whether it HELPS, and never asks whether it HARMS. Those
are different axes and no amount of hedging on the first reaches the second. The
hedge is what makes it dangerous rather than merely wrong: it reads as due
diligence, so a reviewer sees the epistemic care and endorses, and the
endorsement generalises a single act into practice.

THE SPECIMEN. This session cancelled a superseded CI run on a two-commit-stale
head to return a fleet slot during saturation, and reported it with the explicit
caution "I have no evidence it helps rather than merely stops waste". The manager
verified the reasoning and endorsed it as "the only lever any of us actually
has". Both wrong, and the authority was one grep away: gunbc.witness_floor_workflow
emits `cancel-in-progress: false` with `NeverSupersedeRunning`, because a
cancelled run SIGKILLs cargo mid-build and every borrowed jobserver permit is
lost for the daemon's lifetime -- the pool decays monotonically toward zero and a
host at zero permits still accepts jobs that simply WAIT. So the act is the
mechanism that produces the exhaustion symptom it was taken to relieve, and the
stale runs it "reclaimed" are that policy's DECLARED cost.

THE RECOGNITION RULE IS MECHANICAL, which is what makes this worth a row rather
than a lesson: read the hedge and ask which axis it is on. "I cannot show this
helps" is a BENEFIT hedge; if the action touches shared state at all, its presence
is evidence the harm question was never asked. The reviewer's tell: an actor who
has examined harm NAMES THE MECHANISM by which harm would occur and then says why
it does not apply. An actor who has only hedged benefit names no mechanism,
because there is none to name.

RUNG AND CEILING, stated honestly. Found at 1, mitigatable, and by nothing
structural -- caught after the fact by reading an authority nobody was required to
read. Ceiling 2, not higher: whether side effects are harmful is not decidable
from the action alone, since the governing policy may live in any authority. What
IS decidable is whether that policy was CONSULTED. Next trigger names the
capability, not an artifact: a typed standing binding each fleet-affecting
operator action to the authority governing it, so taking the action without
consulting it refuses rather than relying on the actor to grep. Until then this is
review discipline and citing it as coverage is rung inflation.

APPENDED AS ONE LINE, which is the repair `merge_region_excludes_shared_tail`
prescribes for this exact carrier: a multi-line unit ending in a shared
`evidence: [],` / `}` tail makes every two-lane append resolve wrong by default.

Both projections regenerated from the authority: the DESIGN.md index line and the
docs/design-ledgers.md content, one line each.

CAUGHT WHILE WRITING IT, and it is the same class of mistake one layer down: my
first roster-list append anchored on `one_refusal_two_destinations,\n]` and
SILENTLY NO-OPPED, because main had gained `liveness_probe_read_as_currency`
underneath me. A `grep -c` showing 1 where 2 was owed is what caught it. An
anchored edit that misses is indistinguishable from one that lands unless you
count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
…the receipt, do not mint a row

Manager ruling, and it is the §2 test applied to me for the third time tonight: a
second row for a mechanism that already has one is re-invention. The failure I
found while writing the harm-axis row belongs to `positional_citation`, whose rule
already is that a positional handle decays invisibly because anything above it
invalidates it while the containment tree names the same thing stably.

WHAT THE APPEND ADDS is the form the existing row did not cover. An edit anchored
on NEIGHBOURING TEXT -- sed on a surrounding line, a replace keyed on the adjacent
entry, an insert before a named sibling -- is a positional handle by another
spelling, and it decays identically. The DIFFERENCE IS THE CONSEQUENCE, and it is
strictly worse: a stale citation misleads a reader who can still see the text; a
stale ANCHOR silently does nothing, because a replace that matches zero times
reports success.

RECEIPT, this carrier, this session. Appending the harm-axis row needed two edits,
the declaration and the roster-list entry. The roster append was anchored on
`one_refusal_two_destinations,` plus the closing bracket. Between reading the file
and writing it, main gained `liveness_probe_read_as_currency` as the new last
entry; the anchor stopped matching, the declaration landed, the roster entry did
not, and NOTHING FAILED. The result would have been a declared-but-unrostered
class -- the half-applied state a roster cannot detect about itself.

WHAT CAUGHT IT IS THE TRANSFERABLE PART AND IT IS NOT VIGILANCE: `grep -c`
returning 1 where 2 was owed. The instruction is CHECK THE RESULT RATHER THAN THE
ACTION, because a no-op edit and a successful one are indistinguishable at the
actuator. The same discipline caught the mirror-image failure on this exchange
from the other direction: a manager reading `git show origin/main:<path>` rather
than a pinned worktree avoided refuting a correct citation. Two instances, one
root -- a read and a write separated by someone else's push.

RECOGNITION RULE: any edit whose anchor is text the edit does not own. State the
expected post-condition as a COUNT before applying it, and assert the count
afterwards; if you cannot say what the count should become, the edit is not
specified.

THE RULE CAUGHT MY OWN STATEMENT OF IT WHILE I APPLIED IT. I predicted the
harm-axis identity would appear twice after this edit and measured THREE -- the
third being the receipt above naming the row it cites. The prediction was wrong
and the count is right; had I asserted "2" mechanically I would have "found" a
defect that does not exist. A post-condition count is only as good as the reason
attached to it, which is why the rule says state it, not automate it.

Names kept as they happened, per the manager's instruction: a specimen whose
actors are anonymised is one nobody can falsify, and the endorsement is the step
that shows how a single act becomes practice.

docs/design-ledgers.md regenerated from the authority; DESIGN.md's index is
unchanged because no new identity was minted, which is the point.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
… names the TREE it holds of

Third roster row from this session, and the manager confirmed no other lane is
carrying the class. Checked both neighbours before minting, since the last two
findings belonged to rows that already existed:
  execution_provenance_loss asks whether the observation RAN. Here it ran, was
  valid, and then STOPPED being valid because its subject moved underneath it.
  positional_citation is a decaying HANDLE -- a pointer that no longer names what
  it named. Here the handle is fine; the EVIDENCE BINDING decayed.
Different invalid states, different repairs, so a third row rather than an append.

THE CLASS. Evidence of the form "regenerating produces the committed bytes", "the
diff is zero", "the census matches", "the fixed point is reached" is a claim about
A BASE, recorded and cited as a claim about A CHANGE. When the base moves it keeps
reading as valid while describing a tree that no longer exists.

IT IS THE WORST-SIGNALLING MEMBER OF ITS FAMILY, which is why it earns a row
rather than a caution. A stale citation still shows the reader its text; an
incompatible merge still conflicts; an absent observation still reports an
absence. AN EXPIRED GREEN PRODUCES NOTHING. Worse, the author's own careful
re-verification RE-AFFIRMS it, because re-running the same command against the
same stale tree reproduces the same true-of-nothing answer.

SPECIMEN, and it is this session's own PR #10036. Its entire acceptance test is
that the emitted witnesses.yml is byte-identical after regeneration -- a
behaviour-identical refactor whose whole claim is that the bytes do not move. It
was re-verified after each of three commits and ALL THREE were measured against a
base main had already left: #10024 had landed, changing the exact authority the PR
rewrites by 151 lines and the exact artifact under test by 140. No signal at any
point. The expiry was found by asking what had changed in main, not by any check.

RECOGNITION RULE, mechanical: before citing a zero-diff receipt as merge evidence,
ask whether the base moved ON THE FILES THE RECEIPT IS ABOUT --
`git diff --stat $(git merge-base HEAD origin/main) origin/main -- <those paths>`.
Scoped to those paths deliberately: a receipt about generated bytes is invalidated
only by changes to the authorities that produce them, and a whole-repository
is-my-branch-behind check is both too coarse to act on and too noisy to keep.

TWO OBLIGATIONS THE OBVIOUS READING MISSES, both carried in the row. Re-derive the
discriminating CONTROL too -- one taken before the merge proves the zero was
readable on a tree that is not the one shipping, so a re-derived green beside a
stale control is half a measurement. And rebuild the PRODUCER, because the tree
and the emitter are two inputs and a regeneration with one stale proves nothing
about the pair.

AND THE UPSIDE IS IN THE ROW, because it changes whether anyone bothers: a receipt
re-derived after a FOREIGN change to the same artifact is STRICTLY STRONGER than
the original. The first green asks only whether a projection reproduces its own
base; the second asks whether it reproduces a base someone else moved, which is
where a projection that disagreed with the hand-kept form it replaced would
actually surface.

RUNG 1, CEILING 2, per the manager's framing and for a derived reason: a receipt
cannot make its subject stop moving, but it can make the movement OBSERVABLE.
Trigger names the capability, not an artifact -- a receipt CARRIES THE BASE IT WAS
MEASURED ON, so a producer compares it against HEAD and refuses an expired
receipt, rather than a person remembering to. The row says outright that until
then it is the mitigation and citing it as coverage is rung inflation.

MAIN MERGED BEFORE APPENDING, which is the rule the previous commit filed rather
than a coincidence: the last roster append silently no-opped against a moved base.
Post-condition stated before measuring -- 2 occurrences, one declaration and one
roster entry, no other row citing the identity -- and measured 2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
…s one surface over

The row named its next-rung trigger as a capability -- a receipt carries the base
it was measured on, so a producer refuses an expired one rather than a person
remembering to. True, but weaker than it needed to be: a trigger nobody can
picture reads as aspirational, and §4b's honesty obligation is about separating
"cannot climb" from "can climb but unbuilt".

IT IS THE SECOND, AND THE PROOF WAS IN THIS SESSION'S OWN REVIEW LOG. The review
runner ALREADY binds each review to the sha it was taken on and refuses on
mismatch. On #10036, review 58587 failed with

  worktree freshness check failed: HEAD is ad38818 but PR #10036 head is
  c80c9d8 -- refusing to review a stale/wrong checkout

and a refused review is exactly the signal an expired receipt does not produce.
So reviews carry their base and regeneration receipts do not. The gap is an
UNAPPLIED mechanism, not a missing one, and the trigger is discharged by giving
receipts the binding reviews already have.

That matters for the row's rung honesty rather than its prose: a trigger whose
capability is demonstrably implemented twenty lines of config away is a tracked
stall someone can close, where the same sentence without the receipt is a wish.

Found while confirming #10036's merge criteria -- the failed review in that
listing is not noise, it is the mechanism working.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
@gunbai-bot gunbai-bot Bot changed the title Fix #9886's two compiler_tests failures on main, THEN promote rust-unit-tests into the required aggregate (operator-approved 2026-09-02) File two recurring failure classes found in this session, and give positional_citation its edit-form instance Sep 2, 2026
…ct at the AUTHORITY and derive both projections

Three conflicts, and only ONE of them was mine to resolve by hand.

THE AUTHORITY CONFLICT WAS PURELY ADDITIVE: my two rows against main's three
(ambient_process_state_read_by_a_concurrent_reader,
predicate_vacuously_true_on_an_empty_domain,
check_subject_narrower_than_its_declared_claim), appended at the same point.
Resolution keeps both sides in both regions.

CHECKED FOR THE SEVERING HAZARD FIRST, because this carrier is the specimen for
`merge_region_excludes_shared_tail`: a multi-line unit ending in a shared
`evidence: [],` / `}` tail makes the naive both-sides resolution leave one row's
closing lines belonging to the next. Both regions here are self-contained -- the
preceding row's `}` sits above `<<<<<<<` and main's last `}` above `>>>>>>>` --
so no tail was factored out. My own rows are one-line units, which is that row's
prescribed repair.

VERIFIED BY IDENTITY JOIN, NOT BY COUNT. Equal counts is exactly the check that
row warns stays green over a severed unit, so all three axes were joined:
  declared 56, identity-fields 56, rostered 56
  declared not rostered: none / rostered not declared: none
  declaration name != its own identity field: none
56 = main's 54 + my 2, which was the post-condition stated before resolving.

THE TWO GENERATED PROJECTIONS WERE NOT HAND-RESOLVED. DESIGN.md and
docs/design-ledgers.md are derived from the authority; their conflict markers were
discarded and both files re-emitted by main_wet over the MERGED .dag, then
confirmed at the fixed point by a second run leaving nothing unstaged. Taking
either side of a generated file is how a projection ends up asserting something
its authority does not say.

AND THE EXPIRING-RECEIPT ROW NOW CARRIES ITS OWN BASE, which is the row applying
its own rule to itself. It cites the review runner's freshness refusal as evidence
that its trigger's capability is already implemented. That citation was taken four
trees ago, so it is now split into what is durable and what is not:
  DURABLE -- review 58587 is retrievable from the reviews API with status=failed,
  sha c80c9d8 and its error text. A completed review is a historical fact.
  OBSERVATION AT A TIME -- that the runner STILL behaves this way. Measured
  2026-09-02 and not re-confirmed, because the runner's source lives outside this
  repository and is not visible from a session container. Its absence from a local
  filesystem is not evidence either way, and six later PRs showing no freshness
  failure is expected rather than reassuring: the check fires only when a head
  moves mid-checkout.
That uncheckability is now stated as the REASON the trigger names a capability
rather than an artifact -- a trigger pointing at a file this repository cannot see
would be satisfied or falsified by nothing observable here, which is the
restoration-promise failure one row over.

Merged rather than rebased: merge policy prefers a merge commit, and the dashboard
notice asking for a rebase is the one instruction in it I did not follow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 2, 2026 13:25
gunbc-ci-auto-heal and others added 2 commits September 2, 2026 13:40
…ojection (review 58732)

REQUEST_CHANGES from codex/gpt-5.6-sol, verified against DESIGN §6 and accepted in
full. §6: "Name the instrument, never transcribe its output... never by copying
its numbers into prose... a transcribed number is unreachable from the thing that
owns it, so it rots without anyone touching either end."

TWO SITES, AND THEY WERE NOT THE SAME DEFECT, which is why they get different
repairs.

1. THE SPECIMEN'S DIFFSTAT. The row said #10024 changed the authority by 151 lines
and the artifact by 140. Those describe two immutable commits, so they cannot rot
in §6's stated sense -- but §6's other half applies squarely: "if a measurement is
worth re-deriving it is worth an entry point." It now names the producer,
`git show --stat 7c310ed -- <the two paths>`, and drops the numbers. The row
keeps its force by saying the change was MATERIAL rather than by quantifying it,
and the reader gets a command that answers the same forever.

2. THE SIX-PR SAMPLE. This one was a genuine §6 violation with no mitigation: a
live population, no named producer, and a count doing no work. The row already
conceded the sample discriminates nothing -- the freshness check fires only when a
head moves mid-checkout, so a sample drawn where the condition never arose is
uninformative by construction. The count is now gone, the producer
(`dashboard-ops reviews <pr>`, reading status and error) is named, and the reason
the sample is uninformative is stated as the point rather than as a caveat.

The generated projection at docs/design-ledgers.md carried the same text because
it is derived; it was re-emitted from the authority rather than edited, and the
transcribed figures are gone from both.

WHAT REMAINS IN THE ROW IS IDENTIFIERS, NOT MEASUREMENTS -- PR numbers, review
58587, a commit sha, dates. Those are names, which is what §6 asks a citation to
carry.

NOTED FOR SOMEONE, NOT ARGUED HERE: 42 of this carrier's 56 rows embed measured
figures in their specimens. If §6 binds a RecurringFailureMode specimen as
strictly as it binds ordinary prose, that is a corpus-wide finding with its own
owner rather than a defect unique to this row. This row complies either way,
because complying cost nothing and produced better prose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Verified against DESIGN §6 and accepted in full — fixed in d1a... (pushed). Both sites, plus the generated projection at docs/design-ledgers.md that carried the same text because it is derived.

They were not the same defect, so they get different repairs.

1. The specimen's diffstat (by 151 lines / 140). These describe two immutable commits, so they cannot rot in §6's stated sense — git show --stat on a fixed sha answers the same forever. But §6's other half applies squarely: "if a measurement is worth re-deriving it is worth an entry point." The row now names the producer — git show --stat 7c310ed74c5 -- dag/gunbc/witness/witness_floor_workflow.dag .github/workflows/witnesses.yml — and drops the figures. It keeps its force by saying the change was material rather than by quantifying it.

2. The six-PR sample. This one was a plain §6 violation with no mitigation: a live population, no named producer, and a count doing no work. The row already conceded the sample discriminates nothing — the freshness check fires only when a head moves mid-checkout, so a sample drawn where the condition never arose is uninformative by construction. The count is gone, the producer (dashboard-ops reviews <pr>, reading each review's status and error) is named, and the reason the sample is uninformative is now the point rather than a caveat.

What remains in the row is identifiers, not measurements — PR numbers, review 58587, a commit sha, dates. Those are names, which is what §6 asks a citation to carry.


One observation offered as information rather than as a defence, since I complied either way: 42 of this carrier's 56 rows embed measured figures in their specimens (admitted_module_without_judged_standing carries 4,419 admitted, 947 build-judged, 2,158 witnesses-judged; disagreement_census_blind_to_agreed_wrong carries 76 rows: 74 Agrees, 2 IdentityUnavailable). If §6 binds a RecurringFailureMode specimen as strictly as it binds ordinary prose, that is a corpus-wide finding with its own owner rather than a defect unique to this row — and worth someone's attention in either direction, because the alternative reading is that a class row's specimen is exactly the place a measured incident should be recorded. I have not acted on that either way; this row now complies under the strict reading.

— sent from sunny-gull-270

…derived

Second additive roster merge on this branch in one session -- #10053 landed
`recurrence_ledger_scoped_below_the_recurrence` while this PR was waiting on its
own checks, so the merge failed on ORDERING rather than on readiness. Both sources
had agreed the PR was ready; main simply moved between the readiness call and the
land.

Only the roster LIST conflicted this time; the declarations auto-merged. Kept both
sides, main's row first.

VERIFIED BY IDENTITY JOIN AGAINST A POST-CONDITION STATED FIRST: main declared 57,
so 59 was owed. Measured declared 59, identity-fields 59, rostered 59, with
declared-not-rostered, rostered-not-declared and name-disagrees-with-its-own-
identity all empty. Equal counts alone would not have caught a severed unit, which
is the failure `merge_region_excludes_shared_tail` records for this exact carrier.

DESIGN.md and docs/design-ledgers.md were NOT hand-resolved. Both are derived; the
markers were discarded and both re-emitted by main_wet over the merged .dag, then
confirmed at the fixed point by a second run leaving nothing unstaged. All three
new rows -- main's one and my two -- verified present in the projection by content
rather than by the merge reporting success. Taking the ours side of a generated
file has silently dropped authority-derived bytes twice on this repo.

Merged, not rebased: the dashboard notice asks for a rebase and merge policy
forbids it on a squash-merge repo. Same divergence as the previous merge, flagged
for the same reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VSP89XiSm2YMnUvSwSR1ct
@gunbai-bot
gunbai-bot Bot merged commit cfb448d into main Sep 2, 2026
6 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/sunny-gull-270-harm-axis branch September 2, 2026 15:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants