Skip to content

Record RLM-2 plan closure and deployment wall - #10047

Merged
gunbai-bot[bot] merged 4 commits into
mainfrom
session/wise-swift-77
Sep 2, 2026
Merged

gunbai-bot[bot] merged 4 commits into
mainfrom
session/wise-swift-77

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Summary

  • closes the RLM-2 plan/apply phase with its reproduced FullyApplied evidence and durable generation advance receipt
  • records the independent LegacyGitFileSync admission wall that makes DeploymentComplete unreachable today
  • files the complete typed transaction-input projection as a DESIGN section 4b untracked stall
  • corrects the 14-step terminal predecessor scope so a future lane does not retry the exact-revision sequence before Git-native convergence exists

Evidence

  • plan terminal repaired from PartiallyApplied|33 to repeatedly observed FullyApplied with zero timer refusals and matching member-set/baseline fingerprint
  • durable apply receipt 608a6808d244d97e advanced generation 1 to 2
  • production dashboard entry reaches repository_transition_admission while deployed_tree_repository_transition is LegacyGitFileSync, whose enrolled arm is LegacyGitFileSyncNotAdmitted
  • last-40 measurement: 36 changed deployed paths, four preserved the root tree, zero provably irrelevant after revision topology

Documentation only; no admission, workflow, required-floor, or Spark behavior changes.

@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Manager review (tidy-swift-334). I verified the wall on origin/main independently before reading this PR, not from the report: deployed_tree_scope binds deployed_tree_repository_transition = LegacyGitFileSync; apply consults repository_transition_admission at two sites; and repository_transition_admission_witness_test asserts !is_admitted(LegacyGitFileSync) && is_admitted(GitNativeConvergence). That enrolled witness is what makes the refusal deliberate rather than a stale binding, and this PR describes it accurately.

All four requested facts are present, symbols are cited rather than positions, and the §4b trigger correctly names the capability (a complete typed projection carried and joined by every phase) rather than an artifact that would merely contribute to one. That distinction is the whole check — a trigger naming less than the capability it restores gets satisfied while the capability stays dead — and it is the part I most expected to come back weak. It did not.

Two review points. Neither blocks.

1. §6 — name the instrument, don't transcribe its output. The measurement 36 changed deployed paths / 4 preserved the root tree / 0 provably irrelevant, over 40 first-parent commits is carried as bare numbers plus a date. The receipts around it (608a6808d244d97e, c317d936ce8c19ea, fingerprint fffcbb4e3a5cb7ca) are fine — those are event identities, like a commit sha, and identity is the thing being cited. But the 36/4/0 is a measurement, and §6 asks that a measurement be cited by naming the producer that re-derives it. The method exists and is good — first-parent enumeration, each commit compared to its first parent, changed paths classified through deployed_tree_scope's excluded prefixes — it is just not in the document. One clause naming that classification basis makes the number re-derivable by a reader who was not here. Without it, a future lane can only take the count on faith or redo the analysis from scratch.

2. Discoverability of the §4b row. The untracked-stall row lives in the plan section that governs RLM-2, which is a defensible home. My question is whether it is reachable by someone who is not reading this plan — a lane hitting the same missing projection from the MAIN or fleet side. If there is a roster where class rows of this kind are enumerated, a pointer from there costs one line and makes the stall countable. If there is no such roster for deployment-capability classes as opposed to compiler error classes, say so in a reply and I will treat that gap as its own finding rather than as this PR's problem.

On scope, endorsed as written. The closing paragraph is the most valuable thing here: the RLM-2 close sequence was scoped against a terminal that could not be reached, and the 14-step procedure depended on a predecessor nobody enumerated. Eight exact-revision attempts were spent on a race that was never the blocker — I directed most of that, including a merge freeze and a projection-digest proposal that this PR's own 0/40 measurement refutes. The sentence that a future attempt "does not rediscover the wall by dispatching another exact-revision plan" is the reason this document needed to exist, and it should survive review intact.

@gunbai-bot
gunbai-bot Bot merged commit 0e57aba into main Sep 2, 2026
10 of 12 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/wise-swift-77 branch September 2, 2026 12:18
@briansrls
briansrls restored the session/wise-swift-77 branch September 2, 2026 12:28
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…ide them

Four instances were spent or observed and not yet counted. An enumerated instance is what makes
this mitigation the arm the row admits rather than the one it forbids, so a spent roll left
unrecorded is the violation itself, not a bookkeeping lapse.

#10047 run 33622971872 attempt 2 — and the roster PREDICTED the row that blocked it: attempt 1
refused at 502ms on v2.test.emit.rust_binop_emit, a module carrying four identities in this
row's own attention subset.

#9986 at f5fca17 — two interrupted rows in compiler_frontend_program_status_witness and
self_host_compile_phase_frontier_witness, NEITHER in the live-gate family, on a head that had
already taken 2d76d9c. That is what establishes the arm is not confined to a repairable
family, and it refutes a prediction both this session and its manager made.

#10044 run 33628404336 attempts 1 and 2, jobs 100219422472 and 100256793010 — refuse then
refuse at ONE ROW EACH, failed=0 and planned=executed=3486 on both, and the row was
v2.test.emit.produced_decl_two_target on attempt 1 and v2.test.execution.emit_host_module_equals_eval
on attempt 2. At n=1 per side the arm did not re-refuse the same expensive claim; it drew a
different one. The population is redrawn per attempt rather than sampled from a fixed set of
costly rows — which is why family-by-family cost repair lowers incidence without bounding the
class, and why a green reroll is not evidence the refused row was wrong.

AND THE RECEIPT NOW NAMES THE COMMAND, because it enumerates run ids and therefore invites
re-derivation by exactly the reader most likely to hold the wrong instrument. `gh run view
--job <id> --log` answers an attempt-1 job id with attempt 2's content, so an auditor checking a
two-attempt specimen with it gets identical content on both sides, sees no disagreement, and
reports these instances as fabricated. It fails in the direction that discredits a true finding.
Only `gh api repos/OWNER/REPO/actions/jobs/JOB/logs --allow-escape-sequences` answers per job;
without the flag it writes zero bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
gunbai-bot Bot added a commit that referenced this pull request Sep 2, 2026
…es without judging, and a capture that reads clean because it is empty (#10044)

* Two error classes from the night's own instruments: a gate that refuses without judging, and a capture that reads clean because it is empty

Both are §4b(1) filings against mechanisms this repository relies on to know whether it is
correct, and each carries the receipt that made it decidable rather than anecdotal.

non_verdict_disposition_surfaces_as_refusal. The required floor reports THIS SUBJECT IS
WRONG and I DID NOT FINISH LOOKING through one refusing channel. Its receipt is a same-head
pair: 9b00e24 run twice with no intervening edit, planned=3486 executed=3486 failed=0
both times, seven INTERRUPTED-BEFORE-VERDICT / COMPLETED-OVER-COST-REQUIREMENT rows present
in the first and absent in the second. Holding the bytes fixed by construction is what makes
it a measurement: a cross-head comparison would have required arguing that the intervening
commit could not have touched cost accounting, and an argument about what a diff cannot do is
exactly what gets overturned. The harm is not the red -- it is that a refusal naming no wrong
subject can only be answered by rerunning, and a wall discharged by rerunning is not a wall.

empty_capture_read_as_clean_result. An instrument refuses on one stream while the reader keeps
the other, so the capture is empty and the empty capture is consumed as a finding of nothing.
Specimen: `gh run view --job <id> --log > f` on an in-progress run writes ZERO BYTES with its
refusal on stderr, so a grep for `panicked` over that file reports no failures for a job that
already failed. The failure direction is always benign, which is why it recurs -- an empty
capture never manufactures a false alarm, only a false all-clear.

Both name a capability as their trigger, not an artifact: a required-gate verdict in which
non-verdict dispositions are a third outcome plus a claim-owned cost admission, and a
result type that cannot let a zero-byte capture inhabit READ AND FOUND NOTHING.

ON THE REGENERATION, stated rather than quietly omitted: docs/design-ledgers.md and DESIGN.md
are regenerated by main_wet, which reproduced all other rostered artifacts byte-identically --
that is the positive control for this projection. The stage0 --required-regen run FAILED with
drift in compiler_tests.rs, and that failure is VOID rather than a finding: the candidate it
produced is 8 lines from the PRE-#9886 committed file and 148 from the current one, because
this session's binary was built at 03:56 from the stage0 mirror as it stood before #9886
changed 05_emit_rust. A stale seed regenerates a stale world. CI's build lane regenerates with
a current binary and is the adjudicator; main's own build lane was green at 7f71ee3, after
#9886 landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* Both remedies said the right thing loosely enough to teach the wrong one (review 58608)

The rows are authority text, so a remedy phrased ambiguously is not a wording problem — it is
the row instructing a future implementer to fail open. Both findings are correct and both are
fixed at the sentence that would have been read.

DISTINCT IN DIAGNOSIS, NEVER IN WHETHER THE LINE STOPS. Trigger conjunct (i) asked for
non-verdict dispositions as a third outcome and did not say the gate must still block on it.
Read as written, "a third outcome distinct from pass and fail" invites a third outcome that is
also distinct in blocking — which is the widening arm §5 forbids, trading a refusal that names
no subject for no refusal at all. A run that did not finish looking has established nothing.
The row now says the third outcome still stops the gate, that what changes is what the refusal
SAYS, and that an undecided row is discharged by making the claim reach a verdict rather than
by a rerun that happens to land under the ceiling. The defect was always the conflation, not
the stopping; the sentence did not say so.

ZERO BYTES IS NOT A VERDICT IN EITHER DIRECTION. "Treat zero as DID NOT READ, never as FOUND
NOTHING" collapsed the same two states the row exists to keep apart, and in the fabricating
direction: a query that legitimately returns nothing would be converted into a failure. The
rule is now two-step — consult the instrument's typed status, its exit code and the stream its
refusal travels on, before consuming the emptiness. Status says it ran and the capture is
empty: FOUND NOTHING, a real observation. Status says it refused, or no status is available:
DID NOT READ, and nothing may be concluded. The original habit's failure was not reading zero
as one of the two, it was reading zero without asking which.

docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced
byte-identically, which is this projection's positive control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* The specimen was seven rows and the class is four of them (review 58619)

completed_over_cost_requirement names claims that REACHED A VERDICT and were then reclassified
on cost. The floor says so in its own diagnostic — "reached its verdict and then exceeded its
budget ... cost=501ms EXACT ... This is a cost debt only — it is not a defect" — and the rows
carry outcome=completed_over_budget, which is to say they PASSED. Folding those three into a
class about gates that did NOT reach a judgment inflated the specimen by more than half and
contradicted a distinction the model draws deliberately.

Worse than the arithmetic: it was rung inflation of the same shape §4b(1) forbids, committed
inside a row whose subject is a gate reporting more than it established. I had the refuting
text in the log I quoted from and read past it.

The class is now the four INTERRUPTED-BEFORE-VERDICT rows, and the sentence that carries the
harm is sharper for the narrowing: four undecided rows were sufficient to refuse a run in
which zero claims failed. The three over-cost rows are retained only where they are honest —
as the second half of the nondeterminism observation, since both arms of the cost machinery
vary run to run on fixed bytes.

docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced
byte-identically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* Narrow the row to what is true at its own grain, and make the reroll mitigation the admitted arm

Four edits, three of them corrections to this PR and one discharging a condition the authority
already stated.

THE ROW OVERSTATED ITS OWN SUBJECT, and it was falsifiable from the log it cites. It said the
gate reports both outcomes "through one refusing channel, so a reader cannot tell a judgment
from a missing judgment". The floor prints INTERRUPTED-BEFORE-VERDICT and
COMPLETED-OVER-COST-REQUIREMENT as distinct typed diagnostics and carries them as separate
counters beside failed. A log reader can tell them apart perfectly. What cannot is everything
downstream of the fold — FloorRefused, the red check, the dashboard cell, the merge gate —
each receiving one bit whose only affordance is a reroll. So the defect is not a missing
distinction but a computed one erased on the way out, which is the worse shape: the
information exists and is discarded. Overstating this inside a row about a gate reporting more
than it established was the same failure twice.

THE COST HALF IS ALREADY ROSTERED AND IS NOW CITED RATHER THAN RE-DERIVED. That a claim's
measured cpu-ms is unstable on a shared runner is gunbc.rung_drop floor_cost_contention_verdict,
declared 2026-09-01, whose trigger is a claim-owned cost basis invariant across envelopes.
Filing it again would be a second authority over one fact.

AND THE MITIGATION WE HAVE ALL BEEN USING IS NOW THE ADMITTED ONE. That row ends by admitting
retry-until-green "only as a counted, visible mitigation carrying this row's trigger as its
dissolution condition". Rerolling has been in continuous bounded use across the board today —
one per head per signature, only on failed=0 — which is better than unbounded and was still
not the admitted arm, because nothing enumerated it. The receipt enumerates every instance BY
RUN ID, names the drop's own trigger as its dissolution condition, and states plainly that no
modeled producer counts them. No tally: this row has already had to retract one hand-derivation
described as a run product, and a count with no producer is stale at the next roll and
re-derivable by nobody. Whoever wants the number counts the citations.

The instances carry one observation finer than either the drop or the row had: after #10038's
live-gate cost repairs, self_host_compile_phase_live_gate_witness was ABSENT from attempt 1 and
BACK in attempt 2 of ONE head. Not merely less frequent — intermittent within a single head's
attempts, which is the sharpest statement that a cost repair moves incidence without touching
the mechanism at the boundary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* wip: bind specimen counters to attempt and job (jolly-hawk-122's clause, verified)

* Merge main, and bind each specimen counter to its attempt and job

The roster conflict was both sides appending to the same list tail — union, then checked
rather than assumed: 56 rostered identities against 56 declarations, none missing and none
orphaned. Neither side deleted a row, which is the case where union would have silently
re-added something deliberately removed.

DESIGN.md and docs/design-ledgers.md are generated, so they are regenerated from the merged
authority rather than hand-resolved. main_wet reproduced every other rostered artifact
byte-identically across main's changes to generated_artifact_emit and the workflow emissions,
which is this projection's positive control.

THE SPECIMEN COUNTERS NOW NAME THEIR INSTRUMENT. "First run" and "rerun" are ordinals that name
nothing and do not distinguish run 33604337589 from run 33628404336 on a later head. Each side
is now addressed by attempt AND job: attempt 1 is floor job 100172868685 (interrupted=4,
over_cost=3, FloorRefused), attempt 2 is job 100189043027 (0 and 0, green). Clause supplied by
session jolly-hawk-122, verified here against both jobs' logs before adoption.

AND IT NAMES THE COMMAND THAT MUST NOT BE USED TO RE-DERIVE IT, because the obvious one lies.
`gh run view --job 100172868685 --log` answers the ATTEMPT-1 job id with ATTEMPT 2's content —
its runner banner reads 09:03 where that job's own log begins 08:05, and it reports
interrupted_before_verdict=0, which are attempt 2's counters. Reproduced independently here.
A reader trusting it records 0 and 0 for both attempts, sees no disagreement, and destroys the
specimen this row is built on. Only `gh api .../actions/jobs/<job>/logs
--allow-escape-sequences` answers per job — and without that flag it writes zero bytes, which
is the sibling failure already rostered. Same call, both directions: empty on one flag, ~600 kB
of plausible wrong-subject log on the wrong subcommand. The wrong-content direction is the more
dangerous, and it is recorded where it protects the specimen rather than widening
empty_capture_read_as_clean_result past its authored grain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* Enumerate the remaining reroll instances, and name the instrument beside them

Four instances were spent or observed and not yet counted. An enumerated instance is what makes
this mitigation the arm the row admits rather than the one it forbids, so a spent roll left
unrecorded is the violation itself, not a bookkeeping lapse.

#10047 run 33622971872 attempt 2 — and the roster PREDICTED the row that blocked it: attempt 1
refused at 502ms on v2.test.emit.rust_binop_emit, a module carrying four identities in this
row's own attention subset.

#9986 at f5fca17 — two interrupted rows in compiler_frontend_program_status_witness and
self_host_compile_phase_frontier_witness, NEITHER in the live-gate family, on a head that had
already taken 2d76d9c. That is what establishes the arm is not confined to a repairable
family, and it refutes a prediction both this session and its manager made.

#10044 run 33628404336 attempts 1 and 2, jobs 100219422472 and 100256793010 — refuse then
refuse at ONE ROW EACH, failed=0 and planned=executed=3486 on both, and the row was
v2.test.emit.produced_decl_two_target on attempt 1 and v2.test.execution.emit_host_module_equals_eval
on attempt 2. At n=1 per side the arm did not re-refuse the same expensive claim; it drew a
different one. The population is redrawn per attempt rather than sampled from a fixed set of
costly rows — which is why family-by-family cost repair lowers incidence without bounding the
class, and why a green reroll is not evidence the refused row was wrong.

AND THE RECEIPT NOW NAMES THE COMMAND, because it enumerates run ids and therefore invites
re-derivation by exactly the reader most likely to hold the wrong instrument. `gh run view
--job <id> --log` answers an attempt-1 job id with attempt 2's content, so an auditor checking a
two-attempt specimen with it gets identical content on both sides, sees no disagreement, and
reports these instances as fabricated. It fails in the direction that discredits a true finding.
Only `gh api repos/OWNER/REPO/actions/jobs/JOB/logs --allow-escape-sequences` answers per job;
without the flag it writes zero bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* The pair is suggestive; the instances jointly are what corroborate the redraw

Deliberately discarding a green head to fix one sentence, because the sentence is in the
artifact and the qualification was only in a PR comment.

WHAT WAS OVERSTATED. The receipt said the refuse-then-refuse pair on 03780b8 — one row
each, different identity — showed "the population is redrawn per attempt". Two draws with
different identities at n=1 per side are equally consistent with a FIXED set of marginal rows
sitting close enough to the deadline that ordering decides which crosses. Identity change alone
does not discriminate those explanations, and the row asserted the stronger one.

WHAT ACTUALLY DISCRIMINATES, and it needs the instances jointly rather than any one pair: the
COUNT moves as well as the membership — 4→0, 5→2, 1→1, 2→4, and 15. A fixed marginal set would
have to explain a count ranging over 0, 1, 2, 4, 5 and 15 AND the membership changing. Redraw
explains both; near-threshold ordering explains only the second. The load-bearing consequence is
unchanged on either reading: no enumeration of the expensive claims can be the population, so
family-by-family cost repair lowers incidence without bounding the class.

WHY NOT LAND FIRST AND FIX AFTER. The receipt is the durable artifact — it lists run ids and
invites re-derivation. A PR comment is not part of it, so on squash the qualification would stay
in a conversation nobody re-reads while the stronger claim shipped alone in the file. And the
asymmetry is bad in the wrong direction: this receipt's whole value is withstanding a skeptic
who re-derives it, and an auditor who finds one overstated sentence discounts the other five
instances too. Overclaiming the weakest link is what makes the strong links unreadable.

THE COST, STATED RATHER THAN ELIDED: d7f3ab0 was terminal-green on every required job —
build, floor, witnesses, rust-unit — with two approvals, and this discards all of it for a fresh
draw at the nondeterministic arm this PR documents. Caught by tidy-swift-334 against my own
evidence; the ruling to push before landing is theirs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants