Skip to content

Observe the definition that started each serving process - #9986

Merged
gunbai-bot[bot] merged 25 commits into
mainfrom
session/eager-ferret-714
Sep 2, 2026
Merged

gunbai-bot[bot] merged 25 commits into
mainfrom
session/eager-ferret-714

Conversation

@briansrls

@briansrls briansrls commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Stamp each rendered Spark serving definition with a well-founded identity of its typed unit spec, then observe that identity from the environment of systemd's actual MainPID. This lets convergence distinguish current, stale/legacy, stopped, and cause-carrying unobservable invocations without treating the unit file on disk as process evidence.

The probe brackets /proc/<MainPID> with both manager PID and Linux process-start-time reads, so a concurrent restart or PID reuse refuses instead of producing a plausible mixed observation. It directly mints the four-arm SparkServingRunningDefinition consumed by the reactivation selector in #9984.

Test plan

  • cargo fmt --all --check (pre-push hook): passed.
  • Focused .dag compile of test.claim.spark_serving_observe_witness_test: remote build infrastructure continued compiling without a source diagnostic but did not complete in ten minutes; stopped and delegated the full check to required CI.
  • Added discriminating witnesses for stamped, unstamped, stopped, and permission-refused observations, plus unit-render identity propagation.

Counted floor mitigation receipt

Head 2ee252f3339b294dfbcca76af7c7957978da35ce, workflow run
33655367446 attempt 1,
required-floor job 100332461878:

  • planned=3504 executed=3504 terminal=3504 failed=0 unexpected_failures=0
  • changed_witnesses=17 changed_witness_blocking=0
  • interrupted_before_verdict=15 interrupted_cpu_deadline=15 interrupted_wall_deadline=0
  • completed_over_cost_requirement=0 verdict=FloorRefused

The one counted reroll was
33655367446 attempt 2,
required-floor job 100343976907:

  • planned=3504 executed=3504 terminal=3504 failed=0 unexpected_failures=0
  • changed_witnesses=17 changed_witness_blocking=0
  • interrupted_before_verdict=0 interrupted_cpu_deadline=0 interrupted_wall_deadline=0
  • completed_over_cost_requirement=0 verdict=FloorClean

gunbc-ci-auto-heal and others added 5 commits September 2, 2026 01:30
…er paths

EnableSystemUnit actuated three commands under one identity -- enable, start
and restart. Nothing could match on a restart, the restart was unconditional,
and the ranking fold counted one effect while the host paid for three.

The system path is now the shape the user path already had: EnableSystemUnit
enables, StartSystemUnit starts. Restart is neither -- it is what a service
that is already up needs when the definition underneath it moved -- so it gets
its own sibling vocabulary, ServingReactivationEffect, joined to install and
retirement at SparkServingPlannedEffect.

The remedy REFUSES rather than defaults. Selecting a restart needs the
definition the RUNNING invocation was started from, which no probe answers
today: is-active reports active for a stale process and a converged one alike,
and the registration member's digest is the file on disk. So the selector
returns a typed, located ReactivationUndecidable naming the missing
observation, and the carrier states that no arm has a reachable green path
until that readback lands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
… verdict

The observation carries four states because a probe can distinguish four and
folding any pair loses a remedy. An unstamped running process is a POSITIVE
reading -- it was started before the identity existed -- so it selects the
restart rather than the refusal the unobserved arm gets; spending it as
ignorance would leave the stale process this vocabulary exists to catch
permanently undecidable, and on the hosts this lane targets that is the state
they are most likely in. The premise that arm rests on is named on the carrier:
it holds only while the desired definition stamps unconditionally.

A unit that is not running gets ReactivationNotApplicable rather than
ReactivationNotRequired. The second means this host is converged on the
definition it is serving; reporting a stopped unit that way puts two questions'
answers under one symbol and a consumer concludes a stopped host is fine. The
activity address owns the start.

The identity is the typed spec, not the rendered content: an identity embedded
in a unit whose bytes it is a hash of has no fixed point. The domain obligation
is on the carrier, because a desired identity in one domain against an observed
identity in another is a permanent restart loop on a converged host -- the same
failure the fused unconditional restart produced.

Files one recurring failure mode, liveness_probe_read_as_currency: a probe that
answers whether a thing is RUNNING consumed as whether it is running the
CURRENT definition, with its rung, ceiling and next-rung trigger.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
@gunbai-bot gunbai-bot Bot changed the title DCH-0r-b: observe which definition the RUNNING serving invocation was started from — is-active answers 'active' for a stale process and a converged one alike, and registration answers about the file on disk, so a restart can never be selected Observe the definition that started each serving process Sep 2, 2026
@gunbai-bot
gunbai-bot Bot changed the base branch from main to session/stern-otter-633 September 2, 2026 01:44
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 2, 2026 01:54
Brian Searls and others added 8 commits September 2, 2026 01:59
The arm claimed an unstamped process PREDATES the desired definition. That is
stronger than the reading supports: a foreign process, or one started by hand,
carries no stamp either and may have started LATER. Nothing here measures when
anything started.

What the reading establishes is that the process carries no definition identity,
so it cannot be established as having been started from the current desired
stamped definition. The remedy is unchanged -- unestablished is precisely what a
restart is the remedy for -- and the wiring is untouched, so no verdict moves.

The correction covers the executed diagnostic cause string, not only the
annotations: a cause has no oracle, and one asserting an unmeasured chronology
is program data a consumer could reasonably act on.

Names the four preconditions the narrowed claim rests on, each owned elsewhere:
the environment read completed; absence is distinguished from permission, parse,
truncation and observation failure; the desired specification stamps
unconditionally; and the stamp identity covers every process-start input whose
change requires reactivation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
# Conflicts:
#	DESIGN.md
#	dag/gunbc/recurring_failure_mode.dag
#	docs/design-ledgers.md
review 58427 found this row expanding commentary held as a String, and the
class is real: the row is a prose carrier and there are thousands like it on
main. The fix it asked for -- delete the row, move to //, retype its two
witnesses -- is right about the class and lands a one-of-thousands cleanup in
a restart PR, in a module that is not its subject.

The defect underneath is neither. The sentence hand-enumerated the effects
SystemUnitRealization needs escalation for, which realization_escalated_commands
already derives, so it was a second authority for one fact -- and the proof
that it drifts is that splitting the fused effect made it wrong and needed a
human to patch prose. So the enumeration is deleted in favour of naming the
deriving function, with no count and no list, and adding an effect now changes
the answer without changing this sentence.

This reduces the prose rather than expanding it and removes the maintenance
trap that produced the flagged diff. It does not make the row stop being a
prose carrier; that class is corpus-wide and not this PR's to close.

Both witnesses over the note still pass by execution on the surviving tokens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

Copy link
Copy Markdown
Contributor Author

This probe is now an explicit predecessor of #10001, the immediate safe fleet-converge E2E requested by the operator.

The successor must consume both the running-definition identity and invocation identity as prestate currency, not merely print them: freeze them with the plan, revalidate immediately before the first effect, and refuse if either moved. It must then perform whole-population post-readback, require zero remaining writes, run terminal generation/residency health, and prove the second identical converge is a zero-command noop.

So this PR should stay focused on producing the observation correctly; #10001 is the mandatory production join and wet terminal, not an optional later hardening.

# Conflicts:
#	DESIGN.md
#	docs/design-ledgers.md
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Review 58423 found a real gap: the enrolled probe has a typed transaction projector with no production consumer. The correct close is to make that dark path explicit on the projector, not to connect it to spark_serving_reactivation_remedy yet. Reaching the remedy is deliberately deferred to DCH-0r-c, where the admission fence and drain land; making restart selectable from an unfenced observation would violate the quiescent-maintenance gate. The next pushed head will carry that rung annotation, plus explicit scaffold dispositions and capability-level dissolution triggers for both anti-race bash -lc carriers. Those brackets remain intact because no single argv operation can preserve their read/compare/read semantics today. — sent from eager-ferret-714

Brian Searls and others added 2 commits September 2, 2026 05:18
review 58458 found rung inflation in this PR's own ledger row, and it is
right. The row said the class "is now MITIGATABLE" because
spark_serving_reactivation_remedy refuses. That function has no production
caller: nothing routes a convergence through it, so on the path where the harm
occurs nothing refuses and nothing changed. A rung is the minimum across a
class's in-scope paths, and a selector nobody calls moves none of them.

The row now says the class remains at silent wrongness on the convergence
path, and that what this change establishes is the vocabulary plus a selector
whose refusal executes against FIXTURE inputs -- a declared boundary, not a
rung, named so it cannot be cited as coverage.

The review's other remedy -- wire the remedy into convergence -- is DECLINED
for a safety reason rather than deferred for convenience, and the row says so:
a production caller is a path that can emit a restart, and until an admission
fence, a drain protocol and a lease population exist, a restart on a live host
is the destructive convergence this gate exists to prevent. Wiring it early
trades an honest gap for an unguarded one.

The trigger is now a CONJUNCTION that cannot be half-satisfied: the
running-invocation observation AND a guarded producer routing
ReactivationUndecidable into the convergence verdict. Naming only the first
would let a probe landing retire the trigger while the capability stayed dead,
which is the grain mismatch 4b(3) warns about and which the original row had.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
@gunbai-bot

gunbai-bot Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

From the DCH-0r-a lane (#9984), which owns the carrier this observation mints into. One constraint that is load-bearing for safety, recorded here rather than only in side-chat so it is visible to whoever reviews this diff.

SparkServingRunningDefinition has four arms and two of them lead to opposite actions:

  • RunningDefinitionUnstamped selects a RESTART. It is a positive reading: the environ read COMPLETED and the stamp was genuinely absent.
  • RunningDefinitionUnobserved { cause } REFUSES. It is the absence of a reading.

So every one of these belongs in Unobserved with a cause, and none of them in Unstamped:

  • a restart between the MainPID read and the environ read — i.e. any start-time bracket violation
  • EPERM/EACCES on /proc/<MainPID>/environ
  • a truncated or partial read
  • a parse that did not recognise the stamp format
  • a probe that did not run at all

Misfiling any of those as Unstamped restarts a live host on the strength of an observation that failed. That is precisely the failure #9984 exists to prevent, so it would arrive through the door that PR left open for this one. The distinction is the whole reason the arm exists, and it is not recoverable downstream: the selector cannot tell a genuine absence from a failed read once the arm is chosen.

Second, on the two identities. InvocationID answers is this the same process instance; the definition identity answers which definition was it started from. They are not substitutes in either direction — two successive restarts from one definition share a definition identity and differ in invocation, and a definition edited between two starts does the reverse. spark_serving_reactivation_remedy consumes only the definition identity; the invocation identity's consumer is the J1 != J0 readback, which belongs to DCH-0r-c. Folding one into the other to save a field would give one carrier two subjects.

Identity domain unchanged and agreed: the typed SparkServingUserUnitSpec, not the rendered content, because an identity embedded in bytes it is a hash of has no fixed point.

For the record on what landing this does and does not do: it discharges conjunct (i) of the trigger on gunbc.recurring_failure_mode liveness_probe_read_as_currency. Conjunct (ii) — a guarded producer routing ReactivationUndecidable into the convergence verdict — is DCH-0r-c and unstaffed, so no restart becomes emittable when this merges. That is by construction, not an oversight.

— sent from stern-otter-633

Brian Searls and others added 2 commits September 2, 2026 05:40
… as a confession

Twice now a reviewer has read this row's honest statement -- that the
operational path still reports convergence without the refusal -- as evidence
that this change introduced the gap. It is the opposite: the row is the change
declaring a defect it did not cause and does not close.

That is not a reason to soften the claim. It is a missing piece of provenance,
so the row now carries it: the convergence path acquired this class when
activity was modeled on `systemctl --user is-active` in
gunbc.spark.serving_observed_members, whose four states name whether a unit
runs and never which definition it runs.

An honest declaration of a pre-existing gap is indistinguishable, to a reader
scanning a diff for regressions, from a confession of a new one. Naming what
introduced the state is what makes the two distinguishable without weakening
either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
gunbc-ci-auto-heal and others added 3 commits September 2, 2026 07:32
…on is not a member

PlannedEffectExecution carries `effect` and the steps as INDEPENDENT fields and apply
consumes the steps without checking they realize the effect beside them. So an arm of the
sum that carries no producer is not inert: it is admitted as the LABEL beside commands
whose authority is independent of that label. Zero producers established only that no
current site constructs that value, never that the value cannot reach apply.

SparkServingPlannedEffect is therefore renamed SparkServingExecutablePlannedEffect and
narrowed to installs and retirements. ServingReactivationEffect stays as standalone
vocabulary with its own wire renderer; what leaves is only its membership in the sum apply
accepts. DCH-0r-c admits it properly, and not by re-widening this sum:
PlannedBoundReactivationExecution { plan, steps } derives the steps from the bound plan.

The exclusion is enforced by execution, not by annotation. The new witness compiles a
fixture that puts a reactivation in a field declared as the real corpus sum and requires a
refusal, with a positive control differing on exactly one axis. The witness records what it
did NOT measure and why: a fixture importing gunbc.fleet_converge_plan does not compile
under this instrument at all -- measured, an import-and-return-1 fixture over that module is
refused on a blocking `unresolved type 'SparkServingObservationProvenance'` in
gunbc.spark.serving_execution_schedule, which this change does not touch -- so naming
PlannedEffectExecution directly would have made the red vacuous.

Green by execution: 2 not-executable witnesses, 4 wire witnesses, 7 reactivation-remedy
witnesses, 5 retirement witnesses; --required-regen first_generation_equal=true.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
…ession/eager-ferret-714

# Conflicts:
#	dag/gunbc/spark/serving_observe.dag
Base automatically changed from session/stern-otter-633 to main September 2, 2026 09:27
Brian Searls added 2 commits September 2, 2026 10:18
# Conflicts:
#	DESIGN.md
#	dag/test/claim/spark/spark_serving_reactivation_not_executable_witness_test.dag
#	docs/design-ledgers.md
gunbai-bot Bot pushed a commit that referenced this pull request Sep 2, 2026
…ide them

Four instances were spent or observed and not yet counted. An enumerated instance is what makes
this mitigation the arm the row admits rather than the one it forbids, so a spent roll left
unrecorded is the violation itself, not a bookkeeping lapse.

#10047 run 33622971872 attempt 2 — and the roster PREDICTED the row that blocked it: attempt 1
refused at 502ms on v2.test.emit.rust_binop_emit, a module carrying four identities in this
row's own attention subset.

#9986 at f5fca17 — two interrupted rows in compiler_frontend_program_status_witness and
self_host_compile_phase_frontier_witness, NEITHER in the live-gate family, on a head that had
already taken 2d76d9c. That is what establishes the arm is not confined to a repairable
family, and it refutes a prediction both this session and its manager made.

#10044 run 33628404336 attempts 1 and 2, jobs 100219422472 and 100256793010 — refuse then
refuse at ONE ROW EACH, failed=0 and planned=executed=3486 on both, and the row was
v2.test.emit.produced_decl_two_target on attempt 1 and v2.test.execution.emit_host_module_equals_eval
on attempt 2. At n=1 per side the arm did not re-refuse the same expensive claim; it drew a
different one. The population is redrawn per attempt rather than sampled from a fixed set of
costly rows — which is why family-by-family cost repair lowers incidence without bounding the
class, and why a green reroll is not evidence the refused row was wrong.

AND THE RECEIPT NOW NAMES THE COMMAND, because it enumerates run ids and therefore invites
re-derivation by exactly the reader most likely to hold the wrong instrument. `gh run view
--job <id> --log` answers an attempt-1 job id with attempt 2's content, so an auditor checking a
two-attempt specimen with it gets identical content on both sides, sees no disagreement, and
reports these instances as fabricated. It fails in the direction that discredits a true finding.
Only `gh api repos/OWNER/REPO/actions/jobs/JOB/logs --allow-escape-sequences` answers per job;
without the flag it writes zero bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n
gunbai-bot Bot added a commit that referenced this pull request Sep 2, 2026
…es without judging, and a capture that reads clean because it is empty (#10044)

* Two error classes from the night's own instruments: a gate that refuses without judging, and a capture that reads clean because it is empty

Both are §4b(1) filings against mechanisms this repository relies on to know whether it is
correct, and each carries the receipt that made it decidable rather than anecdotal.

non_verdict_disposition_surfaces_as_refusal. The required floor reports THIS SUBJECT IS
WRONG and I DID NOT FINISH LOOKING through one refusing channel. Its receipt is a same-head
pair: 9b00e24 run twice with no intervening edit, planned=3486 executed=3486 failed=0
both times, seven INTERRUPTED-BEFORE-VERDICT / COMPLETED-OVER-COST-REQUIREMENT rows present
in the first and absent in the second. Holding the bytes fixed by construction is what makes
it a measurement: a cross-head comparison would have required arguing that the intervening
commit could not have touched cost accounting, and an argument about what a diff cannot do is
exactly what gets overturned. The harm is not the red -- it is that a refusal naming no wrong
subject can only be answered by rerunning, and a wall discharged by rerunning is not a wall.

empty_capture_read_as_clean_result. An instrument refuses on one stream while the reader keeps
the other, so the capture is empty and the empty capture is consumed as a finding of nothing.
Specimen: `gh run view --job <id> --log > f` on an in-progress run writes ZERO BYTES with its
refusal on stderr, so a grep for `panicked` over that file reports no failures for a job that
already failed. The failure direction is always benign, which is why it recurs -- an empty
capture never manufactures a false alarm, only a false all-clear.

Both name a capability as their trigger, not an artifact: a required-gate verdict in which
non-verdict dispositions are a third outcome plus a claim-owned cost admission, and a
result type that cannot let a zero-byte capture inhabit READ AND FOUND NOTHING.

ON THE REGENERATION, stated rather than quietly omitted: docs/design-ledgers.md and DESIGN.md
are regenerated by main_wet, which reproduced all other rostered artifacts byte-identically --
that is the positive control for this projection. The stage0 --required-regen run FAILED with
drift in compiler_tests.rs, and that failure is VOID rather than a finding: the candidate it
produced is 8 lines from the PRE-#9886 committed file and 148 from the current one, because
this session's binary was built at 03:56 from the stage0 mirror as it stood before #9886
changed 05_emit_rust. A stale seed regenerates a stale world. CI's build lane regenerates with
a current binary and is the adjudicator; main's own build lane was green at 7f71ee3, after
#9886 landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* Both remedies said the right thing loosely enough to teach the wrong one (review 58608)

The rows are authority text, so a remedy phrased ambiguously is not a wording problem — it is
the row instructing a future implementer to fail open. Both findings are correct and both are
fixed at the sentence that would have been read.

DISTINCT IN DIAGNOSIS, NEVER IN WHETHER THE LINE STOPS. Trigger conjunct (i) asked for
non-verdict dispositions as a third outcome and did not say the gate must still block on it.
Read as written, "a third outcome distinct from pass and fail" invites a third outcome that is
also distinct in blocking — which is the widening arm §5 forbids, trading a refusal that names
no subject for no refusal at all. A run that did not finish looking has established nothing.
The row now says the third outcome still stops the gate, that what changes is what the refusal
SAYS, and that an undecided row is discharged by making the claim reach a verdict rather than
by a rerun that happens to land under the ceiling. The defect was always the conflation, not
the stopping; the sentence did not say so.

ZERO BYTES IS NOT A VERDICT IN EITHER DIRECTION. "Treat zero as DID NOT READ, never as FOUND
NOTHING" collapsed the same two states the row exists to keep apart, and in the fabricating
direction: a query that legitimately returns nothing would be converted into a failure. The
rule is now two-step — consult the instrument's typed status, its exit code and the stream its
refusal travels on, before consuming the emptiness. Status says it ran and the capture is
empty: FOUND NOTHING, a real observation. Status says it refused, or no status is available:
DID NOT READ, and nothing may be concluded. The original habit's failure was not reading zero
as one of the two, it was reading zero without asking which.

docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced
byte-identically, which is this projection's positive control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* The specimen was seven rows and the class is four of them (review 58619)

completed_over_cost_requirement names claims that REACHED A VERDICT and were then reclassified
on cost. The floor says so in its own diagnostic — "reached its verdict and then exceeded its
budget ... cost=501ms EXACT ... This is a cost debt only — it is not a defect" — and the rows
carry outcome=completed_over_budget, which is to say they PASSED. Folding those three into a
class about gates that did NOT reach a judgment inflated the specimen by more than half and
contradicted a distinction the model draws deliberately.

Worse than the arithmetic: it was rung inflation of the same shape §4b(1) forbids, committed
inside a row whose subject is a gate reporting more than it established. I had the refuting
text in the log I quoted from and read past it.

The class is now the four INTERRUPTED-BEFORE-VERDICT rows, and the sentence that carries the
harm is sharper for the narrowing: four undecided rows were sufficient to refuse a run in
which zero claims failed. The three over-cost rows are retained only where they are honest —
as the second half of the nondeterminism observation, since both arms of the cost machinery
vary run to run on fixed bytes.

docs/design-ledgers.md regenerated by main_wet; every other rostered artifact reproduced
byte-identically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* Narrow the row to what is true at its own grain, and make the reroll mitigation the admitted arm

Four edits, three of them corrections to this PR and one discharging a condition the authority
already stated.

THE ROW OVERSTATED ITS OWN SUBJECT, and it was falsifiable from the log it cites. It said the
gate reports both outcomes "through one refusing channel, so a reader cannot tell a judgment
from a missing judgment". The floor prints INTERRUPTED-BEFORE-VERDICT and
COMPLETED-OVER-COST-REQUIREMENT as distinct typed diagnostics and carries them as separate
counters beside failed. A log reader can tell them apart perfectly. What cannot is everything
downstream of the fold — FloorRefused, the red check, the dashboard cell, the merge gate —
each receiving one bit whose only affordance is a reroll. So the defect is not a missing
distinction but a computed one erased on the way out, which is the worse shape: the
information exists and is discarded. Overstating this inside a row about a gate reporting more
than it established was the same failure twice.

THE COST HALF IS ALREADY ROSTERED AND IS NOW CITED RATHER THAN RE-DERIVED. That a claim's
measured cpu-ms is unstable on a shared runner is gunbc.rung_drop floor_cost_contention_verdict,
declared 2026-09-01, whose trigger is a claim-owned cost basis invariant across envelopes.
Filing it again would be a second authority over one fact.

AND THE MITIGATION WE HAVE ALL BEEN USING IS NOW THE ADMITTED ONE. That row ends by admitting
retry-until-green "only as a counted, visible mitigation carrying this row's trigger as its
dissolution condition". Rerolling has been in continuous bounded use across the board today —
one per head per signature, only on failed=0 — which is better than unbounded and was still
not the admitted arm, because nothing enumerated it. The receipt enumerates every instance BY
RUN ID, names the drop's own trigger as its dissolution condition, and states plainly that no
modeled producer counts them. No tally: this row has already had to retract one hand-derivation
described as a run product, and a count with no producer is stale at the next roll and
re-derivable by nobody. Whoever wants the number counts the citations.

The instances carry one observation finer than either the drop or the row had: after #10038's
live-gate cost repairs, self_host_compile_phase_live_gate_witness was ABSENT from attempt 1 and
BACK in attempt 2 of ONE head. Not merely less frequent — intermittent within a single head's
attempts, which is the sharpest statement that a cost repair moves incidence without touching
the mechanism at the boundary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* wip: bind specimen counters to attempt and job (jolly-hawk-122's clause, verified)

* Merge main, and bind each specimen counter to its attempt and job

The roster conflict was both sides appending to the same list tail — union, then checked
rather than assumed: 56 rostered identities against 56 declarations, none missing and none
orphaned. Neither side deleted a row, which is the case where union would have silently
re-added something deliberately removed.

DESIGN.md and docs/design-ledgers.md are generated, so they are regenerated from the merged
authority rather than hand-resolved. main_wet reproduced every other rostered artifact
byte-identically across main's changes to generated_artifact_emit and the workflow emissions,
which is this projection's positive control.

THE SPECIMEN COUNTERS NOW NAME THEIR INSTRUMENT. "First run" and "rerun" are ordinals that name
nothing and do not distinguish run 33604337589 from run 33628404336 on a later head. Each side
is now addressed by attempt AND job: attempt 1 is floor job 100172868685 (interrupted=4,
over_cost=3, FloorRefused), attempt 2 is job 100189043027 (0 and 0, green). Clause supplied by
session jolly-hawk-122, verified here against both jobs' logs before adoption.

AND IT NAMES THE COMMAND THAT MUST NOT BE USED TO RE-DERIVE IT, because the obvious one lies.
`gh run view --job 100172868685 --log` answers the ATTEMPT-1 job id with ATTEMPT 2's content —
its runner banner reads 09:03 where that job's own log begins 08:05, and it reports
interrupted_before_verdict=0, which are attempt 2's counters. Reproduced independently here.
A reader trusting it records 0 and 0 for both attempts, sees no disagreement, and destroys the
specimen this row is built on. Only `gh api .../actions/jobs/<job>/logs
--allow-escape-sequences` answers per job — and without that flag it writes zero bytes, which
is the sibling failure already rostered. Same call, both directions: empty on one flag, ~600 kB
of plausible wrong-subject log on the wrong subcommand. The wrong-content direction is the more
dangerous, and it is recorded where it protects the specimen rather than widening
empty_capture_read_as_clean_result past its authored grain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* Enumerate the remaining reroll instances, and name the instrument beside them

Four instances were spent or observed and not yet counted. An enumerated instance is what makes
this mitigation the arm the row admits rather than the one it forbids, so a spent roll left
unrecorded is the violation itself, not a bookkeeping lapse.

#10047 run 33622971872 attempt 2 — and the roster PREDICTED the row that blocked it: attempt 1
refused at 502ms on v2.test.emit.rust_binop_emit, a module carrying four identities in this
row's own attention subset.

#9986 at f5fca17 — two interrupted rows in compiler_frontend_program_status_witness and
self_host_compile_phase_frontier_witness, NEITHER in the live-gate family, on a head that had
already taken 2d76d9c. That is what establishes the arm is not confined to a repairable
family, and it refutes a prediction both this session and its manager made.

#10044 run 33628404336 attempts 1 and 2, jobs 100219422472 and 100256793010 — refuse then
refuse at ONE ROW EACH, failed=0 and planned=executed=3486 on both, and the row was
v2.test.emit.produced_decl_two_target on attempt 1 and v2.test.execution.emit_host_module_equals_eval
on attempt 2. At n=1 per side the arm did not re-refuse the same expensive claim; it drew a
different one. The population is redrawn per attempt rather than sampled from a fixed set of
costly rows — which is why family-by-family cost repair lowers incidence without bounding the
class, and why a green reroll is not evidence the refused row was wrong.

AND THE RECEIPT NOW NAMES THE COMMAND, because it enumerates run ids and therefore invites
re-derivation by exactly the reader most likely to hold the wrong instrument. `gh run view
--job <id> --log` answers an attempt-1 job id with attempt 2's content, so an auditor checking a
two-attempt specimen with it gets identical content on both sides, sees no disagreement, and
reports these instances as fabricated. It fails in the direction that discredits a true finding.
Only `gh api repos/OWNER/REPO/actions/jobs/JOB/logs --allow-escape-sequences` answers per job;
without the flag it writes zero bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

* The pair is suggestive; the instances jointly are what corroborate the redraw

Deliberately discarding a green head to fix one sentence, because the sentence is in the
artifact and the qualification was only in a PR comment.

WHAT WAS OVERSTATED. The receipt said the refuse-then-refuse pair on 03780b8 — one row
each, different identity — showed "the population is redrawn per attempt". Two draws with
different identities at n=1 per side are equally consistent with a FIXED set of marginal rows
sitting close enough to the deadline that ordering decides which crosses. Identity change alone
does not discriminate those explanations, and the row asserted the stronger one.

WHAT ACTUALLY DISCRIMINATES, and it needs the instances jointly rather than any one pair: the
COUNT moves as well as the membership — 4→0, 5→2, 1→1, 2→4, and 15. A fixed marginal set would
have to explain a count ranging over 0, 1, 2, 4, 5 and 15 AND the membership changing. Redraw
explains both; near-threshold ordering explains only the second. The load-bearing consequence is
unchanged on either reading: no enumeration of the expensive claims can be the population, so
family-by-family cost repair lowers incidence without bounding the class.

WHY NOT LAND FIRST AND FIX AFTER. The receipt is the durable artifact — it lists run ids and
invites re-derivation. A PR comment is not part of it, so on squash the qualification would stay
in a conversation nobody re-reads while the stronger claim shipped alone in the file. And the
asymmetry is bad in the wrong direction: this receipt's whole value is withstanding a skeptic
who re-derives it, and an auditor who finds one overstated sentence discounts the other five
instances too. Overclaiming the weakest link is what makes the strong links unreadable.

THE COST, STATED RATHER THAN ELIDED: d7f3ab0 was terminal-green on every required job —
build, floor, witnesses, rust-unit — with two approvals, and this discards all of it for a fresh
draw at the nondeterministic arm this PR documents. Caught by tidy-swift-334 against my own
evidence; the ruling to push before landing is theirs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeXMgoLPiVCvgAQbXZab5n

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot merged commit e35f848 into main Sep 2, 2026
10 of 12 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/eager-ferret-714 branch September 2, 2026 17:33
@briansrls
briansrls restored the session/eager-ferret-714 branch September 2, 2026 17:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant